The most frequently used 5,000 words in English, French, and Spanish, plus Chinese character frequency lists.
| # | Source | File | Status | Notes |
|---|---|---|---|---|
| 1 | TV and Movie Scripts (2006) | Removed | Contains patterns like "Ivy or ivy" as one entry, and like "W ." | |
| 2 | Wikipedia (2016) | wiki.txt | Included | Contains multi-word phrases like "the most" as single entries |
| 3 | hermitdave FrequencyWords | en_5k.txt | ✅ Good | License: CC BY-SA 4.0 |
Source #2 is based on: D. Goldhahn, T. Eckart & U. Quasthoff. "Building Large Monolingual Dictionaries at the Leipzig Corpora Collection: From 100 to 200 Languages." Proceedings of the 8th International Language Resources and Evaluation (LREC'12), 2012.
| # | Source | File | Status | Notes |
|---|---|---|---|---|
| 1 | Lexique 4 | lexique4_5000.tsv | ✅ Good | Generated via code/lexique4.sql. Has attributes. License: CC BY-SA |
| 2 | OpenSubtitles 5000 | opensubtitles_5000.txt | Included | Based on movie subtitles from opensubtitles.org, compiled by User:Hermitd |
Source #1 is based on: New, B., Pallier, C., Schalchli, G., Bourgin, J., & Gimenes, M. (2026). "Lexique 4: A major upgrade of the 'Lexique' French lexical database." Behavior Research Methods.
| # | Source | File | Status | Notes |
|---|---|---|---|---|
| 1 | Mixed 730K | mix_730k.txt | Included | Contains multi-word phrases. Based on Wortschatz Leipzig 2021 Wikipedia & 2022 News 1M Sentence corpora |
| 2 | Subtitles 10K (5 parts) | subtitles5k.txt | ✅ Good | Combined from raw files in es/subtitles5k_raw/ via code/combine_spannish.py. ~27.4 million words from movie/TV subtitles. Has lemma forms. License: GFDL & LGPL |
| 3 | hermitdave FrequencyWords | es_5k.txt | Included | License: CC BY-SA 4.0 |
Subtitles raw source links (source #2)
Source #1 is based on: D. Goldhahn, T. Eckart & U. Quasthoff. "Building Large Monolingual Dictionaries at the Leipzig Corpora Collection: From 100 to 200 Languages." Proceedings of the 8th International Language Resources and Evaluation (LREC'12), 2012.
| # | Source | File | Status | Notes |
|---|---|---|---|---|
| 1 | BLCU 25亿字语料汉字字频表 by 邢红兵 | blcu_5000.csv | Included | 仅供研究者进行汉字及相关研究之用 (for research use only) |
| 2 | MTSU 汉字单字字频总表 by 笪骏 | mtsu_5000.tsv | Included | Combined Classical and Modern Chinese. For research/teaching only; commercial use requires written permission |
| 3 | 通用规范汉字表 (一级字表) | hanzi_1_3500.txt | ✅ Good | 3,500 most common simplified Chinese characters (not ordered by frequency). Issued by State Council notice 国发〔2013〕23号 |
Source #3: 根据2013年6月18日《国务院关于公布〈通用规范汉字表〉的通知》(国发〔2013〕23号)印发 原文件.
code/lexique4.sql— Generatesfr/lexique4_5000.tsvfrom the Lexique 4 databasecode/combine_spannish.py— Combines raw Spanish subtitle files intoes/subtitles5k.txt