Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 3 additions & 1 deletion datasets.tsv
Original file line number Diff line number Diff line change
Expand Up @@ -115,6 +115,8 @@ Zalizniak-2020-SemanticShifts Anna Zalizniak and Anna Smirnitskaya and Maksim Ru
Mochizuki-2026-Sensorimotor Mochizuki, Masaya and Ota, Naoto 2026 ratings Japanese Japanese https://doi.org/10.3758/s13428-025-02939-1 Mochizuki2026 This list includes ratings of body-object interaction (BOI), i.e., the extent to which the word refers to an object or thing a human body can easily interact with, for 5,637 Japanese words. Items were rated on a 7-point scale (1 = low body–object interaction, 7 = high body–object interaction).
Dulay-2026a-AoA Dulay, Katrina May and Mirković, Jelena and Fua, Margaret Mary Rosary Carmel and Prabhu, Deeksha and Nag, Sonali 2026 ratings Filipino Filipino https://doi.org/10.3758/s13428-025-02876-z Dulay2026 This list contains age of acquisition (AoA) ratings for 885 Filipino words, collected from parents, teachers, and experts. In addition to the ratings, the original dataset includes the age band of first occurrence of a given word in a children’s book corpus, as well as expert ratings of the age suitability of the source books. While the original study was conducted with Filipino words, the present mapping is based on the English translations given for these items. The original study further includes AoA ratings for Kannada words, which can be found in a separate list, see Dulay-2026b-AoA.
Dulay-2026b-AoA Dulay, Katrina May and Mirković, Jelena and Fua, Margaret Mary Rosary Carmel and Prabhu, Deeksha and Nag, Sonali 2026 ratings Kannada Kannada https://doi.org/10.3758/s13428-025-02876-z Dulay2026 This list contains age of acquisition (AoA) ratings for 885 Kannada words, collected from parents, teachers, and experts. In addition to the ratings, the original dataset includes the age band of first occurrence of a given word in a children’s book corpus, as well as expert ratings of the age suitability of the source books. While the original study was conducted with Kannada words, the present mapping is based on the English translations given for these items. The original study further includes AoA ratings for Filipino words, which can be found in a separate list, see Dulay-2026a-AoA.
List-2018-Colexifications List, Johann-Mattis and Greenhill, Simon J. and Anderson, Cormac and Mayer, Thomas and Tresoldi, Tiago and Forkel, Robert 2018 relations English Global http://clics.clld.org List2018 This is the concept list underlying the second release of the [Database of Cross-Linguistic Colexifications (CLICS2)](http://clics.clld.org) (for the first version see, [List 2014](:bib:List2014). Based on the CLDF framework, 15 data sets were included in the second version of CLICS. The data comprises 1220 language varieties and 2486 distinct Concepticon concept sets. Instead of providing full graphs for the networks, we provide links to the major clusters for each concept in CLICS. The column ``CENTRAL_CONCEPT`` lists the most central node in each of the community clusters in CLICS2. The community identifier is listed in column ``COMMUNITY``. The columns ``FAMILY_FREQUENCY``, ``LANGUAGE_FREQUENCY`` and ``WORD_FREQUENCY`` list the number of times a translation for a given concept occurs in the CLICS2 data. The columns ``DEGREE``, ``WEIGHTED_FAMILY_DEGREE`` and ``WEIGHTED_LANGUAGE_DEGREE`` contain the unweighted and weighted degree for each node. List-2018-1105
Rzymski-2020-Colexifications Christoph Rzymski and Tiago Tresoldi and Simon J. Greenhill and Mei-Shin Wu and Nathanael E. Schweikhard and Maria Koptjevskaja-Tamm and Volker Gast and Timotheus A. Bodt and Abbie Hantgan and Gereon A. Kaiping and Sophie Chang and Yunfan Lai and Natalia Morozova and Heini Arjava and Nataliia Hübler and Ezequiel Koile and Steve Pepper and Mariann Proos and Briana Van Epps and Ingrid Blanco and Carolin Hundt and Sergei Monakhov and Kristina Pianykh and Sallona Ramesh and Russell D. Gray and Robert Forkel and Johann-Mattis List 2020 relations English Global http://clics.clld.org Rzymski2020 This is the concept list underlying the third release of the [Database of Cross-Linguistic Colexifications (CLICS3)](http://clics.clld.org) (for the first version see, [List 2014](:bib:List2014). The data comprises 3156 language varieties and 2906 distinct Concepticon concept sets. Instead of providing full graphs for the networks, we provide links to the major clusters for each concept in CLICS. The column ``CENTRAL_CONCEPT`` lists the most central node in each of the community clusters in CLICS3. The community identifier is listed in column ``COMMUNITY``. The columns ``FAMILY_FREQUENCY``, ``LANGUAGE_FREQUENCY`` and ``WORD_FREQUENCY`` list the number of times a translation for a given concept occurs in the CLICS3 data. The columns ``DEGREE``, ``WEIGHTED_FAMILY_DEGREE`` and ``WEIGHTED_LANGUAGE_DEGREE`` contain the unweighted and weighted degree for each node. Rzymski-2020-1624
Wang-2026-Utility Wang, Andrew and Brysbaert, Marc and Günther, Fritz 2026 ratings English English https://doi.org/10.3758/s13428-026-03051-8 Wang2026 This list contains ratings of utility of 80,000+ English words and multiword expressions (MWE) given on a best-worst scale (BWS). Participants were asked to select the most useful and the least useful word or multiword expression out of a 6-tuple. Utility scores were computed using the Value Learning algorithm [(Hollis 2018](:bib:Hollis2018). For each item, the number of times it was selected as “most useful” was summed and the number of times it was selected as “least useful” was subtracted. This best–worst difference score was then scaled across all items to a 0–1 range. The list also includes familiarity scores by an LLM (GPT-4). Originally, these scores were given on a 7-point scale (1 = completely unfamiliar, 7 = completely familiar), however only items with a familiarity score >3.5 were included in the analysis with the exception of a small number of catch trials that fall below this threshold. Further, some items exceed the threshold of 7 slightly. Multilex frequencies are provided in the original dataset but not included here.
Jiang-2026-Frequency Jiang, Zehua R. and Siew, Cynthia S.Q. 2026 norms Singapore English Singapore English https://doi.org/10.3758/s13428-026-03012-1 Jiang2026 This list contains frequency and contextual diversity measures for words in the Singapore English National Speech Corpus (NSC) [(Koh et al. 2019)](:bib:Koh2019).
Song-2026-Emotions Song, Dangui 2026 ratings Chinese Chinese https://doi.org/10.3758/s13428-026-03010-3 Song2026 This list contains ratings for Chinese words on ten discrete emotion categories: happiness, anger, sadness, fear, disgust, anxiety, surprise, contentment, amusement and serenity. Participants were native speakers of Chinese and rated words on a 5-point scale (1 = not at all, 5 = extremely).
Song-2026-Emotions Song, Dangui 2026 ratings Chinese Chinese https://doi.org/10.3758/s13428-026-03010-3 Song2026 This list contains ratings for Chinese words on ten discrete emotion categories: happiness, anger, sadness, fear, disgust, anxiety, surprise, contentment, amusement and serenity. Participants were native speakers of Chinese and rated words on a 5-point scale (1 = not at all, 5 = extremely).
Loading