Knowledge collections and datasets (English): Difference between revisions
Jump to navigation
Jump to search
No edit summary |
|||
| Line 13: | Line 13: | ||
* [http://wordnet.princeton.edu/ WordNet] | * [http://wordnet.princeton.edu/ WordNet] | ||
* [http://www.cs.technion.ac.il/~gabr/resources/data/wordsim353/wordsim353.html WordSimilarity-353 Test Collection] | * [http://www.cs.technion.ac.il/~gabr/resources/data/wordsim353/wordsim353.html WordSimilarity-353 Test Collection] | ||
* [[TEASE]] - Acquisition of Entailment Relations from the Web | |||
== Additional Dataset Collections == | == Additional Dataset Collections == | ||
Revision as of 14:20, 12 December 2006
Datasets for Computational Linguistics and Natural Language Processing.
- Clustering by Committee - terms clustered and organized using the Distributional Hypothesis
- DIRT Paraphrase Collection - Discovery of Inference Rules from Text
- Edinburgh Associative Thesaurus (EAT)
- FrameNet
- MRC Psycholinguistic Database
- Noun Compound Repository
- Reuters-21578 Text Categorization Collection
- Spam filtering datasets
- University of South Florida Free Association Norms
- VerbOcean - verbs organized by semantic relation, including temporal precedence and strength
- WordNet
- WordSimilarity-353 Test Collection
- TEASE - Acquisition of Entailment Relations from the Web