Resources for Bulgarian: Difference between revisions
Jump to navigation
Jump to search
0 |
HamleDT |
||
| (3 intermediate revisions by 2 users not shown) | |||
| Line 2: | Line 2: | ||
===Free software=== | ===Free software=== | ||
* [https://apertium.svn.sourceforge.net/svnroot/apertium/trunk/apertium-mk-bg apertium-mk-bg] RBMT system between Macedonian and Bulgarian | |||
===Proprietary=== | ===Proprietary=== | ||
| Line 17: | Line 19: | ||
== Grammars == | == Grammars == | ||
===Proprietary=== | ===Proprietary=== | ||
* [http://dcl.bas.bg/BulNet/general_en.html BulNet WordNet] (21,444 synonym sets) | * [http://dcl.bas.bg/BulNet/general_en.html BulNet WordNet] (21,444 synonym sets) | ||
* [[Generation grammars|KPML generation grammar]] | |||
==Corpora== | ==Corpora== | ||
| Line 27: | Line 29: | ||
===Free=== | ===Free=== | ||
* [http://www.statmt.org/setimes/ Southeast European Times] | * [http://www.statmt.org/setimes/ Southeast European Times], sentence aligned corpus, Albanian, Bulgarian, English, Greek, Macedonian, Romanian, Serbo-Croatian, Turkish — approximately 4.5 million words per language | ||
* [http://www.statmt.org/europarl Europarl corpus], sentence aligned with English | |||
* [http://ufal.mff.cuni.cz/hamledt HamleDT], harmonized dependency treebanks of many languages, common annotation style. | |||
===Proprietary=== | ===Proprietary=== | ||
Latest revision as of 15:36, 26 May 2014
Machine translation systems
Free software
- apertium-mk-bg RBMT system between Macedonian and Bulgarian
Proprietary
Lexical resources
Morphological analysis
Free software
- Morphological analyser 8,581 lemmata, ~88% coverage over SETimes
Proprietary
Grammars
Proprietary
- BulNet WordNet (21,444 synonym sets)
- KPML generation grammar
Corpora
Free
- Southeast European Times, sentence aligned corpus, Albanian, Bulgarian, English, Greek, Macedonian, Romanian, Serbo-Croatian, Turkish — approximately 4.5 million words per language
- Europarl corpus, sentence aligned with English
- HamleDT, harmonized dependency treebanks of many languages, common annotation style.