Turkish and Kyrgyz/Kymorph article
Jump to navigation
Jump to search
Outline
Morphotactica
Morphophonologia
Corpora
- Which corpora to use?
- Wikipedia
- punktgen.py ky.crp.txt ky.pickle
- aq-wikicrp -x -t ky.pickle kywiki-20110923-pages-articles.xml kywp.xml
- Azattyk
- Wikipedia
- concerns
- Wikipedia is messy; should we have an automated cleaning process or get stats as-is?
- Use aq-wikicrp, this way it is reproducible .
- Wikipedia is messy; should we have an automated cleaning process or get stats as-is?
Numbers
wikipedia | azattyk | |
---|---|---|
num words | 271005 | |
xml file size | >3.8MB |