Difference between revisions of "Languages of the Volga-Kama region"
(10 intermediate revisions by 3 users not shown) | |||
Line 31: | Line 31: | ||
|| [[HFST|HFST (lexc+twol)]] |
|| [[HFST|HFST (lexc+twol)]] |
||
|| development |
|| development |
||
|align="right"| {{#lst:Apertium-myv-fin/stats| |
|align="right"| {{#lst:Apertium-myv-fin/stats|myv_stems}} |
||
|align="center"| |
|align="center"| |
||
|| [[apertium-myv-fin]] ([[incubator]]) |
|| [[apertium-myv-fin]] ([[incubator]]) |
||
Line 41: | Line 41: | ||
|| <code>tat</code> |
|| <code>tat</code> |
||
|| [[HFST|HFST (lexc+twol)]] |
|| [[HFST|HFST (lexc+twol)]] |
||
|| {{#lst:Apertium-tat/stats|state}} |
|||
|| working |
|||
|align="right"| {{#lst:Apertium-tat/stats|stems}} |
|align="right"| {{#lst:Apertium-tat/stats|stems}} |
||
|align="center"| [[Apertium-tat#Current_State|~{{:Apertium-tat/stats/average}}%]] |
|align="center"| [[Apertium-tat#Current_State|~{{:Apertium-tat/stats/average}}%]] |
||
|| {{#lst:Apertium-tat/stats|location}} |
|||
|| [[apertium-tat]] ([[languages]]) |
|||
|| {{#lst:Apertium-tat/stats|authors}} |
|||
|| [[User:Ilnar.salimzyan|Ilnar]], [[User:Francis Tyers|Fran]], [[User:Firespeaker|Jonathan]], Milli |
|||
|- |
|- |
||
| <code>[[apertium-chv]]</code> |
| <code>[[apertium-chv]]</code> |
||
Line 75: | Line 75: | ||
|| [[HFST|HFST (lexc+twol)]] |
|| [[HFST|HFST (lexc+twol)]] |
||
|| development |
|| development |
||
|align="right"| {{#lst:Apertium-mrj-fin/stats| |
|align="right"| {{#lst:Apertium-mrj-fin/stats|mrj_stems}} |
||
|align="center"| |
|align="center"| |
||
|| [[apertium-mrj-fin]] ([[incubator]]) |
|| [[apertium-mrj-fin]] ([[incubator]]) |
||
Line 86: | Line 86: | ||
|| [[HFST|HFST (lexc+twol)]] |
|| [[HFST|HFST (lexc+twol)]] |
||
|| prototype |
|| prototype |
||
|align="right"| {{#lst:Apertium-udm-rus/stats| |
|align="right"| {{#lst:Apertium-udm-rus/stats|udm_stems}} |
||
|align="center"| |
|align="center"| |
||
|| |
|| [[apertium-udm-rus]] ([[nursery]]) |
||
|| [[User:Francis_Tyers|Fran]], [[User:Trondtr|Trond]], [[User:Andrewboltachev|Andrey]], Лукерья, Алексей |
|| [[User:Francis_Tyers|Fran]], [[User:Trondtr|Trond]], [[User:Andrewboltachev|Andrey]], Лукерья, Алексей |
||
|- |
|- |
||
Line 97: | Line 97: | ||
|| [[HFST|HFST (lexc+twol)]] |
|| [[HFST|HFST (lexc+twol)]] |
||
|| prototype |
|| prototype |
||
|align="right"| {{#lst:Apertium-kpv-mhr/stats| |
|align="right"| {{#lst:Apertium-kpv-mhr/stats|kpv_stems}} |
||
|align="center"| |
|align="center"| |
||
|| [[apertium-kpv-mhr |
|| [[apertium-kpv-mhr]] ([[incubator]]) |
||
|| [[User:Francis_Tyers|Fran]], [[User:Trondtr|Trond]], Fedina, Andrei Chemyshev |
|| [[User:Francis_Tyers|Fran]], [[User:Trondtr|Trond]], Fedina, Andrei Chemyshev |
||
|- |
|- |
||
Line 108: | Line 108: | ||
|| [[HFST|HFST (lexc+twol)]] |
|| [[HFST|HFST (lexc+twol)]] |
||
|| prototype |
|| prototype |
||
|align="right"| {{#lst:Apertium-kpv-mhr/stats| |
|align="right"| {{#lst:Apertium-kpv-mhr/stats|mhr_stems}} |
||
|align="center"| |
|align="center"| |
||
|| [[apertium-kpv-mhr]] ([[incubator]]) |
|| [[apertium-kpv-mhr]] ([[incubator]]) |
||
Line 115: | Line 115: | ||
=== Existing language pairs === |
=== Existing language pairs === |
||
==== Volga-Kama–Volga-Kama pairs ==== |
|||
Text in ''italic'' denotes language pairs under development / in the incubator. Regular text denotes a functioning language pair in staging, while text in '''bold''' denotes a stable well-working language pair in trunk. |
Text in ''italic'' denotes language pairs under development / in the incubator. Regular text denotes a functioning language pair in staging, while text in '''bold''' denotes a stable well-working language pair in trunk. |
||
{| style="text-align: center;" class="wikitable" |
{| style="text-align: center;" class="wikitable dixtable" |
||
|- style="background: #ececec" |
|- style="background: #ececec" |
||
! |
! !! tat !! chv !! bak !! mrj !! udm !! mhr !! myv !! kpv |
||
|- |
|- |
||
| '''tat''' || - |
| '''tat''' || - || ''[[Apertium-chv-tat|chv-tat]]''<br>{{#lst:Apertium-chv-tat/stats|chv-tat_stems}} || [[Apertium-tat-bak|tat-bak]]<br>{{#lst:Apertium-tat-bak/stats|tat-bak_stems}} || || || || || |
||
|- |
|||
| '''chv''' || ''[[apertium-cv-tt|cv-tt]]'' || - || || || || || |
|||
|- |
|||
| '''bak''' || ''[[apertium-tat-bak|tat-bak]]'' || || - || || || || |
|||
|- |
|||
| '''udm''' || || || || - || || || |
|||
|- |
|||
| '''mhr''' || || || || || - || || ''[[apertium-kpv-mhr|kpv-mhr]]'' |
|||
|- |
|||
| '''myv''' || || || || || || - || |
|||
|- |
|||
| '''kpv''' || || || || || ''[[apertium-kpv-mhr|kpv-mhr]]'' || || - |
|||
|} |
|||
==== Pairs with non–Volga-Kama languages ==== |
|||
{| style="text-align: center;" class="wikitable" |
|||
|- style="background: #ececec" |
|||
! !! tat !! chv !! bak !! udm !! mhr !! myv !!kpv |
|||
|- |
|||
| '''ru''' || ''[[apertium-tt-ru|tt-ru]]'' || ''[[apertium-cv-ru|cv-ru]]'' || || ''[[apertium-udm-rus|udm-rus]]'' || || || |
|||
|- |
|||
| '''ky''' || ''[[apertium-tt-ky|tt-ky]]'' || || || || || || |
|||
|- |
|||
| '''tr''' || || ''[[apertium-cv-tr|cv-tr]]'' || || || || || |
|||
|- |
|||
| '''fin''' || || || || ''[[apertium-fin-udm|fin-udm]]'' || || ''[[apertium-myv-fin|myv-fin]]'' || ''[[apertium-kpv-fin|kpv-fin]]'' |
|||
|} |
|||
==== Table of dix progress ==== |
|||
{| style="text-align: center;" class="wikitable" |
|||
|- style="background: #ececec" |
|||
! !! tat !! chv !! bak !! udm !! mhr !! myv !! kpv |
|||
|- |
|- |
||
| ''' |
| '''chv''' || ''[[Apertium-chv-tat|chv-tat]]''<br>{{#lst:Apertium-chv-tat/stats|chv-tat_stems}} || - || || || || || || |
||
|- |
|- |
||
| ''' |
| '''bak''' || [[Apertium-tat-bak|tat-bak]]<br>{{#lst:Apertium-tat-bak/stats|tat-bak_stems}} || || - || || || || || |
||
|- |
|- |
||
| ''' |
| '''mrj''' || || || || - || || || || |
||
|- |
|- |
||
| '''udm''' || |
| '''udm''' || || || || || - || || || |
||
|- |
|- |
||
| '''mhr''' || || || || |
| '''mhr''' || || || || || || - || || ''[[Apertium-kpv-mhr|kpv-mhr]]''<br>{{#lst:Apertium-kpv-mhr/stats|kpv-mhr_stems}} |
||
|- |
|- |
||
| '''myv''' || || |
| '''myv''' || || || || || || || - || |
||
|- |
|- |
||
| '''kpv''' || |
| '''kpv''' || || || || || || ''[[Apertium-kpv-mhr|kpv-mhr]]''<br>{{#lst:Apertium-kpv-mhr/stats|kpv-mhr_stems}} || || - |
||
|- |
|- |
||
| || || |
| || || || || || || || || |
||
|- |
|- |
||
| ''' |
| '''fin''' || || || || ''[[Apertium-mrj-fin|mrj-fin]]''<br>{{#lst:Apertium-mrj-fin/stats|mrj-fin_stems}} || ''[[Apertium-fin-udm|fin-udm]]''<br>{{#lst:Apertium-fin-udm/stats|fin-udm_stems}} || || ''[[Apertium-myv-fin|myv-fin]]''<br>{{#lst:Apertium-myv-fin/stats|myv-fin_stems}} || ''[[Apertium-kpv-fin|kpv-fin]]''<br>{{#lst:Apertium-kpv-fin/stats|kpv-fin_stems}} |
||
|- |
|- |
||
| ''' |
| '''kaz''' || '''[[Apertium-kaz-tat|kaz-tat]]'''<br>'''{{#lst:Apertium-kaz-tat/stats|kaz-tat_stems}}''' || || || || || || || |
||
|- |
|- |
||
| ''' |
| '''kir''' || ''[[Apertium-tat-kir|tat-kir]]''<br>{{#lst:Apertium-tat-kir/stats|tat-kir_stems}} || || || || || || || |
||
|- |
|- |
||
| |
| '''rus''' || [[Apertium-tat-rus|tat-rus]]<br>{{#lst:Apertium-tat-rus/stats|tat-rus_stems}} || ''[[Apertium-cv-ru|cv-ru]]''<br>{{#lst:Apertium-cv-ru/stats|cv-ru_stems}} || || || [[Apertium-udm-rus|udm-rus]]<br>{{#lst:Apertium-udm-rus/stats|udm-rus_stems}} || || || |
||
|- |
|- |
||
| |
| '''tur''' || [[Apertium-tur-tat|tur-tat]]<br>{{#lst:Apertium-tur-tat/stats|tur-tat_stems}} || ''[[Apertium-cv-tr|cv-tr]]''<br>{{#lst:Apertium-cv-tr/stats|cv-tr_stems}} || || || || || || |
||
|} |
|} |
||
Line 200: | Line 166: | ||
=== Volga-Kama language vulnerability === |
=== Volga-Kama language vulnerability === |
||
The following table shows information about Volga-Kama varieties |
The following table shows information about Volga-Kama varieties. |
||
{|class="wikitable sortable" |
{|class="wikitable sortable" |
Latest revision as of 23:25, 22 December 2014
The languages of the Volga-Kama region include several Turkic and Uralic languages spoken in the Volga-Kama region (along the Volga and Kama rivers) in Russia. These include [varieties of] Tatar, Bashqort, Chuvash, Mari, Komi, Mordvin, and Udmurt (and linguistically, to some extent, Russian).
The master plan involves generating independent finite-state transducers for each language, and then making individual dictionaries and transfer rules for every pair. The current status of these goals is listed below.
Status[edit]
The ultimate goal is to have multi-purposable transducers for a variety of Volga-Kama languages. These can then be paired for X→Y translation with the addition of a CG for language X and transfer rules / dictionary for the pair X→Y. Below is listed development progress for each language's transducers and dictionary pairs.
Transducers[edit]
Once a transducer has ~80% coverage on a range of medium-large corpora we can say it is "working". Over 90% and it can be considered to be "production".
Existing language pairs[edit]
Text in italic denotes language pairs under development / in the incubator. Regular text denotes a functioning language pair in staging, while text in bold denotes a stable well-working language pair in trunk.
tat | chv | bak | mrj | udm | mhr | myv | kpv | |
---|---|---|---|---|---|---|---|---|
tat | - | chv-tat 198 |
tat-bak 2,941 |
|||||
chv | chv-tat 198 |
- | ||||||
bak | tat-bak 2,941 |
- | ||||||
mrj | - | |||||||
udm | - | |||||||
mhr | - | kpv-mhr 127 | ||||||
myv | - | |||||||
kpv | kpv-mhr 127 |
- | ||||||
fin | mrj-fin 273 |
fin-udm 93 |
myv-fin 401 |
kpv-fin 1 | ||||
kaz | 'kaz-tat ' |
|||||||
kir | tat-kir |
|||||||
rus | tat-rus 5,999 |
cv-ru 75 |
udm-rus 148 |
|||||
tur | tur-tat 3,317 |
cv-tr 100 |
The languages[edit]
Volga-Kama languages by subgroup[edit]
- Uralic → Finno-Ugric → Finno-Permic
Volga-Kama language vulnerability[edit]
The following table shows information about Volga-Kama varieties.
language | iso | num speakers | UNESCO classification |
---|---|---|---|
Tatar | tat |
6500K | 0. none |
Bashqort | bak |
1379K | 1. vulnerable |
Chuvash | chv |
1325K | 1. vulnerable |
Udmurt | udm |
0464K | 2. definitely endangered |
Mari - Eastern | mhr |
0414K | 2. definitely endangered |
Mordvin - Erzya | myv |
0400K | 2. definitely endangered |
Komi - Zyryan | kpv |
0217K | 2. definitely endangered |
Mordvin - Moksha | mdf |
0200K | 2. definitely endangered |
Komi - Permyak | koi |
0094K | 2. definitely endangered |
Mari - Western | mrj |
0037K | 3. severely endangered |
Komi - Yazva | koi |
0000K | 3. severely endangered |
Existing general resources[edit]
Grammars[edit]
Dictionaries[edit]
Existing computational resources[edit]
Corpora and corpora projects[edit]
Spell-checkers[edit]
Text-to-speech and speech-to-text systems[edit]
Keyboards[edit]
- Xkb includes keyboards for the following languages:
- Tatar
- Chuvash
- ...?