Kodava TakkA translator, a dictionary, and the corpus underneath them.
Kodava takk had no translation model, no parallel corpus, and no dataset on any public repository. This is the project that changed that — and the tools are free, in the open, and in use today.
What it does
Four tools, one purpose.
Translator
Type English, read Kodava. Built on a retrieval system that feeds a dictionary, grammar notes and verified sentence pairs into the model for every request — because there is no trained Kodava model to call. Disagree with a translation? Correct it in place.
Dictionary
About 14,000 words, searchable, with meanings, transliteration and the Kannada script form. Assembled from print sources that existed nowhere in machine-readable form until someone sat down and typed them.
Lessons
Structured practice for the person who understands their parents but can't answer back — the single most common Kodava story of the last two generations.
The corpus
Every correction, every suggested word, every verified pair goes into a growing parallel corpus. That corpus is the point: it is what a dedicated Kodava translation model will eventually be trained on.
The plan
How you build a translator for a language with no data.
You cannot train a model on a dataset that does not exist. So the first version is built to earn the dataset — and the training work now runs alongside the collecting rather than waiting for it to finish.
Built on a fraction of the usual data.
Gathering four million pairs in Kodava takk would ask something like a thousand native speakers for a great deal of their time — a tall order for a language this size. So we built with a fraction of it instead.
The trade-off is that some translations still need a second look, and that is where a fluent speaker makes more difference than any amount of engineering: every correction goes back in, and the next version is better for it.
If you speak Kodava takk and would like to be part of it, write to us at info.kodavame@gmail.com.
- Phase 1Live
Retrieval-powered translation
Dictionary, grammar and examples retrieved and passed to a large model at request time. Good enough to be useful today, and it doubles as the collection pipeline.
- Phase 2Under way
Community correction at scale
Native speakers correct what the translator produces. Target: tens of thousands of verified English–Kodava sentence pairs — the dataset that does not currently exist.
- Phase 3Under way
A dedicated Kodava model
Fine-tuning an open translation model on the collected corpus, in parallel with the collecting. Fast, cheap, self-hostable, and — for the first time — a machine that speaks Kodava without renting anyone's API.
The language
What Kodava takk actually is.
A South Dravidian language of its own — not a dialect of Kannada, though it is written in Kannada script by most people who write it at all. Its closest well-resourced relative is Kannada, which is what makes transfer learning plausible; its grammar, vocabulary and sound are its own.
It has been written in at least four ways: Kannada script, the Kodava Lipi adopted in 2022, the older Coorgi-Cox alphabet, and Latin transliteration — both the plain kind people type in family WhatsApp groups and the ISO 15919 standard. Our tools handle all of them, because our speakers use all of them.
At a glance
- ISO 639-3
- kfa
- Family
- Dravidian · South Dravidian I
- Speakers
- 114,000 – 200,000
- Scripts
- Kannada script · Kodava Lipi · Latin transliteration, ISO 15919
- Home
- Kodagu district, Karnataka
Next
Everything here is also in the app.
Arivāme carries the translator, the dictionary and the lessons on your phone — with the heritage library, the songs and the community alongside them.