தமிழ்AI

Status

Where this honestly stands

Every number on this site comes from one file, with the date it was checked. This page is that file, read out. If something here is stale, the fault is ours and worth telling us about.

9
tools live
266
tests passing
94/108
origin correct
11
honest unknowns

Measured on a 108 word everyday sweep. One answer is wrong and we name it. Verified 2026-08-11 against CODE-STATUS.md.

What works

Nine tools, live
Origin, root, meaning, formation, grammar, equivalents, and the two that grow the cache. All over one engine.
Provenance on every claim
Each field carries its source, that source’s evidential grade, and when it was retrieved.
நூற்பா quoted at answer time
Grammar claims quote the verse from the pinned editions rather than naming a chapter.
Honest gaps
A field no source can ground comes back as a gap. Nothing is filled in to look complete.
Four front doors
MCP for assistants, REST for apps, a browser page, and a command line. The browser page needed no engine changes.
Open source, under a nonprofit
Apache-2.0. Anyone may run it, and nothing is built on it commercially by us.

The measurement, in full

A sweep of 108 everyday words, re-run against the live build. It is the only honest read on quality we have, and the expected answers inside it are our own assessments rather than an authority, so a Tamil reader disagreeing with one of them is useful to us.

WhatResultReading
Origin classified correctly 94/108 Where the word came from, decided from orthography, the native parse and attestation.
Returned an honest unknown 11 Mostly modern loans absent from every source we hold. A gap, not an error.
Wrong 1 One word. We know which one, and it is in the sweep output rather than hidden.
Formation decoded 26/29 Of the words in scope for பகுபத உறுப்பிலக்கணம். The rest are non-finite forms.
Automated tests 266 5 skip without a live foma install. They include a check that every cited நூற்பா resolves in the pinned text, verbatim.
Classical verses pinned 1948 தொல்காப்பியம் 1486 · நன்னூல் 462, checksummed.

Verified 2026-08-11. Source: CODE-STATUS.md.

What is not done

  • Non-finite FST coverage
  • The full புணரியல் sandhi engine
  • Storage backend abstraction
  • The morphological-lift evaluation, paused
  • No CI in the code repo yet
  • No release rung shipped: version is still 0.1.0

Nothing has been released yet. The version number is still 0.1.0 and no installable artifact exists, so at the moment the only way to run this is to clone it. That is the single biggest gap on this list, and it is the distribution problem.

Debts we say out loud

Sandhi joins are sometimes unnamed
Where no confident classical rule fires, the join is left unnamed rather than guessed. That is correct behaviour and it is below the standard the product should eventually meet.
Origin is unknown more often than we would like
Orthography can prove a word is not native. It can never say which language it came from. The remaining unknowns mostly need a lexicon we do not have.
One word list has no licence
It is graded D, declared in every answer that rests on it, and named on the sources page. The fix is an authenticated glossary, not a better disclaimer.
The lift has not been measured
Whether a model actually answers Tamil questions better with these tools attached is the number that would justify the project, and we do not have it yet. We will publish it whichever way it comes out.

What comes after this →