Status
Where this honestly stands
Every number on this site comes from one file, with the date it was checked. This page is that file, read out. If something here is stale, the fault is ours and worth telling us about.
Measured on a 108 word everyday sweep. One answer is wrong and we name it. Verified 2026-08-11 against CODE-STATUS.md.
What works
- Nine tools, live
- Origin, root, meaning, formation, grammar, equivalents, and the two that grow the cache. All over one engine.
- Provenance on every claim
- Each field carries its source, that source’s evidential grade, and when it was retrieved.
- நூற்பா quoted at answer time
- Grammar claims quote the verse from the pinned editions rather than naming a chapter.
- Honest gaps
- A field no source can ground comes back as a gap. Nothing is filled in to look complete.
- Four front doors
- MCP for assistants, REST for apps, a browser page, and a command line. The browser page needed no engine changes.
- Open source, under a nonprofit
- Apache-2.0. Anyone may run it, and nothing is built on it commercially by us.
The measurement, in full
A sweep of 108 everyday words, re-run against the live build. It is the only honest read on quality we have, and the expected answers inside it are our own assessments rather than an authority, so a Tamil reader disagreeing with one of them is useful to us.
| What | Result | Reading |
|---|---|---|
| Origin classified correctly | 94/108 | Where the word came from, decided from orthography, the native parse and attestation. |
Returned an honest unknown | 11 | Mostly modern loans absent from every source we hold. A gap, not an error. |
| Wrong | 1 | One word. We know which one, and it is in the sweep output rather than hidden. |
| Formation decoded | 26/29 | Of the words in scope for பகுபத உறுப்பிலக்கணம். The rest are non-finite forms. |
| Automated tests | 266 | 5 skip without a live foma install. They include a check that every cited நூற்பா resolves in the pinned text, verbatim. |
| Classical verses pinned | 1948 | தொல்காப்பியம் 1486 · நன்னூல் 462, checksummed. |
Verified 2026-08-11. Source: CODE-STATUS.md.
What is not done
- Non-finite FST coverage
- The full புணரியல் sandhi engine
- Storage backend abstraction
- The morphological-lift evaluation, paused
- No CI in the code repo yet
- No release rung shipped: version is still 0.1.0
Nothing has been released yet. The version number is still 0.1.0 and no installable artifact exists, so at the moment the only way to run this is to clone it. That is the single biggest gap on this list, and it is the distribution problem.
Debts we say out loud
- Sandhi joins are sometimes unnamed
- Where no confident classical rule fires, the join is left unnamed rather than guessed. That is correct behaviour and it is below the standard the product should eventually meet.
- Origin is unknown more often than we would like
- Orthography can prove a word is not native. It can never say which language it came from. The remaining unknowns mostly need a lexicon we do not have.
- One word list has no licence
- It is graded D, declared in every answer that rests on it, and named on the sources page. The fix is an authenticated glossary, not a better disclaimer.
- The lift has not been measured
- Whether a model actually answers Tamil questions better with these tools attached is the number that would justify the project, and we do not have it yet. We will publish it whichever way it comes out.