Layer 3
மூலம் sources
What this project grounds on, what we still need, and what we may not use. Every source answers two questions that are kept apart on purpose, because collapsing them causes mistakes in both directions.
Two questions, never one
How good is the evidence?
- A
- Primary authority — the classical text itself, or a rule table derived from it and carrying its நூற்பா.
- B
- Scholarly or institutional publication — peer-reviewed, or published by a government body; named editors, stable citation, versioned release.
- C
- Community-edited and revisable — anyone may change it, so a claim is only as good as the revision it was read from. Citable per revision, not per work.
- D
- Unattributed compilation — no named editor, no stated terms, no active maintenance. Usable as attested evidence, never as authority.
What may we do with it?
- redistribute
- Ship the data as a version-locked artifact in the public repo.
- serve-with-attribution
- Cache and serve; attribution travels with every claim; never relicensed.
- consult-and-cite
- Consult to establish a FACT; store the fact plus the citation, never the source's own wording. Excluded from gold-corpus export.
The Madras Tamil Lexicon is why these are separate axes. It is the most authoritative lexicon in existence for Tamil and the most restrictively licensed. Grade A, and consult-and-cite. Both true at once, and a system that had only one field for "how good is this source" would have to lie about one of them.
A grade D source is fine to ship, as long as the answer says it is grade D. Authenticity comes from every claim carrying a graded, citable provenance a reader can check, rather than from every source being impeccable.
The ledger
Eleven sources, copied from the server's own registry. Where a source has its own page, the name links to it.
| Source | Grade | Tier | What we may do | Licence |
|---|---|---|---|---|
| Tholkappiyam | A | anchor | redistribute | The work is public domain (classical). The pinned Project Madurai etext is distributable provided its header is kept intact. |
| Nannūl | A | anchor | redistribute | The work is public domain (classical). The pinned Project Madurai etext is distributable provided its header is kept intact. |
| thamizh-mcp cited rule tables | A | anchor | redistribute | Apache-2.0 (our own work) |
| ThamizhiMorph | B | anchor | redistribute | Apache-2.0 |
| thamizh-mcp curated verb paradigms | B | anchor | redistribute | Apache-2.0 (our own work) |
| Google Dakshina (attested romanizations) | B | anchor | serve-with-attribution | CC BY-SA 4.0 — inherited from Dakshina. NOT Apache-2.0 and never relicensed. (The dwyl/english-words wordlist component is Unlicense.) |
| Tamil Wiktionary | C | evolving | serve-with-attribution | CC BY-SA 4.0 / GFDL |
| English Wiktionary (etymology) | C | evolving | serve-with-attribution | CC BY-SA 4.0 / GFDL |
| Sanskrit-To-Pure-Tamil (community தனித்தமிழ் lists) | D | evolving | serve-with-attribution | UNSTATED — no LICENSE file and no terms in the README. |
| TVA கலைச்சொல் | B | anchor | consult-and-cite | Government of Tamil Nadu / Tamil Virtual Academy — no redistribution grant located. |
| Madras Tamil Lexicon | A | anchor | consult-and-cite | CC BY-NC-ND 2.0 — © University of Madras (original 1924–1936); DSAL digitization refreshed September 2023. |
Registry last updated 2026-08-11, copied here 2026-08-19 by
scripts/sync-sources.py. It is not retold, so it cannot read softer here than it
does in the code.
The one licence gap we ship
The தனித்தமிழ் word lists we use for native equivalents have no stated licence. No LICENSE file, no terms in the README, and no upstream commit since 2020. An earlier claim that they were MIT was withdrawn when nobody could find the basis for it.
We could quietly drop the sentence and most readers would never know. It stays because the whole argument of this project is that a claim you cannot check is worth nothing, and that has to apply to our own claims first. The fix is not a better disclaimer. It is an authenticated glossary, which is exactly what we are asking the Tamil Virtual Academy for.
What we are asking for
Three open asks. Each has a page saying what we want, what we would do with it, and what we would give back, so that a letter can link one address instead of explaining itself.
தமிழ் இணையக் கல்விக்கழகம் Tamil Virtual Academy
The கலைச்சொல் glossaries, in machine-readable form, and a licence statement for the archive.
It would retire the one genuine licence gap this project ships.
ஆலமரம் Aalamaram
Where the treebank is distributed, and under what terms.
Ten thousand annotated sentences would let us check our morphology at scale instead of one word at a time.
ILAKKANAM
The benchmark dataset, or word when it publishes.
It is how we would measure whether any of this actually helps.
If you hold one of these, or know who does, we would rather hear from you than keep guessing. Write to thamizh@ief-global.org.
Closed, with the reason published
The Madras University Tamil Lexicon is the source we most want and cannot have. Publishing why, rather than staying vague about which lexicon we consulted, is the same discipline as returning an honest gap.