தமிழ்AI

Layer 3 · மூலம்

Madras Tamil Lexicon

The most authoritative Tamil lexicon in existence, compiled at the University of Madras between 1924 and 1936. We may consult it and cite it. We may not bundle a line of it, and we have not.

Closed, with a reason Grade A consult-and-cite anchor tier

Compiled

University of Madras, 1924–1936

Digitised by

Digital South Asia Library, University of Chicago

Licence

CC BY-NC-ND 2.0

Our status

Adapter exists as a stub and is off. Nothing has been scraped.

Two separate walls

The licence. CC BY-NC-ND has two clauses that both bite. NonCommercial bites because our code is Apache-2.0 and anybody may run this server commercially, which we cannot prevent and would not want to. NoDerivatives bites because a restructured, indexed copy inside a knowledge store is plausibly a derivative work, and "plausibly" is not a standard we are willing to ship on.

The robots file. The only query endpoint is disallowed to automated clients. So even the narrow reading, look a word up at query time and store nothing but the fact and the citation, is not available to us without asking.

We noticed the rule and did not scrape. That is worth stating plainly, because it is the difference between a request and a fait accompli, and because the alternative is available to anyone with an afternoon and no scruples.

Why this source is the reason for our whole grading model

This lexicon is grade A and consult-and-cite at the same time. It is the most authoritative source we could possibly have and the most restrictively licensed one we hold.

A system with a single field for "how good is this source" has to lie about one of those. So ours has two: evidential grade, and what we are legally allowed to do with the bytes. They move independently. That design came directly out of running into this entry, and it later caught a more embarrassing case in the other direction, a word list we depended on whose licence turned out to be nothing at all.

What we asked, and what we already know

A letter to the Digital South Asia Library is drafted, with one question and one request:

  1. Does the digitised edition carry machine-readable etymology fields? We cannot check, because those pages are the ones we are not fetching. This question matters more than the permission does.
  2. Permission for rate-limited programmatic queries, storing derived claims and citations only, never the lexicon's own wording.

The first question is first for a reason. We assessed the other digitisation of the same lexicon and found glosses without etymology. If this edition is the same, permission buys us very little, and it would be a waste of an institution's goodwill to ask for something we do not need.

Meanwhile the honest position holds: where a borrowed word carries no orthographic marker, this project returns unknown rather than a plausible etymology. Some of those unknowns would close immediately if this door opened. thamizh@ief-global.org

What the registry records

This is copied from the server's own source registry rather than restated, so it cannot be softer here than it is in the code.

Licence
CC BY-NC-ND 2.0 — © University of Madras (original 1924–1936); DSAL digitization refreshed September 2023.
What we may do with it
Consult to establish a FACT; store the fact plus the citation, never the source's own wording. Excluded from gold-corpus export.
Evidential grade
Primary authority — the classical text itself, or a rule table derived from it and carrying its நூற்பா. Confidence is capped at 0.95.
Pin
not vendored — query-time only, by licence
Maintenance
digitization maintained by DSAL; the lexicon itself is a fixed historical work
Attribution
Madras University Tamil Lexicon, dsal.uchicago.edu/dictionaries/tamil-lex/.
Status in the server
stub — raises NotImplementedError

⚠️ GRADE A AND consult-and-cite AT THE SAME TIME — this entry is the reason grade and redistribution are separate axes. It is the most authoritative lexicon we have and the most restrictively licensed. NC bites because thamizh-mcp is Apache-2.0 and anyone may run it commercially; ND bites because a restructured store is plausibly a derivative. When wired: no bundled copy, query-time lookup, cache in the gitignored knowledge.sqlite3 only, store the derived FACT and citation never the entry text, excluded from gold-corpus export, opt-in and disabled by default. Also robots.txt-blocked (Disallow: /cgi-bin/ is the only endpoint), so integration awaits written permission. Recorded honestly: this is the cautious reading, and Saran is approaching DSAL for permission.