The Means of Valid Knowledge
Classical Indian epistemology distinguishes carefully between the elements involved in an act of knowing.
Pramā: Valid cognition.
Pramātā: The knower in whom that cognition arises.
Prameya: The object toward which the cognition is directed.
Pramāṇa: The means by which valid cognition is produced [1].
The Nyāya school enumerates four such means: Pratyakṣa (perception), Anumāna (inference), Upamāna (comparison), and Śabda (verbal testimony) [2]. Other schools count differently—Sāṃkhya accepts three, Advaita Vedānta six, the Cārvāka materialists only one. That the schools disagree is itself significant: this is a living tradition of rigorous argument about the foundations of knowledge, not a settled catechism [3].
Contemporary discussions of artificial intelligence have no comparably developed vocabulary at the pramāṇa layer. We see extensive work on whether model outputs are accurate, but remarkably little on what kind of knowledge-source a model purports to be. That gap is the main subject of this essay.
Śabda and the Condition of Āpta
Every school that admits śabda as a valid means of knowledge admits it conditionally, and the condition of admission is key. The Nyāya Sūtra defines verbal testimony as āptopadeśa—the instruction of a trustworthy person [4]. The commentarial tradition specifies what trustworthiness requires: an āpta is one who possesses direct knowledge of the matter at hand and who communicates that knowledge faithfully, without distortion or intent to mislead [5]. Both conditions are strictly necessary. One who knows but misreports is not āpta; neither is one who is sincere but ignorant. Śabda is therefore not equivalent to credulity. It is a system with an explicit trust anchor and an explicit failure condition—a considerably more rigorous account of testimony than much of what presently circulates under the heading of information literacy.
Testimony Without a Testifier
In today's world: The output of a large language model is structurally śabda.
It does not reach the user as perception; the user has observed nothing. It does not reach the user as inference; the user has derived nothing, and neither, in any substantive sense, has the model. It arrives as authoritative verbal testimony—fluent, confident, well-formed, in the register of one who knows. Every surface property by which testimony is ordinarily credited, the model reproduces faithfully.
Behind it though, there is no āpta. There is only a training distribution.
This is a much stronger objection than the familiar issue of model hallucination, though hallucination is among its symptoms [6]. When an LLM reports what a classical text states on a given question, there is no party who knows and is reporting faithfully. There is only a weighted aggregate of ingested material—nineteenth-century colonial indology, contemporary academic commentary, devotional writing, polemic, and summaries of summaries—collapsed into a single confident voice. It leaves no visible seams and no mechanism by which any assertion may be traced to a source.
The tradition possesses a precise term for the alternative: Paramparā, the traceable chain of transmission.
Where a traditional teacher advances a claim about a text, it is in principle possible to ask from whom was this received?and follow the answer backwards. The chain is auditable. On this account, auditability is not incidental to the validity of knowledge; it is partly constitutive of it. A language model cannot be interrogated in this way—not because the capability has yet to be engineered, but because there is no corresponding fact to retrieve. Provenance was dissolved in training.
From Philosophy to Measurement
We are conscious of a genre this argument might be mistaken for: the genre that discovers management principles in the epics and leadership doctrine in the Gītā. That genre treats a contemporary framework as the reference point and mines the tradition for ornamental parallels. We regard it as a poor use of a serious inheritance.
Our claim runs in the opposite direction.
We do not suggest that Nyāya epistemology resembles an authentication problem. We suggest that Nyāya offers a fully worked theory of the conditions under which testimony confers warrant; that this theory is more developed than the criteria presently employed in model evaluation; and that applying it to language models yields specific, testable measures.
The āpta condition decomposes into four concrete audit criteria:
Provenance. Can the model distinguish the claims of a tradition from claims made about that tradition? Where it reports on a contested passage, does it identify whose reading it is presenting, or does it advance one commentarial position — frequently the position best represented in English-language sources — as simply what the text states?
Faithful transmission. Does the model attribute correctly? Verse citations, commentarial attributions, and the assignment of doctrinal positions to schools are objectively verifiable, and error rates on such tasks would be disqualifying in any professional context.
Calibration under disagreement. Where a question is genuinely disputed — between sampradāyas, or between traditional and academic readings — does the model represent the dispute as a dispute, or manufacture a false settlement?
Symmetry across traditions. Are these failure rates consistent, or do they vary systematically by tradition? Existing work has documented that refusal behaviour, stereotype association and representational omission differ measurably across religious categories.7
None of these measures requires us to determine which sampradāya is correct — a question we are not competent to adjudicate and which is not ours to settle. They require only that the model represent the landscape of positions accurately and attribute them honestly. That is a standard to which a scholar of any persuasion, or none, may consent to be held.
None of these measures requires us to determine which sampradāya is correct—a question we are not competent to adjudicate and which is not ours to settle. They require only that the model represent the landscape of positions accurately and attribute them honestly. That is a standard to which a scholar of any persuasion, or none, may consent to be held.
The Benchmark Program
We are undertaking a systematic evaluation of how frontier language models handle the internal self-understanding of religious traditions, beginning with the Indic traditions and designed from the outset for comparative application across others.
The intended output is not commentary on algorithmic bias, but an empirical benchmark: a published methodology, reproducible results, and documented failure cases.
We will be mistaken about some things. The surest route to failure would be for researchers trained in computer science to issue confident judgments on material they are not qualified to arbitrate. We therefore impose an explicit design constraint: every benchmark item touching textual or commentarial ground is reviewed by scholars who hold that learning, and the review panel is published alongside the results.
This is the paramparā performing the exact function a paramparā exists to perform. A pramāṇa has now been constructed that speaks with complete confidence and answers to no one. It will teach a generation. The least that can be done is to measure how well it performs the task.
References
Footnotes & References
Stephen H. Phillips, Epistemology in Classical India: The Knowledge Sources of the Nyāya School (Routledge, 2012), ch. 1.
Nyāya Sūtra 1.1.3 (Gautama). See also B. K. Matilal, Perception: An Essay on Classical Indian Theories of Knowledge (Clarendon Press, 1986).
Jonardon Ganeri, Philosophy in Classical India: The Proper Work of Reason (Routledge, 2001).
Nyāya Sūtra 1.1.7.
Vātsyāyana, Nyāya Bhāṣya on Nyāya Sūtra 1.1.7.
Ziwei Ji et al., "Survey of Hallucination in Natural Language Generation," ACM Computing Surveys 55, no. 12 (2023).
Auditing studies of religious representation, omission, and stereotype benchmarks in frontier models (e.g., Indian-BhED).
