科研速览 · Science Skim继续刷下去 · Keep skimming →
◇ Knowledge Commons (Lakehead University)2026-08-09· Decipherment

The Indus Operating System: A Functional Decipherment of the Indus Valley Administrative Protocol

ADRIAN SHARMAN

原始摘要(英文原文)· Original abstract
The Indus Valley script has resisted phonetic decipherment for a century. This paper argues that phonetic decipherment was the wrong target, and demonstrates a different kind of result: a functional decipherment — a complete, testable model of what Indus inscriptions do, validated on documents the model never saw. From a 2,543-object corpus (11,280 sign tokens, ICIT-derived) we derive a four-rule record grammar: every inscription is one or more [HEADER][CONTENT][TERMINAL] records, with two forbidden class transitions, two bound sign compounds, and a prohibition on document-initial numerals. Trained on 80% of the corpus, the grammar parses 78.8% ± 1.9% of held-out inscriptions against a shuffled-order null of 27.1% ± 2.2%. Frozen and applied to 959 inscriptions from an independently compiled expansion corpus, entirely absent from training, it parses 80.3% against a 38.2% null. Characterised in this revision against a graded ladder of null models, the four rules prove to be a compact and legible specification of the protocol's end-conditioned first-order structure: a first-order Markov surrogate that additionally knows how documents terminate reproduces the corpus's grammaticality in full, and slightly exceeds it. That is the expected profile of a rigid administrative form and not of prose. A separate length-matched test finds that the class sequences carry genuine second-order dependence that the four rules do not capture. The system described is a two-channel administrative protocol: seal iconography carries institutional identity (WHO), the sign sequence carries transaction content (WHAT and HOW MUCH), and no institution holds exclusive rights to any sign. Quantity marking is real and economically responsive: the modal numeral G2 — the standard consignment unit — covaries with commodity department, material, and document class (p ≤ 0.019 on four independent axes), collapsing from 68.0% in commerce documents to 27.3% in audit documents. Physical evidence anchors the statistics: 42% of homogeneous duplicate-formula groups cluster in museum accession space beyond a per-site random null, several at literally consecutive catalogue numbers, and commodity departments occupy named excavation trenches at Harappa (p < 0.0005). We report size-match-corrected effect sizes throughout, maintain a falsification registry of 34 killed or null hypotheses including two entire interpretive branches, and release every script, the database, and the findings ledger for one-command reproduction. The Indus script emerges not as a silenced language but as a working information system: a rigid, low-entropy, network-wide protocol for authorising and recording transactions — an operating system, four and a half millennia before the term. Version 2.0 (9 August 2026) is a corrective revision. Version 1.0 reported the grammar validation against a single order-shuffle null, which cannot distinguish real structure from the corpus's own low-order statistics. Re-measured against a ladder of progressively better-informed nulls, an end-conditioned first-order chain parses 82.9% against the corpus's 79.2% in-corpus and 88.0% against 80.3% out-of-corpus; the claim that the four rules encode structure beyond first-order transition statistics is therefore killed and entered in the falsification registry. The parse rates, the out-of-corpus replication, the derivation controls, and the non-grammar findings are unchanged. Version 1.0 remains retrievable and is not silently superseded.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

The Indus Operating System: A Functional Decipherment of the Indus Valley Administrative Protocol — 科研速览 Science Skim