The plan divided into 148 numbered units across four stages; the sources the corpus is assembled from; and an honest account of what is loaded, what is queued, and what has still to be acquired.
The work divides into four stages. They are thematic rather than strictly sequential — collection continues while the lexicon is built, and the lexicon deepens while the applications are written.
Veda, Vedāṅga, Upāṅga and Upaveda; the Purāṇas, Itihāsa, Kāvya and the ancillary literature, together with their commentaries. Where a text is not already proof-read, it enters through the digitisation route.
Loading is not mere ingestion. The Bhāṣya is processed so that its Vyākaraṇa and śāstra references are linked to their targets; the related portions are brought together, so a Mantra or Brāhmaṇa passage stands beside its Pada-pāṭha and its Bhāṣyam; and the MAP analysis is produced for each mantra on demand rather than stored. Throughout, a master list is maintained that can absorb further texts and references without being restructured.
Three vocabularies, built by different means because they are different in kind.
Rūḍha — the conventional vocabulary, drawn from the kośas. Being non-compositional by definition, it must be enumerated rather than derived.
Vaidika — every available Pada-pāṭha ingested, Pada-pāṭha generated where none exists, and the unique words with their meanings extracted from the Veda Bhāṣyam. A Pada-pāṭha processor brings uniformity at load time and keeps the processed padas separately, so that the Krama and Vikṛti sandhi layers can be built upon them.
Yaugika — the Vyākaraṇa Vocabulary Engine: the Pāṇinian word-space, over a hundred million forms, each carrying its complete sūtra-by-sūtra derivation, with meanings in Telugu, Kannada, Hindi and English. Telugu leads, because the team's native expertise gives the subtlest sense.
Alongside these: Nirukta etymology, with traditional nirvacana set beside the formal Pāṇinian vyutpatti; the liṅga rules from Nāma-liṅgānuśāsana, Mahābhāṣya and Kāśikā; Krama and the eight Vikṛti forms from Jaṭā to Ghana, with rule-verified sandhi and svara at every junctura; and Varṇa Krama — akṣara-by-akṣara decomposition with svara and mātrā, built directly on Śikṣā and Prātiśākhya.
Data synthesised across the verticals to establish Vākya analysis, and the cross-references that run through Veda, Vedāṅga, Upāṅga and Upaveda connected so the corpus reads as one body rather than many.
Anuvāda — translation of the mantras into Indian languages founded solely on the Bhāṣya; where no Bhāṣya exists, on the principles of Sāyaṇācārya. Samānatā — an engine determining similarity in śabda and in artha.
On these rest the applications: Śrauta Prayoga generation with every step carrying a traceable sūtra citation; the Mīmāṃsā Nyāya engine; Sāmaveda stotra generation with viṣṭuti patterns and stobha insertion; Tarka sentence structure in the śābda-bodha style with pariṣkāra; Sāma sound analysis, audio in and notation out, with the Ūha engine that re-frames melody onto new chandas; and the Chatur-Darśana engine, which interprets one sentence in parallel through Tarka, Vyākaraṇa, Mīmāṃsā and Vedānta — the śāstrārtha assembly, computed.
The work reaches its audience: integration with general search; the Vedic perspective offered on any topic on demand, with societal reference; APIs for researchers seeking Vedic references with plain meanings; video, audio and web publication; webinars; the VVS textbooks; short-form material across thousands of topics; a daily presence; and a considered śāstra perspective on contemporary events.
The plan is deliberately granular. Each unit carries a permanent number, a definition of done, and its dependencies — so progress is a matter of record rather than impression.
Unit numbers (SVU-001 onward) never change and never encode a stage, so work can be resequenced without renumbering. The same discipline governs every identifier in the project: an identifier records what a thing is, never where it currently sits.
The corpus is assembled from the Foundation's own digitised holdings, the open Sanskrit lexical corpora, and a Pāṇinian derivation engine that generates grammatical forms rather than storing them.
| Source | What it contributes | Items | Size |
|---|---|---|---|
| Veda mūla | Saṃhitā text of the four Vedas, with svara | 30 | 86 MB |
| Pada-pāṭha | Word-separated recitational text | 7 | 7 MB |
| Bhāṣya | Traditional commentary — Sāyaṇa, Bhaṭṭa Bhāskara and others | 94 | 2,324 MB |
| Vedāṅga | Śikṣā, Prātiśākhya, Kalpa and the ancillary śāstras | 63 | 233 MB |
| Prayoga | Ritual manuals and performance texts | 33 | 321 MB |
| Lexical — Cologne Digital Sanskrit Lexicon | 44 dictionaries: Monier-Williams, Apte, Amarakośa, Vācaspatyam, Śabdakalpadruma and others | 44 | — |
| Lexical — indic-dict collections | 19 further kośa builds used as independent cross-check witnesses | 19 | — |
| Grammatical engine | Pāṇinian derivation engine with the Dhātupāṭha — 2,229 dhātus, 5,160 sūtras, 128 kṛt and 181 taddhita pratyayas | 1 | — |
| Pāṇinian reference corpora | Aṣṭādhyāyī commentary and annotation sets surveyed for reuse | 4 | — |
Every held artefact is classified by what stands between it and the pipeline. Nothing has been loaded to the production corpus yet — the schema is still under scholarly review — so the whole holding is queued.
| State | Meaning | Files | Size |
|---|---|---|---|
| Ready to load | Text-bearing files that need no conversion. | 156 | 569 MB |
| Awaiting assembly | Recognition already complete; output not yet assembled. | 30 | 1,300 MB |
| Awaiting triage | To be checked for a text layer before any recognition. | 34 | 1,085 MB |
| Awaiting conversion | Legacy word-processor formats. | 2 | 7 MB |
| Superseded | A better copy of the same work is held. | 5 | 10 MB |
The seventy-two volume plan names 390 distinct works. Reconciling that requirement against the holdings shows precisely what has still to be sourced.
| Position | Distinct works |
|---|---|
| Held | 19 |
| Held, pending verification | 68 |
| Not yet acquired | 303 |
| Category | Works to acquire |
|---|---|
| Darśana / Philosophy | 34 |
| Saṃhitā | 30 |
| Saṅgīta / Nāṭya | 23 |
| Purāṇa | 21 |
| Dharmasūtra / Smṛti | 19 |
| Kāvya / Alaṃkāra | 18 |
| Brāhmaṇa | 15 |
| Stotra / Nāmāvali | 14 |
| Upaniṣad | 14 |
| Gṛhyasūtra | 13 |
The holdings are strongest in Veda mūla, bhāṣya and vedāṅga — the Foundation's own scholarly territory — and thinnest in the darśana, purāṇa, kāvya and applied-śāstra divisions. Acquisition is therefore sequenced by what the volumes actually require, not by what is easiest to obtain.
Four stages, 148 units. Stage numbering is thematic; several stages run concurrently. Items marked Awaiting decision are held pending a scholarly determination and are deliberately not started, because the determination may change the work.
Infrastructure, schema, governance and the registers that everything else is tracked against.
| Unit | Task | Track | Status |
|---|---|---|---|
| SVU-001 | Corpus schema v1 sign-off and apply | Governance | Awaiting decision |
| SVU-002 | Migration 007 and the pada_class enum | Governance | Awaiting decision |
| SVU-003 | GPU box commissioning | Infra | Complete |
| SVU-004 | Core data services deployed | Infra | Complete |
| SVU-005 | Scholarly review panel constituted | Governance | Planned |
| SVU-006 | Licence declarations for the three own-artefacts | Governance | Awaiting decision |
| SVU-007 | Repository governance and access control | Governance | Planned |
| SVU-008 | Document numbering and register | Governance | Complete |
| SVU-009 | Dataset register with loading status | Corpus | Complete |
| SVU-010 | Required-books reconciliation vs the 72 volumes | Corpus | Complete |
| SVU-011 | Document pipeline and OOXML validity gate | Tooling | Complete |
| SVU-012 | Corpus resilience and off-site replication | Risk | Planned |
Encoding the accent correctly, then loading the text — mūla, pada-pāṭha, bhāṣya and the ancillary śāstras — with provenance intact.
| Unit | Task | Track | Status |
|---|---|---|---|
| SVU-020 | VijayaDV PUA inventory | Encoding | Complete |
| SVU-021 | PUA three-class classification (draft) | Encoding | Complete |
| SVU-022 | Residual PUA → Unicode bindings | Encoding | Awaiting decision |
| SVU-023 | Accent-collapse principle confirmed | Encoding | Awaiting decision |
| SVU-024 | Full mūla conversion, all four Vedas | Encoding | Awaiting decision |
| SVU-025 | Round-trip gate extended to the full corpus | Encoding | Awaiting decision |
| SVU-026 | Load group 1 — 120 text-ready files | Loading | Awaiting decision |
| SVU-027 | Assemble the already-OCR'd volumes | Loading | Complete |
| SVU-028 | Triage the 34 un-OCR'd PDFs (gate O1) | Loading | Ready to start |
| SVU-029 | OCR the genuine scans | Loading | Planned |
| SVU-030 | Plain-text extraction for text-layer PDFs | Loading | Planned |
| SVU-031 | Convert the 2 legacy .doc/.pub files | Loading | Ready to start |
| SVU-032 | Retire the 5 superseded PDFs | Loading | Ready to start |
| SVU-033 | Ingest the Kalpa Sūtra Unicode collection (Track A) | Corpus | Awaiting decision |
| SVU-034 | Citation extractor and quoted-text register | Corpus | Ready to start |
| SVU-035 | Prioritised acquisition list from the 303 gaps | Corpus | Ready to start |
| SVU-036 | Tier-3 normalisation — GRETIL, Muktabodha, VHP | Corpus | Planned |
| SVU-037 | Sentence segmentation and locus tagging | Corpus | Planned |
| SVU-038 | Fix the locus identifier scheme | Architecture | Awaiting decision |
| SVU-039 | Embeddings generated and loaded to Qdrant | Retrieval | Planned |
| SVU-040 | Knowledge graph v1 in Neo4j | Retrieval | Planned |
| SVU-041 | MAP gold exemplars processed | MAP | Awaiting decision |
| SVU-042 | Bhāṣya cross-linked to Vyākaraṇa and Śāstra references | MAP | Planned |
| SVU-043 | Mantra / Brāhmaṇa grouping assembled | MAP | Planned |
| SVU-044 | Dynamic MAP analysis per mantra | MAP | Planned |
| SVU-044.1 | Spec — MAP analysis contract — what the reader is shown and from what | MAP | Ready to start |
| SVU-044.2 | Build — Dynamic MAP analysis per mantra | MAP | Planned |
| SVU-045 | Master list that absorbs new references | Corpus | In progress |
| SVU-046 | Feed vaakya.vedanidhi.in during the work | Delivery | Planned |
| SVU-047 | Digitisation ledger and seven-stage pipeline commissioned | OCR | Awaiting decision |
| SVU-048 | Reconcile filing of recently added documents | Hygiene | Ready to start |
The lexical foundation: the conventional vocabulary drawn from the kośas, the Vedic vocabulary drawn from pada-pāṭha and bhāṣya, and the compositional vocabulary generated from Pāṇini's rules.
| Unit | Task | Track | Status |
|---|---|---|---|
| SVU-050 | Load Compartment A — the kośa corpus | Pada Kośa A | Awaiting decision |
| SVU-051 | Kośa census published | Pada Kośa A | Complete |
| SVU-052 | Access and serving policy for third-party lexical sources | Pada Kośa A | Awaiting decision |
| SVU-053 | Kośa precedence order for the reader | Pada Kośa A | Awaiting decision |
| SVU-054 | Flag the 106,014 single-attestation words | Pada Kośa A | Awaiting decision |
| SVU-055 | Ingest all available Pada Pāṭhas | Pada Kośa V | Awaiting decision |
| SVU-056 | Pada Pāṭha Processor | Pada Kośa V | Awaiting decision |
| SVU-057 | Generate Padapāṭha where none exists | Pada Kośa V | Planned |
| SVU-057.1 | Spec — Rules for generating Pada-pāṭha where none is attested | Pada Kośa V | Ready to start |
| SVU-057.2 | Build — Generate Padapāṭha where none exists | Pada Kośa V | Planned |
| SVU-058 | Vaidika Pada Kośa — words and meanings from Bhāṣya | Pada Kośa V | Planned |
| SVU-058.1 | Spec — Method for extracting the Vaidika lexicon from Bhāṣya | Pada Kośa V | Ready to start |
| SVU-058.2 | Build — Vaidika Pada Kośa — words and meanings from Bhāṣya | Pada Kośa V | Planned |
| SVU-059 | Yaugika generator built and green | Pada Kośa B | Complete |
| SVU-060 | Svara print convention fixed | Pada Kośa B | Awaiting decision |
| SVU-061 | Pracaya / ekaśruti decision | Pada Kośa B | Awaiting decision |
| SVU-062 | Accent-agreement harness against the gold corpus | Pada Kośa B | Awaiting decision |
| SVU-062.1 | Spec — Accent-agreement test design and acceptance threshold | Pada Kośa B | Awaiting decision |
| SVU-062.2 | Build — Accent-agreement harness against the gold corpus | Pada Kośa B | Planned |
| SVU-063 | Form-space target decided | Pada Kośa B | Awaiting decision |
| SVU-064 | Full paradigm sweep | Pada Kośa B | Awaiting decision |
| SVU-065 | FST lexicon compiled | Pada Kośa B | Awaiting decision |
| SVU-066 | Multilingual morpheme glosses | Pada Kośa B | Planned |
| SVU-067 | Nirukta etymology engine | Nirukta | Awaiting decision |
| SVU-067.1 | Spec — Nirvacana model — sources, competing etymologies, presentation | Nirukta | Ready to start |
| SVU-067.2 | Build — Nirukta etymology engine | Nirukta | Planned |
| SVU-068 | Gender rules and the liṅga authorities | Vyākaraṇa | Planned |
| SVU-069 | Sandhi rule engine | Pāṭha | Planned |
| SVU-069.1 | Spec — Sandhi rule inventory and svara behaviour at each junctura | Pāṭha | Ready to start |
| SVU-069.2 | Build — Sandhi rule engine | Pāṭha | Planned |
| SVU-070 | Krama generation | Pāṭha | Planned |
| SVU-070.1 | Spec — Krama construction rules | Pāṭha | Ready to start |
| SVU-070.2 | Build — Krama generation | Pāṭha | Planned |
| SVU-071 | Eight Vikṛti forms | Pāṭha | Planned |
| SVU-071.1 | Spec — The eight Vikṛti forms — construction rules per form | Pāṭha | Ready to start |
| SVU-071.2 | Build — Eight Vikṛti forms | Pāṭha | Planned |
| SVU-072 | Varṇa Krama generator | Pāṭha | Complete |
| SVU-072.1 | Spec — Varṇa Krama decomposition rules from Śikṣā and Prātiśākhya | Pāṭha | Ready to start |
| SVU-072.2 | Build — Varṇa Krama generator | Pāṭha | Planned |
| SVU-073 | Lookup service and hover payload | Pada Kośa | In progress |
| SVU-074 | Lemma link between Veda words and their attestations | Pada Kośa V | Awaiting decision |
Cross-referencing, translation, the generative engines, and the reading application built on top of them.
| Unit | Task | Track | Status |
|---|---|---|---|
| SVU-080 | Cross-reference graph across all verticals | Synthesis | Planned |
| SVU-081 | Anuvāda — Bhāṣya-grounded translation | Synthesis | Planned |
| SVU-081.1 | Spec — Anuvāda method — grounding every rendering in Bhāṣya | Synthesis | Ready to start |
| SVU-081.2 | Build — Anuvāda — Bhāṣya-grounded translation | Synthesis | Planned |
| SVU-082 | Samānatā — śabda and artha similarity engine | Synthesis | Planned |
| SVU-082.1 | Spec — Samānatā — what counts as similarity in śabda and in artha | Synthesis | Ready to start |
| SVU-082.2 | Build — Samānatā — śabda and artha similarity engine | Synthesis | Planned |
| SVU-083 | Vākya analysis | Synthesis | Planned |
| SVU-084 | SFT dataset assembled | Model | Planned |
| SVU-085 | Base model QLoRA fine-tune | Model | Planned |
| SVU-086 | RAG pipeline | Model | Planned |
| SVU-087 | RLHF-lite preference pass | Model | Planned |
| SVU-088 | Re-benchmark the QLoRA epoch estimate | Model | Awaiting decision |
| SVU-089 | Śrauta Prayoga generation | Engine | Planned |
| SVU-089.1 | Spec — Śrauta Prayoga sequence model and citation requirements | Engine | Ready to start |
| SVU-089.2 | Build — Śrauta Prayoga generation | Engine | Planned |
| SVU-090 | Sāmaveda Stotra generation | Engine | Planned |
| SVU-090.1 | Spec — Stotriyā assembly, viṣṭuti patterns and stobha rules | Engine | Ready to start |
| SVU-090.2 | Build — Sāmaveda Stotra generation | Engine | Planned |
| SVU-091 | Vikṛti Pāṭha model | Engine | Planned |
| SVU-092 | Tarka sentence-structure model | Engine | Planned |
| SVU-092.1 | Spec — Śābda-bodha and pariṣkāra representation for Tarka | Engine | Ready to start |
| SVU-092.2 | Build — Tarka sentence-structure model | Engine | Planned |
| SVU-093 | Varṇa Krama model — confirm whether wanted at all | Engine | Awaiting decision |
| SVU-094 | Sāma sound analysis | Engine | Planned |
| SVU-094.1 | Spec — Svara extraction and notation conventions for Sāmagāna | Engine | Ready to start |
| SVU-094.2 | Build — Sāma sound analysis | Engine | Planned |
| SVU-095 | Ūha engine | Engine | Planned |
| SVU-095.1 | Spec — Ūha — how melody is re-framed onto new chandas | Engine | Ready to start |
| SVU-095.2 | Build — Ūha engine | Engine | Planned |
| SVU-096 | Mīmāṃsā Nyāya engine | Engine | Planned |
| SVU-096.1 | Spec — Mīmāṃsā adhikaraṇa structure for machine analysis | Engine | Ready to start |
| SVU-096.2 | Build — Mīmāṃsā Nyāya engine | Engine | Planned |
| SVU-097 | Chatur-Darśana interpretation engine | Engine | Planned |
| SVU-097.1 | Spec — The four lenses — scope and retrieval boundary of each | Engine | Ready to start |
| SVU-097.2 | Build — Chatur-Darśana interpretation engine | Engine | Planned |
| SVU-098 | Layered text display L1–L5 | App | In progress |
| SVU-099 | Grammar hover system | App | In progress |
| SVU-100 | Prayoga Mode interface | App | Planned |
| SVU-101 | Vikṛti generator interface | App | Planned |
| SVU-102 | Sāma Studio | App | Planned |
| SVU-103 | Tarka Workbench and Varṇa Krama viewer | App | Planned |
| SVU-104 | Stotra Library and Nirukta Explorer | App | Planned |
| SVU-105 | Cross-disciplinary notes layer | App | Planned |
| SVU-106 | vLLM serving with batching and caching | Infra | Planned |
| SVU-107 | Institutional pilot | Validation | Planned |
| SVU-108 | Swadharma QA benchmark | Validation | Planned |
| SVU-109 | Prayoga validated against the 2026 Agniṣṭoma record | Validation | Planned |
Publication, the researcher interface, and bringing the material to a general audience.
| Unit | Task | Track | Status |
|---|---|---|---|
| SVU-115 | sarvaveda.info updated with this plan | Site | Ready to start |
| SVU-116 | Static pre-rendering for discoverability | Site | Planned |
| SVU-117 | Contribution mechanism — in-page edit to pull request | Site | Awaiting decision |
| SVU-118 | Public web app with tiered access | Release | Planned |
| SVU-119 | Researcher API | Release | Planned |
| SVU-120 | Search-engine integration and structured data | Reach | Planned |
| SVU-121 | Vedic perspective on demand, with societal references | Reach | Planned |
| SVU-122 | Video, audio and web publication pipeline | Media | Planned |
| SVU-123 | Webinar programme | Media | Planned |
| SVU-124 | VVS textbooks | Media | Planned |
| SVU-125 | AI-produced reels across thousands of topics | Media | Planned |
| SVU-126 | Daily social presence | Media | Planned |
| SVU-127 | Daily śāstra perspective on contemporary events | Media | Planned |
| SVU-128 | Publish the consultation papers | Release | Awaiting decision |
| SVU-129 | Continuous learning cycle | Release | Planned |
| SVU-130 | Video corpus transcription | Corpus | Planned |