SarvaVeda LLM ← Home
Plan of Action

Plan of Action

The plan divided into 148 numbered units across four stages; the sources the corpus is assembled from; and an honest account of what is loaded, what is queued, and what has still to be acquired.

॥ अनन्ता वै वेदाः ॥
Section 1

The plan in four stages

The work divides into four stages. They are thematic rather than strictly sequential — collection continues while the lexicon is built, and the lexicon deepens while the applications are written.

Stage 1 Collect and load the texts

Veda, Vedāṅga, Upāṅga and Upaveda; the Purāṇas, Itihāsa, Kāvya and the ancillary literature, together with their commentaries. Where a text is not already proof-read, it enters through the digitisation route.

Loading is not mere ingestion. The Bhāṣya is processed so that its Vyākaraṇa and śāstra references are linked to their targets; the related portions are brought together, so a Mantra or Brāhmaṇa passage stands beside its Pada-pāṭha and its Bhāṣyam; and the MAP analysis is produced for each mantra on demand rather than stored. Throughout, a master list is maintained that can absorb further texts and references without being restructured.

Stage 2 Pada Kośa — the vocabulary

Three vocabularies, built by different means because they are different in kind.

Rūḍha — the conventional vocabulary, drawn from the kośas. Being non-compositional by definition, it must be enumerated rather than derived.

Vaidika — every available Pada-pāṭha ingested, Pada-pāṭha generated where none exists, and the unique words with their meanings extracted from the Veda Bhāṣyam. A Pada-pāṭha processor brings uniformity at load time and keeps the processed padas separately, so that the Krama and Vikṛti sandhi layers can be built upon them.

Yaugika — the Vyākaraṇa Vocabulary Engine: the Pāṇinian word-space, over a hundred million forms, each carrying its complete sūtra-by-sūtra derivation, with meanings in Telugu, Kannada, Hindi and English. Telugu leads, because the team's native expertise gives the subtlest sense.

Alongside these: Nirukta etymology, with traditional nirvacana set beside the formal Pāṇinian vyutpatti; the liṅga rules from Nāma-liṅgānuśāsana, Mahābhāṣya and Kāśikā; Krama and the eight Vikṛti forms from Jaṭā to Ghana, with rule-verified sandhi and svara at every junctura; and Varṇa Krama — akṣara-by-akṣara decomposition with svara and mātrā, built directly on Śikṣā and Prātiśākhya.

Stage 3 Synthesis and applications

Data synthesised across the verticals to establish Vākya analysis, and the cross-references that run through Veda, Vedāṅga, Upāṅga and Upaveda connected so the corpus reads as one body rather than many.

Anuvāda — translation of the mantras into Indian languages founded solely on the Bhāṣya; where no Bhāṣya exists, on the principles of Sāyaṇācārya. Samānatā — an engine determining similarity in śabda and in artha.

On these rest the applications: Śrauta Prayoga generation with every step carrying a traceable sūtra citation; the Mīmāṃsā Nyāya engine; Sāmaveda stotra generation with viṣṭuti patterns and stobha insertion; Tarka sentence structure in the śābda-bodha style with pariṣkāra; Sāma sound analysis, audio in and notation out, with the Ūha engine that re-frames melody onto new chandas; and the Chatur-Darśana engine, which interprets one sentence in parallel through Tarka, Vyākaraṇa, Mīmāṃsā and Vedānta — the śāstrārtha assembly, computed.

Stage 4 Presentation

The work reaches its audience: integration with general search; the Vedic perspective offered on any topic on demand, with societal reference; APIs for researchers seeking Vedic references with plain meanings; video, audio and web publication; webinars; the VVS textbooks; short-form material across thousands of topics; a daily presence; and a considered śāstra perspective on contemporary events.

Section 2

Where the work stands

The plan is deliberately granular. Each unit carries a permanent number, a definition of done, and its dependencies — so progress is a matter of record rather than impression.

148numbered work units
12complete
24ready to start
227corpus artefacts held
2.9 GBof primary text

Unit numbers (SVU-001 onward) never change and never encode a stage, so work can be resequenced without renumbering. The same discipline governs every identifier in the project: an identifier records what a thing is, never where it currently sits.

Section 3

Data sources

The corpus is assembled from the Foundation's own digitised holdings, the open Sanskrit lexical corpora, and a Pāṇinian derivation engine that generates grammatical forms rather than storing them.

Source What it contributes Items Size
Veda mūlaSaṃhitā text of the four Vedas, with svara3086 MB
Pada-pāṭhaWord-separated recitational text77 MB
BhāṣyaTraditional commentary — Sāyaṇa, Bhaṭṭa Bhāskara and others942,324 MB
VedāṅgaŚikṣā, Prātiśākhya, Kalpa and the ancillary śāstras63233 MB
PrayogaRitual manuals and performance texts33321 MB
Lexical — Cologne Digital Sanskrit Lexicon44 dictionaries: Monier-Williams, Apte, Amarakośa, Vācaspatyam, Śabdakalpadruma and others44
Lexical — indic-dict collections19 further kośa builds used as independent cross-check witnesses19
Grammatical enginePāṇinian derivation engine with the Dhātupāṭha — 2,229 dhātus, 5,160 sūtras, 128 kṛt and 181 taddhita pratyayas1
Pāṇinian reference corporaAṣṭādhyāyī commentary and annotation sets surveyed for reuse4
On the lexical sources. The lexical collections are held for comparison and corroboration: of the distinct Sanskrit words assembled, roughly two-thirds are attested by more than one independent kośa. Where a word rests on a single witness it is recorded as such rather than presented as settled.
Section 4

Pending uploads

Every held artefact is classified by what stands between it and the pipeline. Nothing has been loaded to the production corpus yet — the schema is still under scholarly review — so the whole holding is queued.

State Meaning Files Size
Ready to loadText-bearing files that need no conversion.156569 MB
Awaiting assemblyRecognition already complete; output not yet assembled.301,300 MB
Awaiting triageTo be checked for a text layer before any recognition.341,085 MB
Awaiting conversionLegacy word-processor formats.27 MB
SupersededA better copy of the same work is held.510 MB
The immediate opportunity. A substantial block of text recognition has already been completed and paid for, but the output was never assembled into documents. Assembling it requires no new expenditure and is the largest body of readable text available to the project today. It is unit SVU-027 in the plan below.
Section 5

Acquisition gaps

The seventy-two volume plan names 390 distinct works. Reconciling that requirement against the holdings shows precisely what has still to be sourced.

Position Distinct works
Held19
Held, pending verification68
Not yet acquired303

Where the gaps concentrate

Category Works to acquire
Darśana / Philosophy34
Saṃhitā30
Saṅgīta / Nāṭya23
Purāṇa21
Dharmasūtra / Smṛti19
Kāvya / Alaṃkāra18
Brāhmaṇa15
Stotra / Nāmāvali14
Upaniṣad14
Gṛhyasūtra13

The holdings are strongest in Veda mūla, bhāṣya and vedāṅga — the Foundation's own scholarly territory — and thinnest in the darśana, purāṇa, kāvya and applied-śāstra divisions. Acquisition is therefore sequenced by what the volumes actually require, not by what is easiest to obtain.

Section 6

Project tasks

Four stages, 148 units. Stage numbering is thematic; several stages run concurrently. Items marked Awaiting decision are held pending a scholarly determination and are deliberately not started, because the determination may change the work.

Stage 0 · Foundation & Governance — 12 units, 6 complete

Infrastructure, schema, governance and the registers that everything else is tracked against.

Unit Task Track Status
SVU-001Corpus schema v1 sign-off and applyGovernanceAwaiting decision
SVU-002Migration 007 and the pada_class enumGovernanceAwaiting decision
SVU-003GPU box commissioningInfraComplete
SVU-004Core data services deployedInfraComplete
SVU-005Scholarly review panel constitutedGovernancePlanned
SVU-006Licence declarations for the three own-artefactsGovernanceAwaiting decision
SVU-007Repository governance and access controlGovernancePlanned
SVU-008Document numbering and registerGovernanceComplete
SVU-009Dataset register with loading statusCorpusComplete
SVU-010Required-books reconciliation vs the 72 volumesCorpusComplete
SVU-011Document pipeline and OOXML validity gateToolingComplete
SVU-012Corpus resilience and off-site replicationRiskPlanned

Stage 1 · Collect & Load the Corpus — 31 units, 3 complete

Encoding the accent correctly, then loading the text — mūla, pada-pāṭha, bhāṣya and the ancillary śāstras — with provenance intact.

Unit Task Track Status
SVU-020VijayaDV PUA inventoryEncodingComplete
SVU-021PUA three-class classification (draft)EncodingComplete
SVU-022Residual PUA → Unicode bindingsEncodingAwaiting decision
SVU-023Accent-collapse principle confirmedEncodingAwaiting decision
SVU-024Full mūla conversion, all four VedasEncodingAwaiting decision
SVU-025Round-trip gate extended to the full corpusEncodingAwaiting decision
SVU-026Load group 1 — 120 text-ready filesLoadingAwaiting decision
SVU-027Assemble the already-OCR'd volumesLoadingComplete
SVU-028Triage the 34 un-OCR'd PDFs (gate O1)LoadingReady to start
SVU-029OCR the genuine scansLoadingPlanned
SVU-030Plain-text extraction for text-layer PDFsLoadingPlanned
SVU-031Convert the 2 legacy .doc/.pub filesLoadingReady to start
SVU-032Retire the 5 superseded PDFsLoadingReady to start
SVU-033Ingest the Kalpa Sūtra Unicode collection (Track A)CorpusAwaiting decision
SVU-034Citation extractor and quoted-text registerCorpusReady to start
SVU-035Prioritised acquisition list from the 303 gapsCorpusReady to start
SVU-036Tier-3 normalisation — GRETIL, Muktabodha, VHPCorpusPlanned
SVU-037Sentence segmentation and locus taggingCorpusPlanned
SVU-038Fix the locus identifier schemeArchitectureAwaiting decision
SVU-039Embeddings generated and loaded to QdrantRetrievalPlanned
SVU-040Knowledge graph v1 in Neo4jRetrievalPlanned
SVU-041MAP gold exemplars processedMAPAwaiting decision
SVU-042Bhāṣya cross-linked to Vyākaraṇa and Śāstra referencesMAPPlanned
SVU-043Mantra / Brāhmaṇa grouping assembledMAPPlanned
SVU-044Dynamic MAP analysis per mantraMAPPlanned
SVU-044.1Spec — MAP analysis contract — what the reader is shown and from whatMAPReady to start
SVU-044.2Build — Dynamic MAP analysis per mantraMAPPlanned
SVU-045Master list that absorbs new referencesCorpusIn progress
SVU-046Feed vaakya.vedanidhi.in during the workDeliveryPlanned
SVU-047Digitisation ledger and seven-stage pipeline commissionedOCRAwaiting decision
SVU-048Reconcile filing of recently added documentsHygieneReady to start

Stage 2 · Pada Kośa — 41 units, 3 complete

The lexical foundation: the conventional vocabulary drawn from the kośas, the Vedic vocabulary drawn from pada-pāṭha and bhāṣya, and the compositional vocabulary generated from Pāṇini's rules.

Unit Task Track Status
SVU-050Load Compartment A — the kośa corpusPada Kośa AAwaiting decision
SVU-051Kośa census publishedPada Kośa AComplete
SVU-052Access and serving policy for third-party lexical sourcesPada Kośa AAwaiting decision
SVU-053Kośa precedence order for the readerPada Kośa AAwaiting decision
SVU-054Flag the 106,014 single-attestation wordsPada Kośa AAwaiting decision
SVU-055Ingest all available Pada PāṭhasPada Kośa VAwaiting decision
SVU-056Pada Pāṭha ProcessorPada Kośa VAwaiting decision
SVU-057Generate Padapāṭha where none existsPada Kośa VPlanned
SVU-057.1Spec — Rules for generating Pada-pāṭha where none is attestedPada Kośa VReady to start
SVU-057.2Build — Generate Padapāṭha where none existsPada Kośa VPlanned
SVU-058Vaidika Pada Kośa — words and meanings from BhāṣyaPada Kośa VPlanned
SVU-058.1Spec — Method for extracting the Vaidika lexicon from BhāṣyaPada Kośa VReady to start
SVU-058.2Build — Vaidika Pada Kośa — words and meanings from BhāṣyaPada Kośa VPlanned
SVU-059Yaugika generator built and greenPada Kośa BComplete
SVU-060Svara print convention fixedPada Kośa BAwaiting decision
SVU-061Pracaya / ekaśruti decisionPada Kośa BAwaiting decision
SVU-062Accent-agreement harness against the gold corpusPada Kośa BAwaiting decision
SVU-062.1Spec — Accent-agreement test design and acceptance thresholdPada Kośa BAwaiting decision
SVU-062.2Build — Accent-agreement harness against the gold corpusPada Kośa BPlanned
SVU-063Form-space target decidedPada Kośa BAwaiting decision
SVU-064Full paradigm sweepPada Kośa BAwaiting decision
SVU-065FST lexicon compiledPada Kośa BAwaiting decision
SVU-066Multilingual morpheme glossesPada Kośa BPlanned
SVU-067Nirukta etymology engineNiruktaAwaiting decision
SVU-067.1Spec — Nirvacana model — sources, competing etymologies, presentationNiruktaReady to start
SVU-067.2Build — Nirukta etymology engineNiruktaPlanned
SVU-068Gender rules and the liṅga authoritiesVyākaraṇaPlanned
SVU-069Sandhi rule enginePāṭhaPlanned
SVU-069.1Spec — Sandhi rule inventory and svara behaviour at each juncturaPāṭhaReady to start
SVU-069.2Build — Sandhi rule enginePāṭhaPlanned
SVU-070Krama generationPāṭhaPlanned
SVU-070.1Spec — Krama construction rulesPāṭhaReady to start
SVU-070.2Build — Krama generationPāṭhaPlanned
SVU-071Eight Vikṛti formsPāṭhaPlanned
SVU-071.1Spec — The eight Vikṛti forms — construction rules per formPāṭhaReady to start
SVU-071.2Build — Eight Vikṛti formsPāṭhaPlanned
SVU-072Varṇa Krama generatorPāṭhaComplete
SVU-072.1Spec — Varṇa Krama decomposition rules from Śikṣā and PrātiśākhyaPāṭhaReady to start
SVU-072.2Build — Varṇa Krama generatorPāṭhaPlanned
SVU-073Lookup service and hover payloadPada KośaIn progress
SVU-074Lemma link between Veda words and their attestationsPada Kośa VAwaiting decision

Stage 3 · Synthesis, Models & Applications — 48 units, 0 complete

Cross-referencing, translation, the generative engines, and the reading application built on top of them.

Unit Task Track Status
SVU-080Cross-reference graph across all verticalsSynthesisPlanned
SVU-081Anuvāda — Bhāṣya-grounded translationSynthesisPlanned
SVU-081.1Spec — Anuvāda method — grounding every rendering in BhāṣyaSynthesisReady to start
SVU-081.2Build — Anuvāda — Bhāṣya-grounded translationSynthesisPlanned
SVU-082Samānatā — śabda and artha similarity engineSynthesisPlanned
SVU-082.1Spec — Samānatā — what counts as similarity in śabda and in arthaSynthesisReady to start
SVU-082.2Build — Samānatā — śabda and artha similarity engineSynthesisPlanned
SVU-083Vākya analysisSynthesisPlanned
SVU-084SFT dataset assembledModelPlanned
SVU-085Base model QLoRA fine-tuneModelPlanned
SVU-086RAG pipelineModelPlanned
SVU-087RLHF-lite preference passModelPlanned
SVU-088Re-benchmark the QLoRA epoch estimateModelAwaiting decision
SVU-089Śrauta Prayoga generationEnginePlanned
SVU-089.1Spec — Śrauta Prayoga sequence model and citation requirementsEngineReady to start
SVU-089.2Build — Śrauta Prayoga generationEnginePlanned
SVU-090Sāmaveda Stotra generationEnginePlanned
SVU-090.1Spec — Stotriyā assembly, viṣṭuti patterns and stobha rulesEngineReady to start
SVU-090.2Build — Sāmaveda Stotra generationEnginePlanned
SVU-091Vikṛti Pāṭha modelEnginePlanned
SVU-092Tarka sentence-structure modelEnginePlanned
SVU-092.1Spec — Śābda-bodha and pariṣkāra representation for TarkaEngineReady to start
SVU-092.2Build — Tarka sentence-structure modelEnginePlanned
SVU-093Varṇa Krama model — confirm whether wanted at allEngineAwaiting decision
SVU-094Sāma sound analysisEnginePlanned
SVU-094.1Spec — Svara extraction and notation conventions for SāmagānaEngineReady to start
SVU-094.2Build — Sāma sound analysisEnginePlanned
SVU-095Ūha engineEnginePlanned
SVU-095.1Spec — Ūha — how melody is re-framed onto new chandasEngineReady to start
SVU-095.2Build — Ūha engineEnginePlanned
SVU-096Mīmāṃsā Nyāya engineEnginePlanned
SVU-096.1Spec — Mīmāṃsā adhikaraṇa structure for machine analysisEngineReady to start
SVU-096.2Build — Mīmāṃsā Nyāya engineEnginePlanned
SVU-097Chatur-Darśana interpretation engineEnginePlanned
SVU-097.1Spec — The four lenses — scope and retrieval boundary of eachEngineReady to start
SVU-097.2Build — Chatur-Darśana interpretation engineEnginePlanned
SVU-098Layered text display L1–L5AppIn progress
SVU-099Grammar hover systemAppIn progress
SVU-100Prayoga Mode interfaceAppPlanned
SVU-101Vikṛti generator interfaceAppPlanned
SVU-102Sāma StudioAppPlanned
SVU-103Tarka Workbench and Varṇa Krama viewerAppPlanned
SVU-104Stotra Library and Nirukta ExplorerAppPlanned
SVU-105Cross-disciplinary notes layerAppPlanned
SVU-106vLLM serving with batching and cachingInfraPlanned
SVU-107Institutional pilotValidationPlanned
SVU-108Swadharma QA benchmarkValidationPlanned
SVU-109Prayoga validated against the 2026 Agniṣṭoma recordValidationPlanned

Stage 4 · Presentation & Outreach — 16 units, 0 complete

Publication, the researcher interface, and bringing the material to a general audience.

Unit Task Track Status
SVU-115sarvaveda.info updated with this planSiteReady to start
SVU-116Static pre-rendering for discoverabilitySitePlanned
SVU-117Contribution mechanism — in-page edit to pull requestSiteAwaiting decision
SVU-118Public web app with tiered accessReleasePlanned
SVU-119Researcher APIReleasePlanned
SVU-120Search-engine integration and structured dataReachPlanned
SVU-121Vedic perspective on demand, with societal referencesReachPlanned
SVU-122Video, audio and web publication pipelineMediaPlanned
SVU-123Webinar programmeMediaPlanned
SVU-124VVS textbooksMediaPlanned
SVU-125AI-produced reels across thousands of topicsMediaPlanned
SVU-126Daily social presenceMediaPlanned
SVU-127Daily śāstra perspective on contemporary eventsMediaPlanned
SVU-128Publish the consultation papersReleaseAwaiting decision
SVU-129Continuous learning cycleReleasePlanned
SVU-130Video corpus transcriptionCorpusPlanned
How to read this. A unit is complete only when its definition of done is met and a test asserts it. "Ready to start" means no decision and no prerequisite stands in the way. The plan is a living document and unit numbers are permanent, so this page can be compared honestly against itself over time.