Section 1
Vision & scope
SarvaVeda is a self-hosted foundation platform that encodes, reasons over,
and generates from the entirety of the Sanātana Dharma textual universe — grounded in
the received text, validated by scholars, and owned by the institution.
The knowledge universe — four concentric layers
- Core canonical texts — the four Vedas in Padapāṭha and Saṃhitāpāṭha; the Brāhmaṇas, Āraṇyakas and Upaniṣads; the Śrauta and Gṛhya Sūtras; the Prātiśākhyas. ~100,000 pages of primary text with traditional bhāṣya.
- The complete Ṣaḍaṅga — all six Vedāṅgas in full (Śikṣā, Vyākaraṇa, Chandas, Nirukta, Jyotiṣa, Kalpa), the four Kalpa-sūtra classes (Śrauta, Gṛhya, Dharma, Śulba), and the Upāṅgas. ~100,000 further pages.
- The complete Purāṇa corpus — all 18 Mahāpurāṇas and the principal Upapurāṇas, segmented to adhyāya and śloka, with khaṇḍa/saṃhitā structure.
- Stotra & ancillary — the full Stotra literature; Itihāsa, Āgama and Upaveda texts; and the 2,000-topic Swadharma encyclopaedia. ~1 million pages overall, plus 1,000 hours of recitation and lecture video.
Eight functional objectives
- 1 — Complete knowledge store: all Veda, Śāstra and Purāṇa, retrievable to the Ṛk / unit / pada level.
- 2 — Canonical input processing through the MAP (Meaning–Analysis–Presentation) pipeline.
- 3 — Grammar-aware display: Pāṇinian analysis on hover, without disrupting the reading.
- 4 — Layered delivery: mūla, Padapāṭha, Bhāṣyam and meaning together, with linked passages.
- 5 — Semantic interlinking: a knowledge graph of intertextual relationships.
- 6 — Multilingual generation: Sanskrit, Hindi, Telugu, Kannada, English.
- 7 — Cross-disciplinary annotation: bridges to science, mathematics, philosophy and the arts.
- 8 — Ritual generation: Śrauta Prayoga, the eight Vikṛti forms, and Sāma renderings.
Section 2
Hardware architecture
Self-hosted on two GPU servers at VVV Mysore — full data sovereignty over the sacred
corpus, operating as a primary-training plus inference/validation cluster.
| Component | Specification |
| Servers | 2 × GPU servers (primary training + inference/validation) |
| RAM (each) | 256 GB DDR5 ECC |
| GPU (recommended) | 2–4 × NVIDIA A100 / H100 80 GB per node |
| Storage (training) | 60 TB NVMe RAID-6 — corpus, tensors, checkpoints |
| Storage (inference) | 8 TB NVMe SSD — weights + hot cache |
| Network | 100 GbE interconnect for distributed training |
| OS / stack | Ubuntu 22.04 LTS · CUDA 12.x · cuDNN 9.x |
| Resilience | Shared MinIO/Ceph filesystem · weekly offline backup · 10 kVA online UPS |
Phased start. Where budget constrains the full A100/H100
configuration, training can begin on 2 × NVIDIA RTX 4090 (24 GB) per node for 7B–13B
models, upgrading GPUs as the project scales to the full 70B foundation model. The 256 GB
system RAM suffices for all sizes up to 70B at 8-bit quantisation.
Section 3
Corpus assembly & data pipeline
Four trust tiers
All training data is classified by trust, which governs its weight during fine-tuning.
| Tier | Contents |
| Gold | The 72 VVV volumes (~15,000 pp) — editorially verified, IAST-correct. Highest weight. |
| Silver | The 2,000-topic Swadharma list; curated QA banks (1,400+ questions); the SGSP Kalpa-sūtra Unicode collection (~30 DOCX); existing MAP-format exemplars. |
| Bronze | Digitised mūla from verified sources (GRETIL, Muktabodha, Vedic Heritage); traditional bhāṣya (Sāyaṇa, Mādhava, Uvvaṭa, Skandasvāmin). |
| Reference | 1,000 hours of transcribed lecture video (after ASR); related academic works. Contextual weight. |
Existing holdings — the seed corpus
SGSP already holds a substantial digitised collection, in two processing tracks:
- Track A — Unicode DOCX (immediate): ~30 files including the complete classical Śrauta register (11 Śrauta Sūtras) and 8+ Gṛhya Sūtras; commentated editions that supply mūla↔bhāṣya alignment pairs; and MAP-format exemplars that serve as the gold output templates the model learns from.
- Track B — scanned library (OCR): hundreds of critical and commentary editions across the śāstra spectrum — an estimated 100,000–200,000 pages.
The seven-stage industrial OCR pipeline
- O1 Intake & triage — every book logged in a digitisation ledger (edition, script, layout class, scan grade).
- O2 Preprocessing — deskew, dewarp, denoise, binarise, split spreads.
- O3 Layout analysis — classify zones (mūla, commentary, footnote, apparatus) so text is never read across columns.
- O4 Script-aware OCR — the best engine per script (Devanagari, Grantha, Telugu, Kannada, Roman-IAST), recorded per book.
- O5 Vedic svara recovery — accent marks (which generic OCR silently drops) detected as graphical objects and re-inserted, or transferred from a verified source.
- O6 Lexicon-guided post-correction — the Vyākaraṇa FST doubles as the most complete Sanskrit spell-checker ever built, flagging any non-analysable form.
- O7 Scholar quality gate — 10% sample graded per book; anything below 98% akṣara accuracy loops back. Accepted text is committed with full provenance.
Text preparation & the citation principle
Every document also passes a seven-stage preparation flow — ingest → normalise → segment →
align → annotate → knowledge-graph → quality-gate — with text preserved exactly, including
VijayaDV PUA and Vedic accents. A citation extractor parses all 72 volumes to compile the
register of every text the Series quotes, guaranteeing the corpus covers the sources the
model is most often asked about.
Section 4
Model architecture & training
Rather than training from scratch, SarvaVeda uses retrieval-augmented generation plus
domain fine-tuning — the right instrument for a high-precision, fidelity-first domain.
- Base model: Llama 3.1 70B or Mistral Large 2 (128k context); Gemma 2 27B as a lighter alternative.
- Fine-tuning: QLoRA with 4-bit NF4 quantisation for training, 8-bit for inference.
- Retrieval: a vector database over the full corpus grounds every generation in retrieved source passages.
The MAP pipeline
Every query passes through Meaning-retrieval (semantic match of the top passages),
Analysis (grammar + knowledge graph), and Presentation (the model
assembles a layered, mode-appropriate response).
Nine specialised engines
- Śrauta Prayoga — complete rite sequences with mantra, svara and ṛtvij assignment; every step citing its sūtra.
- Sāmaveda Stotra — stotra assembly (stotriyā, viṣṭuti, stobha) for a given Soma rite, in traditional notation.
- Vikṛti Pāṭha — all eight forms (Jaṭā, Ghana, …) with correct sandhi and svara at every junctura.
- Tarka sentence structure — formal pañcāvayava and Navya-Nyāya analysis; powers śāstrārtha reasoning.
- Varṇa Krama — akṣara-by-akṣara decomposition with per-syllable svara and mātrā.
- Sāma Sound Analysis (audio) — pitch/svara extraction (CREPE / wav2vec 2.0) → automatic notation → Ūha re-framing onto new chandas.
- Vyākaraṇa vocabulary engine — a rule-based Pāṇinian generator producing the full derived lexicon (>100 million forms) in a finite-state transducer, each form carrying its prakriyā.
- Nirukta etymology — Yāska-style nirvacana with the Pāṇinian vyutpatti shown alongside, context-resolved for polysemy.
- Chatur-Darśana engine — one sentence interpreted in parallel through Tarka, Vyākaraṇa, Mīmāṃsā and Vedānta, with a synthesis of where they converge and differ (e.g. svargakāmo yajeta).
Multilingual generation
Six languages in Phase 1 (Sanskrit, Hindi, Telugu, Kannada, English, plus IAST). All meaning
is generated from Sanskrit as the pivot, then translated — so the tradition's
semantic precision is never lost to a direct-to-English rendering.
Section 5
Application layer
Stack: React + TypeScript + Tailwind frontend; FastAPI backend; Qdrant (vectors);
Neo4j (knowledge graph); vLLM / Ollama serving; Keycloak (role-based access); Typesense (lexical search).
The layered reader
- Level 1 (always visible): Devanagari mūla + IAST + English meaning.
- Level 2: Padapāṭha with word-by-word breakdown.
- Level 3: Bhāṣyam excerpt with translation.
- Level 4 (hover any word): dhātu, pratyaya, vibhakti, Pāṇini-sūtra reference — pre-computed, sub-100 ms, no model call.
- Level 5 (sidebar): related mantras, Brāhmaṇa passages, connected Śāstra episodes and Purāṇa parallels from the knowledge graph.
Scholar workbenches
Dedicated interfaces for Prayoga generation, the Vikṛti generator, the Sāma Studio (stotra +
sound analysis + Ūha workbench), the Tarka Workbench and Varṇa-Krama viewer, the four-column
Chatur-Darśana view, the Stotra Library and Nirukta Explorer, and cross-disciplinary notes.
Section 6
Phased implementation — ~30–36 months
Foundation & infrastructure
- Procure, rack and commission the GPU servers, UPS, cooling and network.
- Install the CUDA stack; deploy the vector and graph databases and object storage.
- Full inventory audit of all sources; constitute the 4–6 member scholar review panel.
Corpus assembly & pipeline
- Ingest the 72 volumes and the Kalpa-sūtra Unicode collection first; segment and tag.
- Ingest the complete Ṣaḍaṅga, all four Kalpa classes, the 18 Mahāpurāṇas and the Stotra corpus.
- Generate the Vyākaraṇa FST lexicon v1; deploy grammar and Nirukta annotation.
- Run the citation extractor; launch the industrial OCR programme (15–25 books/month).
Foundation-model fine-tuning
- Prepare the supervised dataset from the Gold and Silver tiers.
- QLoRA fine-tune the 70B base model; stand up the RAG pipeline.
- Scholar-graded evaluation benchmark and a preference-feedback training pass.
The nine-engine suite & application
- Train the Prayoga, Stotra, Vikṛti, Tarka, Varṇa-Krama and Nirukta engines.
- Build the Chatur-Darśana four-lens engine and the Sāma sound-analysis model v1.
- Ship the web application for internal scholarly use; deploy batched inference.
Institutional pilot & validation
- Pilot with SVAMI faculty and invited scholars (25–50 users).
- Score all 2,000 Swadharma topics against scholar reference answers.
- Validate Prayoga against the 2026 Agniṣṭoma records; build and test the Ūha engine.
Public release & continuous learning
- Launch tiered public access — general, scholar, institutional — with an API for partners.
- Onboard the diaspora network of pāṭhaśālās and centres as first subscribers.
- Deploy recitation-video transcription; establish quarterly retraining.
Section 7
Team & governance
Roles
- Principal Investigator — overall vision, corpus authority, Śrauta ritual validation.
- Technical Director (AI/ML) — architecture and delivery (recruiting).
- Vedic NLP Lead — Pāṇinian grammar annotation design; mūla↔bhāṣya alignment.
- Sāmaveda & Prayoga Lead — Sāma training data and Prayoga validation.
- Data Pipeline Engineer — ingestion, embeddings, training infrastructure (recruiting).
- OCR Pipeline Operator — the industrial digitisation programme (recruiting).
- Frontend Developer — the reader and workbenches, with Indic-script rendering (recruiting).
- Scholar Review Panel — 4–6 VVV/SVAMI faculty validating data and grading outputs.
- Institutional Advisor — fundraising and diaspora partnerships.
Ownership & licence
SarvaVeda is a project of Sanatana Guru Sampradaya Pratishthanam (SGSP), Mysore. All model
weights, training data and code are the intellectual property of SGSP. The foundation model
will be released under a restricted, non-commercial research licence for
accredited Vedic institutions and universities. Dharma Poshanam Inc. (USA 501(c)(3)) is the
vehicle for tax-deductible international support.
Section 8
Risks & mitigations
| Risk | Mitigation |
| GPU supply-chain delay | Established Indian vendors, or cloud GPU for early training while hardware ships. |
| Digitisation errors (Bronze tier) | 10% scholar spot-check gate; an automatic IAST/diacritic validator catches most encoding errors. |
| Hallucination in ritual output | Prayoga always cites its source sūtra; any uncited step is flagged unverified and requires scholar review before use. |
| Vedic svara mishandling | Accent handling is rule-based, not ML — derived from the Prātiśākhya rules, eliminating hallucination in accent marking. |
| Loss of institutional knowledge | Every scholarly decision is recorded in a versioned annotation database; provenance is fully traceable. |
| Cost overrun | Early phases run on existing hardware plus consumer GPUs; the full cluster upgrade is deferred to a later phase. |