SarvaVeda LLM ← Home
The complete plan

The Full Roadmap

The complete technical and institutional plan behind SarvaVeda — the corpus, the machine, the data pipeline, the model architecture, the application, the phased build, the team, and the risks. A living document; figures indicative.

॥ अनन्ता वै वेदाः ॥
Section 1

Vision & scope

SarvaVeda is a self-hosted foundation platform that encodes, reasons over, and generates from the entirety of the Sanātana Dharma textual universe — grounded in the received text, validated by scholars, and owned by the institution.

The knowledge universe — four concentric layers

Eight functional objectives

Section 2

Hardware architecture

Self-hosted on two GPU servers at VVV Mysore — full data sovereignty over the sacred corpus, operating as a primary-training plus inference/validation cluster.

ComponentSpecification
Servers2 × GPU servers (primary training + inference/validation)
RAM (each)256 GB DDR5 ECC
GPU (recommended)2–4 × NVIDIA A100 / H100 80 GB per node
Storage (training)60 TB NVMe RAID-6 — corpus, tensors, checkpoints
Storage (inference)8 TB NVMe SSD — weights + hot cache
Network100 GbE interconnect for distributed training
OS / stackUbuntu 22.04 LTS · CUDA 12.x · cuDNN 9.x
ResilienceShared MinIO/Ceph filesystem · weekly offline backup · 10 kVA online UPS
Phased start. Where budget constrains the full A100/H100 configuration, training can begin on 2 × NVIDIA RTX 4090 (24 GB) per node for 7B–13B models, upgrading GPUs as the project scales to the full 70B foundation model. The 256 GB system RAM suffices for all sizes up to 70B at 8-bit quantisation.
Section 3

Corpus assembly & data pipeline

Four trust tiers

All training data is classified by trust, which governs its weight during fine-tuning.

TierContents
GoldThe 72 VVV volumes (~15,000 pp) — editorially verified, IAST-correct. Highest weight.
SilverThe 2,000-topic Swadharma list; curated QA banks (1,400+ questions); the SGSP Kalpa-sūtra Unicode collection (~30 DOCX); existing MAP-format exemplars.
BronzeDigitised mūla from verified sources (GRETIL, Muktabodha, Vedic Heritage); traditional bhāṣya (Sāyaṇa, Mādhava, Uvvaṭa, Skandasvāmin).
Reference1,000 hours of transcribed lecture video (after ASR); related academic works. Contextual weight.

Existing holdings — the seed corpus

SGSP already holds a substantial digitised collection, in two processing tracks:

The seven-stage industrial OCR pipeline

Text preparation & the citation principle

Every document also passes a seven-stage preparation flow — ingest → normalise → segment → align → annotate → knowledge-graph → quality-gate — with text preserved exactly, including VijayaDV PUA and Vedic accents. A citation extractor parses all 72 volumes to compile the register of every text the Series quotes, guaranteeing the corpus covers the sources the model is most often asked about.

Section 4

Model architecture & training

Rather than training from scratch, SarvaVeda uses retrieval-augmented generation plus domain fine-tuning — the right instrument for a high-precision, fidelity-first domain.

The MAP pipeline

Every query passes through Meaning-retrieval (semantic match of the top passages), Analysis (grammar + knowledge graph), and Presentation (the model assembles a layered, mode-appropriate response).

Nine specialised engines

Multilingual generation

Six languages in Phase 1 (Sanskrit, Hindi, Telugu, Kannada, English, plus IAST). All meaning is generated from Sanskrit as the pivot, then translated — so the tradition's semantic precision is never lost to a direct-to-English rendering.

Section 5

Application layer

Stack: React + TypeScript + Tailwind frontend; FastAPI backend; Qdrant (vectors); Neo4j (knowledge graph); vLLM / Ollama serving; Keycloak (role-based access); Typesense (lexical search).

The layered reader

Scholar workbenches

Dedicated interfaces for Prayoga generation, the Vikṛti generator, the Sāma Studio (stotra + sound analysis + Ūha workbench), the Tarka Workbench and Varṇa-Krama viewer, the four-column Chatur-Darśana view, the Stotra Library and Nirukta Explorer, and cross-disciplinary notes.

Section 6

Phased implementation — ~30–36 months

Phase 0
Months 1–4

Foundation & infrastructure

  • Procure, rack and commission the GPU servers, UPS, cooling and network.
  • Install the CUDA stack; deploy the vector and graph databases and object storage.
  • Full inventory audit of all sources; constitute the 4–6 member scholar review panel.
Phase 1
Months 3–12

Corpus assembly & pipeline

  • Ingest the 72 volumes and the Kalpa-sūtra Unicode collection first; segment and tag.
  • Ingest the complete Ṣaḍaṅga, all four Kalpa classes, the 18 Mahāpurāṇas and the Stotra corpus.
  • Generate the Vyākaraṇa FST lexicon v1; deploy grammar and Nirukta annotation.
  • Run the citation extractor; launch the industrial OCR programme (15–25 books/month).
Phase 2
Months 8–16

Foundation-model fine-tuning

  • Prepare the supervised dataset from the Gold and Silver tiers.
  • QLoRA fine-tune the 70B base model; stand up the RAG pipeline.
  • Scholar-graded evaluation benchmark and a preference-feedback training pass.
Phase 3
Months 14–26

The nine-engine suite & application

  • Train the Prayoga, Stotra, Vikṛti, Tarka, Varṇa-Krama and Nirukta engines.
  • Build the Chatur-Darśana four-lens engine and the Sāma sound-analysis model v1.
  • Ship the web application for internal scholarly use; deploy batched inference.
Phase 4
Months 20–28

Institutional pilot & validation

  • Pilot with SVAMI faculty and invited scholars (25–50 users).
  • Score all 2,000 Swadharma topics against scholar reference answers.
  • Validate Prayoga against the 2026 Agniṣṭoma records; build and test the Ūha engine.
Phase 5
Months 26–36

Public release & continuous learning

  • Launch tiered public access — general, scholar, institutional — with an API for partners.
  • Onboard the diaspora network of pāṭhaśālās and centres as first subscribers.
  • Deploy recitation-video transcription; establish quarterly retraining.
Section 7

Team & governance

Roles

Ownership & licence

SarvaVeda is a project of Sanatana Guru Sampradaya Pratishthanam (SGSP), Mysore. All model weights, training data and code are the intellectual property of SGSP. The foundation model will be released under a restricted, non-commercial research licence for accredited Vedic institutions and universities. Dharma Poshanam Inc. (USA 501(c)(3)) is the vehicle for tax-deductible international support.

Section 8

Risks & mitigations

RiskMitigation
GPU supply-chain delayEstablished Indian vendors, or cloud GPU for early training while hardware ships.
Digitisation errors (Bronze tier)10% scholar spot-check gate; an automatic IAST/diacritic validator catches most encoding errors.
Hallucination in ritual outputPrayoga always cites its source sūtra; any uncited step is flagged unverified and requires scholar review before use.
Vedic svara mishandlingAccent handling is rule-based, not ML — derived from the Prātiśākhya rules, eliminating hallucination in accent marking.
Loss of institutional knowledgeEvery scholarly decision is recorded in a versioned annotation database; provenance is fully traceable.
Cost overrunEarly phases run on existing hardware plus consumer GPUs; the full cluster upgrade is deferred to a later phase.