Sanskrit Documents Atlas
← Back to tree

About

sanskritdocuments.org is a crowdsourced archive of mostly small, mostly devotional Sanskrit e-texts. It's been online since the 1990s, making it one of the oldest such collections on the web.

This unofficial “atlas” interface — not affiliated in any way — offers a better way to explore the contents of this collection. It is an "atlas" in the sense that it not only charts the navigable waters — with more transparent browsing by folder, topic, and category — but also plumbs their depths — with visibility into sizes, counts, and changes over time that the source collection doesn't surface on its own. It is not a mirror, as it hosts no text of its own; instead, every entry links back to sanskritdocuments.org for the actual content.

For more perspective on this and sibling Atlas projects, see Sāgarasaṅgama.

code: , on GitHub

Practical Intro to Sanskrit Documents Structure

Files and Folders

Each item on sanskritdocuments.org lives as an HTML file inside a particular folder on the server. For example: the श्रीदुर्गासप्तश्लोकी document exists as durga7.html in the folder doc_devii. Zero document files are duplicated across multiple folders.

There are twenty such document folders, and their names are all simple Roman strings prefixed with "doc_". 13 have simple genre or deity names like "doc_veda" and "doc_rama", while the other 7 are z_misc_* catch-alls, ostensibly prefixed to sort last, for example, "doc_z_misc_major_works."

Navigating to the folder path by itself displays a webpage listing every item in the folder, followed by additional links and information below. For example, /doc_devii/.

Visitors to sanskritdocuments.org are encouraged to browse the site's in-house collection via three main navigation menus: वर्ग ("by category"), देव ("by deity"), and मुख्य ग्रन्थ ("Major Sanskrit Works"). The sitemap provides an overview of these menus and their nesting structure. 136 overlapping menu items resolve to 87 actual pages. One might expect that these menu items link to the folder listings mentioned above, but this is not the case.

Topics

The intermediate structure between folders and menus — the pages that menus actually link to — is topics. In short:

Documents live in folders, topic pages provide cross-cutting views of thematically related documents, and menus provide cross-cutting views of topic pages.

Topic page URLs look like /sanskrit/*/, for example, /sanskrit/mantra/. Exactly like folder pages, they begin with a list of items and often end with a freeform block of additional links and information.

There are 85 such topics (plus 2 redundant ones; see Data Quality below). 13 of these share their name with a folder — the non-z_misc_* ones with a simple genre or deity name — for example, /sanskrit/devii and /doc_devii/. In these cases, the topic page usually contains all the documents listed on the corresponding location page, plus more from other locations. More generally speaking, topics link freely across any number of locations, from one to 16.

  • Which folders each topic draws on
    topic docs folders drawn from
    devii2,14214devii 95%; upanishhat 1%; shiva 1%; +11 more
    stotra2,06116vishhnu 31%; devii 21%; shiva 18%; +13 more
    shiva1,98714shiva 96%; upanishhat 1%; devii 1%; +11 more
    vishnu1,81211vishhnu 97%; deities_misc 1%; upanishhat 1%; +8 more
    deities_misc1,1268deities_misc 98%; vishhnu 1%; upanishhat 0%; +5 more
    ashtaka1,01615vishhnu 25%; deities_misc 22%; devii 22%; +12 more
    misc8525z_misc_general 96%; z_misc_misc 3%; deities_misc 1%; +2 more
    krishna7509vishhnu 90%; deities_misc 3%; devii 2%; +6 more
    gurudev7079deities_misc 95%; z_misc_general 2%; vishhnu 1%; +6 more
    dashamahavidya4404devii 98%; upanishhat 2%; deities_misc 0%; +1 more
    ganesha4137ganesha 94%; purana 3%; giitaa 1%; +4 more
    ashtottarashatanamavali37011devii 34%; deities_misc 18%; shiva 16%; +8 more
    raama3639raama 92%; devii 2%; upanishhat 2%; +6 more
    kavacha3209devii 47%; vishhnu 13%; shiva 12%; +6 more
    dashavatara2815vishhnu 98%; upanishhat 1%; z_misc_major_works 0%; +2 more
    upanishhat25311upanishhat 91%; shiva 3%; veda 2%; +8 more
    major_works2459z_misc_major_works 94%; shiva 2%; raama 2%; +6 more
    sahasranama24510devii 45%; shiva 18%; deities_misc 13%; +7 more
    ashtottarashatanama22612devii 31%; vishhnu 18%; shiva 16%; +9 more
    mantra21315vishhnu 24%; devii 22%; deities_misc 15%; +12 more
    navagraha2084z_misc_navagraha 99%; deities_misc 0%; veda 0%; +1 more
    sahasranamavali19112devii 41%; shiva 20%; vishhnu 12%; +9 more
    subrahmanya1886subrahmanya 95%; devii 3%; veda 1%; +3 more
    nadi1804devii 97%; z_misc_general 2%; raama 1%; +1 more
    panchaka17610deities_misc 28%; devii 16%; vishhnu 14%; +7 more
    shankaracharya17412z_misc_shankara 29%; devii 26%; shiva 13%; +9 more
    sarasvati1576devii 96%; z_misc_general 2%; vishhnu 1%; +3 more
    otherforms1552devii 99%; veda 1%
    advice15411shiva 79%; z_misc_general 9%; raama 3%; +8 more
    vedanta15112z_misc_major_works 24%; z_misc_general 21%; z_misc_shankara 19%; +9 more
    lakshmi1486devii 81%; vishhnu 13%; shiva 3%; +3 more
    durga1432devii 99%; subrahmanya 1%
    hanumaana1415hanumaana 97%; upanishhat 1%; giitaa 1%; +2 more
    gita1386giitaa 96%; z_misc_general 1%; ganesha 1%; +3 more
    parvati1384devii 89%; shiva 9%; z_misc_general 1%; +1 more
    suprabhata13711vishhnu 29%; devii 24%; deities_misc 20%; +8 more
    puja13411shiva 39%; devii 24%; vishhnu 7%; +8 more
    dattatreya1113deities_misc 92%; giitaa 5%; upanishhat 3%
    radha1043devii 93%; vishhnu 6%; upanishhat 1%
    veda968veda 88%; shiva 5%; vishhnu 2%; +5 more
    shataka9410vishhnu 22%; devii 20%; shiva 17%; +7 more
    arati7910devii 25%; vishhnu 24%; shiva 23%; +7 more
    stavaraja728devii 46%; vishhnu 18%; deities_misc 12%; +5 more
    sukta705veda 73%; devii 16%; deities_misc 6%; +2 more
    purana655purana 91%; z_misc_general 3%; deities_misc 3%; +2 more
    venkateshwara631vishhnu 100%
    hridaya5611devii 55%; z_misc_navagraha 9%; vishhnu 9%; +8 more
    pancharatna569devii 23%; deities_misc 20%; shiva 18%; +6 more
    ayyappa522deities_misc 98%; giitaa 2%
    sahitya526z_misc_general 87%; yoga 4%; vishhnu 4%; +3 more
    lalita493devii 96%; shiva 2%; upanishhat 2%
    gayatri484devii 94%; upanishhat 2%; vishhnu 2%; +1 more
    bhujanga458devii 22%; vishhnu 22%; subrahmanya 18%; +5 more
    malamantra409devii 25%; vishhnu 22%; deities_misc 15%; +6 more
    shatkam399deities_misc 18%; devii 18%; vishhnu 15%; +6 more
    shatchakrashakti361devii 100%
    kriti339devii 21%; deities_misc 18%; shiva 18%; +6 more
    trishati335devii 45%; ganesha 21%; subrahmanya 18%; +2 more
    yoga336yoga 76%; shiva 9%; upanishhat 6%; +3 more
    raksha3010vishhnu 23%; shiva 23%; deities_misc 13%; +7 more
    kalidasa292z_misc_major_works 83%; devii 17%
    lahari298devii 55%; z_misc_major_works 10%; deities_misc 10%; +5 more
    vishnu_misc291vishhnu 100%
    tulasi284devii 86%; vishhnu 7%; giitaa 4%; +1 more
    shatinamavali276devii 33%; shiva 19%; subrahmanya 15%; +3 more
    jyotisha262z_misc_sociology_astrology 96%; shiva 4%
    kamakshi262devii 96%; z_misc_general 4%
    subhaashita263z_misc_general 50%; z_misc_subhaashita 38%; z_misc_major_works 12%
    gitam227giitaa 59%; shiva 9%; raama 9%; +4 more
    renuka201devii 100%
    ekashloki188z_misc_general 28%; z_misc_navagraha 22%; raama 22%; +5 more
    minakshi161devii 100%
    sutra165z_misc_major_works 50%; shiva 25%; devii 12%; +2 more
    tantra168devii 56%; subrahmanya 6%; z_misc_general 6%; +5 more
    bhagavadgita121giitaa 100%
    atharvashirsha95deities_misc 22%; upanishhat 22%; ganesha 22%; +2 more
    adhyatmaramayana81raama 100%
    bhajana83deities_misc 75%; vishhnu 12%; raama 12%
    varnamala74shiva 57%; devii 14%; raama 14%; +1 more
    aparadhakshama53devii 40%; deities_misc 40%; ganesha 20%
    shloka41z_misc_general 100%
    aksharamala32shiva 67%; devii 33%
    amarakosha31z_misc_major_works 100%
    natyashastra31z_misc_major_works 100%
    samajashastra21z_misc_sociology_astrology 100%

That is, unlike folders, which possess documents exclusively, topics link to them promiscuously. A given document appears on one to 9 topic pages, 2.08 on average. For example, Śivāṣṭakam is on both /sanskrit/shiva/ and /sanskrit/ashtaka/.

Categories

Each document also carries a comma-separated Category line in the metadata block at the foot of its page, with 2.88 values on average. There are 366 total distinct values across all documents. It is tempting to read this as yet another load-bearing site structure. It is not. These categories are plain text at the bottom of a document.

After spelling normalization (see below), 317 unique categories remain. Virtually any concept is fair game: genres like stuti or nAmAvalI; persons like shrIdharasvAmI or vAsudevAnanda-sarasvatI (authors), or hkmeher (a contributor); works like shivarahasya (entire) or skandapurANa or mudgalapurANa (extract sources).

Similar to how 13 topics share a name with a folder, 85 categories shadow a topic; as one might expect, nearly all the documents linked from a given topic carry the corresponding category value. The remaining 232 categories that do not match a topic name can be thought of as proto-topics which could be promoted to have a full topic page of their own at any time. The most common example of this latter category type, shivarahasya, is used to tag no fewer than 784 documents.

Metadata Fields

There are 17 metadata fields, but not every document uses every field.

  • Metadata field coverage
    field documents coverage what it holds
    title9,756100.0%The English label at the head of every document page.
    language9,756100.0%Sanskrit for all but a hundred or so.
    categories9,756100.0%The free-text tag bag.
    title_deva9,68099.2%The Devanāgarī title — the only field that can be transliterated.
    latest_update9,752100.0%A free-text date string.
    subject9,74499.9%Mostly just "philosophy/hinduism/religion," not helpful.
    proofread_by8,56987.8%The sole name on most documents that carry it, so ostensibly the inputter.
    indexextra7,93281.3%The parenthetical HTML links (archive.org scans, etc.)
    description6,17863.3%A free-text note, often naming the source edition.
    transliterated_by4,33544.4%Ostensibly who converted to ITRANS, usually same as proofread_by.
    author3,67737.7%Free-text names, missing even for many works whose author is well known.
    subdeity3,53936.3%Deity divisions (e.g. krishna, gurudev) within folders (e.g., doc_vishhnu.)
    acknowledge_permission1,47415.1%Whose permission the reproduction rests on, where the site names one.
    texttype1,46515.0%An unhelpful free-text genre tag.
    source5285.4%The printed edition or URL the text was taken from.
    translated_by3213.3%Translator name, for some docs containing translation.
    subcategory1251.3%A rare second tag under category.

The two fields subject and texttype are worth distinguishing from category. subject, where it is used, carries the single value "philosophy/hinduism/religion" 92.3% of the time, and nearly all the rest are permutations of those same three words. As for texttype, it basically just restates information categories already covers, but more coarsely. In short, neither is useful.

No field clearly states who first input a text. proofread_by is the sole name on 4,395 documents, and someone must have input every document, so it seems reasonable to read this as the inputter, not just a checker. transliterated_by then appears to be a second hand where one was involved; of the 4,174 documents with both fields, 85.7% name the same person twice.

Scanned Sources

On topic and folder pages, many document lines include parentheticals with links, and some of these links lead to scans (PDFs) hosted in various places: archive.org (most common), other third-party sites (less common), or the server's own /scannedbooks/ route (least common).

A linked scan is not necessarily the exemplar the e-text was encoded from. The site itself makes no such claim. However, many of them do seem to match. User beware.

Of the external links likely to be exemplar matches, 5,535 links across 5,501 documents reduce to 873 distinct sources which can survive a liveness check; 117 other such scan links are broken.

Transliteration and Generated PDFs

Each item is stored as an ITRANS plaintext original (conventionally using the .itx filename extension). From there, it is rendered automatically into HTML, with transliteration into different schemes performed by Aksaramukha.js — except on 74 legacy pages still served by an older Sanscript-based transliterator stack (sanscript-vedic.js + sanskritdocuments-transliterator-vedic.js). This transformed HTML can be printed to PDF using a "Print" link on the document page.

In addition, nearly every document on the site is also published as a series of formatted, static PDFs containing the same content transliterated with sanscript in five specific Indic scripts: Devanagari, Tamil, Telugu, Kannada, and Gujarati. The prominent "PDF" icons on the site are all these, and they are purely derivative and for reading convenience. They should be sharply distinguished from the more philologically significant links to scanned sources.

Three Names per Document, and Only One Transliterates

Every document carries up to three distinct forms of its own name, and they do not agree: a free-text English label (Akshara Ramayanam — ad hoc Romanization, no scheme, no diacritics), a proper Devanāgarī title (अक्षररामायणम्), and the filename stem (akShararAmAyaNam), which looks like ITRANS but isn't.

Language

The collection is overwhelmingly Sanskrit: 9,646 of 9,756 documents are marked so (98.9%), with the remainder mostly Hindi and Marathi, plus a handful of Gujarati and mixed entries. 14 distinct language strings appear in all.

Obtaining and Representing Sanskrit Documents Data

Scraping

The atlas is built from a snapshot periodically scraped from the site: about 10,000 requests, one request every two seconds, approximately six hours. There are four steps:

  1. Fetch 1 sitemap and parse for 87 topic pages.
  2. Fetch 87 topic pages and parse for 9,781 document pages (and from their paths, 20 folder pages)
  3. Fetch 20 folder pages (for audit/corroboration purposes only).
  4. Fetch 9,781 document pages.

9,756 documents are held here against the 9,781 the site's own topics name, as the site served them on 2026-08-15. The shortfall is 25 dead pages which 404 on the site; each is named under Data Quality below.

Folders, Topics, and Categories

The Atlas presents folders and topics as two alternative browsing methods. The 2 redundant topics — /sanskrit/giitaa/ and /sanskrit/vishhnu/ — are suppressed in favor of /sanskrit/gita/ and /sanskrit/vishnu/.

Unlike the source site, the Atlas also presents categories as a usable filter, which can optionally be overlaid on either of the two browsing methods. For this purpose, spelling irregularities in category values are normalized. Some normalization is easy: dvAdhasha, pAravtI, and kavaha are outright typos, and it's clear what to merge them into. Similarly, alternate transliterations devI, devi and devii, or upanishad, upanishhat, and upaniShat are easily merged. Some normalization is trickier: vishhnu can merge into vishnu, but vishnu_misc is distinct. Similarly, gItA can merge into giitaa, but gitam is distinct.

Calculating Size

The source site does not report sizes. The Atlas calculates a number for each page by taking only the <pre> block — excluding site navigation elements and also the metadata block — and transliterating its contents to IAST.

Most of the corpus consists of small documents: the median document is about 2.7 KB, while the पुराणs and Monier-Williams run past 2 MB each, so just 14 texts carry a sixth of the total 126.5 MB.

Since folders are mutually exclusive in their contents, their respective totals add up to the grand total without complication. By contrast, topics are overlapping in their contents, so their respective totals are simply not used to calculate the grand total.

Data Quality: Structural Errors

Opportunities for Correction

The Atlas codebase also includes an audit pipeline for detecting and concisely indicating problems that should be fixed on sanskritdocuments.org. Every figure below is re-derived by the run that publishes it.

Data Quantity: Historical Trends

measure period totals
Loading…

Every document records its own Latest update, and all but 9 of 9,756 can be placed in a year. More informative would be date of addition, but the site does not record this, so the latest-update date is the best measure of the collection's growth. This means that any modified document drifts forward in time — leaving the year it arrived and landing in the year it was last edited — and the graph as a whole is most likely greatly skewed toward the present, but it's not possible to know by exactly how much.

Partial correctives for this error do exist: two earlier copies of the corpus, each proving that whatever a document's stamp says, it was already there when the copy was taken. One is a complete copy of the site from June 2020 on GitHub; the other is an earlier scrape by the author of this Atlas, from February 2025. Attestation can only ever move a document earlier, never later, so a second copy adds corrections without undoing any. Between the two they pull 153 documents back to a more correct date — 42% of them to somewhere before mid-2020, the remaining 58% into the years between — and they put a date on 9 of the documents whose own field is unusable (see above).