Sanskrit Wikisource Atlas
← Back to tree

About

This web interface offers a better way to explore the Sanskrit e-text content at sa.wikisource.org, with transliteration support, dynamic display of categories (see "वर्गसर्वस्वम्"), more intuitive abstraction to the "text" level, and clear size and date information. It also surfaces insights that will help to gradually improve the source collection itself. The data snapshot is refreshed semi-automatically on a monthly basis.

This is an "atlas" in the sense that it both charts the navigable waters and plumbs their depths, with sizes, counts, and growth dynamics that the source collection doesn't surface on its own.

code: , on GitHub

data last sourced:

Practical Intro to Wikisource Structure

Graph of Equal Pages

Wikisource's most basic ontological unit is the page. Pages can freely link to each other, creating a directed graph, with no inherent hierarchy. In order to make one page link to multiple others, as in a table of contents, all connections must be made manually. Nothing automatically creates such lists, and there is no distinct "subpage" node or relation type.

Pages that are conceptually a "child" or "subpage" can be made to automatically link back to the relevant "parent" page, namely by carefully constructing the child page name so that it consists of the parent's name, then a slash, then the unique child page name. For example, Wikisource will automatically parse the page name "अग्निपुराणम्/अध्यायः १" and place at the top of that page a link to the related page "अग्निपुराणम्." In Wikisource jargon, these auto-generated links are called "breadcrumbs."

Categories: Another Graph

The next basic ontological unit is the Category. Pages can be labeled with any number of Categories. For each Category, Wikisource creates a display page that shows all associated items, e.g., नाटकानि.

Categories of course also relate to each other, creating their own directed graph. By convention, this graph should be constructed as a simple tree, with a single root, exactly one parent per node, and nothing left out. However, since categorization is done manually, this is not guaranteed; one Category may have multiple parents, and there can theoretically even be cycles.

Namespaces

Distinct ontological types exist in different "namespaces." For example, regular content pages are in namespace 0 (technically called "Main"), and Categories constitute namespace 14. This makes it easy to use computational methods to obtain all items of a given type.

OCR "Proofreading" Workflow

Filling in the text content of a Mainspace page is often done manually. Another option provided by Wikisource is the "Proofreading" workflow which is based on OCR. A user uploads an image file, which gets stored as part of a central "Index" item (namespace 106, Sanskrit "अनुक्रमणिका"). For example, काठकोपनिषत्. Based on file metadata, Wikisource then creates a table with links for each page.

At those links, Wikisource creates "Page" items (namespace 104), inherently linked to the Index, and OCRs them with Tesseract. Contributors then use the facing-page editing interface to iteratively proofread each page's text against the source image. When satisfied with a given page, they manually increment the color-coded progress indicator (red → yellow → green).

By convention, Index items can be labeled with Categories, but Pages should not be, in order to avoid clutter when viewing a given Category's associated items.

Transclusion

Content of one item can be virtually injected into the content of another by a process known as "transclusion," similar to symbolic linking. Most notably, a user-facing Mainspace page can transclude relevant ranges of proofreading "Pages." Proofreading can then continue gradually, and the Mainspace page will update automatically.

However, nothing guarantees such use of transclusion. Instead, an editor may well copy-paste the content of OCR-workflow Pages into Mainspace pages, in which case proofreading corrections will no longer propagate automatically. Mainspace pages visually indicate active transclusion with special marginal links to respective Pages (e.g., the "[१]" in this अमरकोषः item).

Obtaining and Representing Sanskrit Wikisource Data

Snapshots

The main page of this Atlas is regularly updated to reflect the latest monthly snapshot drawn from Wikimedia's own "Content File Export" dump for sa.wikisource.org.

By contrast, the historical changelog on this About page (see Data Quantity section below) reaches back to 2012. Two live rolling windows cover the recent past; everything older is reconstructed month by month from Wikimedia's full revision history, taking each page as it stood on the first of that month:

materialized (synthetic) legacy-format live current-format live

Although Internet Archive does also hold additional legacy-format dumps for 74 of the 95 months between 2014-07 and 2022-05, these suffer from the relatively poor organization of past years. By contrast, materialized data benefits from the relatively better organization of the present, which makes it easier to abstract to the text level. For this reason, Internet Archive legacy-format data dumps are not used here.

Pages, Categories, and Texts

After fetching full data for a given point in time, the Main, Category, Index, and Page namespaces are explored in full.

On Mainspace pages, breadcrumbs are mined to establish parent/child relationships. These back-connections allow rolling up child subpages for more concise display. Redirects are followed for this purpose, e.g., the effective parent of रामायणम्/बालकाण्डम् is वाल्मीकिरामायणम्.

Category parent/child relationships are read from Category pages, and the Atlas's top-level category display is created primarily from children of the Category वर्गसर्वस्वम्/ग्रन्थाः (plus the miscategorized वर्गसर्वस्वम्/धर्मशास्त्रम्). Other children of वर्गसर्वस्वम् that serve Wikisource housekeeping functions are excluded here entirely.

Secondarily, any Categories not part of the ग्रन्थाः (+ धर्मशास्त्रम्) central community, along with uncategorized pages, are placed into an artificial category called असम्बद्धवर्गीकृतम् ("improperly categorized"), which can gradually serve to improve the Category structure of sa.wikisource itself.

A "text" in this Atlas can be one of two things: either 1) any Mainspace page with no "/" parent of its own, whether or not it collects real subpages under itself; or 2) an Index item still in the OCR "Proofreading" pipeline that hasn't yet been transcluded into a Mainspace page. The artificial असम्बद्धवर्गीकृतम् category contains the majority of completely unproofread OCR, and its Mainspace pages frequently suffer from poor organization, such that subpages don't "roll up" properly and instead get counted as full "texts." For these reasons, this category may optionally be excluded from size and text/page-count totals.

Transclusion Issues

When OCR content has not been transcluded to Main, the Atlas links to the Index page and flags it with a "Proofing" label. Even so, it remains possible that the editor copy-pasted to Main.

Where any range has been transcluded, the Atlas links to the Mainspace page, and a small "pdf" icon also links to the Index page for source images.

Cross-tree Connections

If a Category (wrongly!) has multiple parents, it will show up in all those places in the Atlas's tree display with an on-hover note ("also filed under...") so that all such occurrences clearly link to each other.

On the other hand, multiple Categories are entirely appropriate for text items, both Main and Index. In this case, a small magnifying-glass icon next to such an item lets a reader jump to a filtered view of the tree showing every other place that same item appears.

An item filed in multiple places risks being double-counted in the Atlas's size and text/page-count stats, so special care is taken that each item be credited to a given category's rollup exactly once. For example, काश्यपसंहिता is filed under both धर्मशास्त्रम् (> काश्यपरचनाः) and वेदाः (> उपवेदाः > आयुर्वेदः > काश्यपरचनाः), so its size and text count contribute to both धर्मशास्त्रम् and वेदाः, but the top-level "All" aggregation still only counts it once.

Calculating Size

The size Wikisource itself reports for a page reflects stored "wikitext," including markup, templates, and Category tags. Importantly, transcluded content is NOT counted. In order to obtain a more meaningful measure of how much "real" text the collection contains, this Atlas parses the wikitext, strips out overhead, populates transcluded content, and transliterates from Devanāgarī to IAST. IAST text takes up less space (approximately 50% of the Devanāgarī byte count) and is more common for cross-collection comparison.

Data Quality: Structural Errors and Provenance Hints

Opportunities for Correction

In addition to the artificial असम्बद्धवर्गीकृतम् category shown on this Atlas website, the codebase also includes an audit pipeline for detecting and concisely indicating problems that should be fixed on sa.wikisource.org:

Provenance Insights

The same audit pipeline also makes note of source links:

Data Quantity: Historical Trends and Deltas

Timeline of data source types. See "Snapshots" above in the "Obtaining and Representing" section for detail.

materialized (synthetic) legacy-format live current-format live

Loading…

Deltas

Loading…