Data and research transparency

How this data is prepared and published

This page distinguishes source text, technical normalization, automated analysis, machine-generated English translations, and human review so each result can be interpreted and cited correctly.

Release
2026-08-15
Source captured
2026-08-15
Public inscriptions
143
Published lines
1,627
Data fingerprint
b9b9284386e8ad04…
Source and permission

The TITUS corpus and its permitted use

Direct permission is recorded for storage, processing, and non-commercial public presentation of the TITUS Old Persian textual corpus in the Parsiandej reference. Parsiandej is not an official TITUS product and is responsible for its processing, labels, and presentation.

Source project names and links remain attached to related results. Public bulk downloads and programmatic third-party access are not provided; the complete machine-readable package remains private for academic review.

View the primary TITUS corpus
Data layers

Source text is not the same as added data

  1. Source text: cuneiform, transliteration, transcription, inscription identifiers, and line locations are kept without semantic rewriting.
  2. Technical normalization: spacing, dividers, and searchable forms are stored separately and never replace the source text.
  3. Tokenization: words, logograms, numerals, dividers, and editorial markers are distinguished.
  4. Dictionary links: automatic, inferred, and reviewed matches retain their own method and confidence.
  5. Generated English: line glosses, catalog descriptions, evidence notes, editorial notes, citation notes, and image descriptions are explicitly labelled Machine-generated and unreviewed.
  6. Sign table: only assigned Unicode Old Persian characters are treated as valid script characters.
Translation and confidence

Generated English is not a reviewed critical translation

English content generated from the active authorized corpus snapshot is stored in an offline, version-bound package with provenance ā€œmachine-generatedā€, review status ā€œgeneratedā€, a source-text fingerprint, and the generator batch. The visible label ā€œMachine-generated and unreviewedā€ remains attached until a human-reviewed replacement is explicitly approved.

Generated line glosses and supporting prose help discovery and reading but do not claim a human-reviewed interpretation of syntax, grammar, history, or literary style. Source cuneiform, identifiers, scholarly references, and official Unicode names are not translated or relabelled as generated content.

Editorial notation

Damage, reconstruction, and special marks

Damage, editorial reconstruction, and text completeness are not conveyed by colour alone: each published line includes a textual status. Numerals, logograms, and dividers retain their sign type.