How this data is prepared and published
This page distinguishes source text, technical normalization, automated analysis, machine-generated English translations, and human review so each result can be interpreted and cited correctly.
- Release
- 2026-08-15
- Source captured
- 2026-08-15
- Public inscriptions
- 143
- Published lines
- 1,627
- Data fingerprint
- b9b9284386e8ad04…
The TITUS corpus and its permitted use
Direct permission is recorded for storage, processing, and non-commercial public presentation of the TITUS Old Persian textual corpus in the Parsiandej reference. Parsiandej is not an official TITUS product and is responsible for its processing, labels, and presentation.
Source project names and links remain attached to related results. Public bulk downloads and programmatic third-party access are not provided; the complete machine-readable package remains private for academic review.
View the primary TITUS corpusSource text is not the same as added data
- Source text: cuneiform, transliteration, transcription, inscription identifiers, and line locations are kept without semantic rewriting.
- Technical normalization: spacing, dividers, and searchable forms are stored separately and never replace the source text.
- Tokenization: words, logograms, numerals, dividers, and editorial markers are distinguished.
- Dictionary links: automatic, inferred, and reviewed matches retain their own method and confidence.
- English translations: currently published English line translations are explicitly labelled Machine-generated and unreviewed.
- Sign table: only assigned Unicode Old Persian characters are treated as valid script characters.
A literal gloss sequence is not a literary translation
English translations generated from the active authorized corpus snapshot are stored as version-bound records with language en, provenance “machine-generated”, and review status “generated”. The visible label “Machine-generated and unreviewed” is never removed.
These translations help discovery and reading but do not claim a human-reviewed interpretation of syntax, grammar, or literary style.
Damage, reconstruction, and special marks
Damage, editorial reconstruction, and text completeness are not conveyed by colour alone: each published line includes a textual status. Numerals, logograms, and dividers retain their sign type.
Exact matches before phonetic approximation
The converter searches complete forms in published data first. A close suggestion is not the same as a documented attestation, and a phonetic approximation is not historical evidence or a translation.