Data and research transparency

How this data is prepared and published

This page distinguishes source text, technical normalization, automated analysis, machine-generated English translations, and human review so each result can be interpreted and cited correctly.

Release
2026-08-15
Source captured
2026-08-15
Public inscriptions
143
Published lines
1,627
Data fingerprint
b9b9284386e8ad04…
Source and permission

The TITUS corpus and its permitted use

Direct permission is recorded for storage, processing, and non-commercial public presentation of the TITUS Old Persian textual corpus in the Parsiandej reference. Parsiandej is not an official TITUS product and is responsible for its processing, labels, and presentation.

Source project names and links remain attached to related results. Public bulk downloads and programmatic third-party access are not provided; the complete machine-readable package remains private for academic review.

View the primary TITUS corpus
Data layers

Source text is not the same as added data

  1. Source text: cuneiform, transliteration, transcription, inscription identifiers, and line locations are kept without semantic rewriting.
  2. Technical normalization: spacing, dividers, and searchable forms are stored separately and never replace the source text.
  3. Tokenization: words, logograms, numerals, dividers, and editorial markers are distinguished.
  4. Dictionary links: automatic, inferred, and reviewed matches retain their own method and confidence.
  5. English translations: currently published English line translations are explicitly labelled Machine-generated and unreviewed.
  6. Sign table: only assigned Unicode Old Persian characters are treated as valid script characters.
Translation and confidence

A literal gloss sequence is not a literary translation

English translations generated from the active authorized corpus snapshot are stored as version-bound records with language en, provenance “machine-generated”, and review status “generated”. The visible label “Machine-generated and unreviewed” is never removed.

These translations help discovery and reading but do not claim a human-reviewed interpretation of syntax, grammar, or literary style.

Editorial notation

Damage, reconstruction, and special marks

Damage, editorial reconstruction, and text completeness are not conveyed by colour alone: each published line includes a textual status. Numerals, logograms, and dividers retain their sign type.