CACCHT

Creating Annotated Corpora of Classical Hebrew Texts

Ancient Semitic
texts, published
as data.

CACCHT prepares linguistically annotated editions of ancient Semitic texts and releases them openly, so that a question about grammar, syntax or textual variation can be answered by running code over a whole corpus.

Six corporaFive scriptsOne annotation system

The CACCHT roundel: aleph, alpha, olaph and the Ugaritic alpa in four quadrants
Hebrew · Greek · Syriac · Ugaritic

Corpora

Six datasets, free to use

Every dataset is a Text-Fabric dataset: text and annotations stored as a graph of nodes and features, read straight into Python. Some carry word-level annotation only; others add phrase and clause structure.

אActive

The Dead Sea Scrolls

Biblical and non-biblical scrolls, transcribed and annotated word by word.

Script
Hebrew square
Source
Transcriptions provided by Martin Abegg
ܐIn progress

The ETCBC Syriac Corpus

Peshitta books, the Syrohexapla Psalms, and post-biblical prose including Ephrem and the Book of the Laws of the Countries.

Script
Syriac, with the consonantal text in Estrangela
Source
The ETCBC database of Syriac literature
In progress

The Samaritan Pentateuch

The Pentateuch with word-level annotations in BHSA style, plus phrase-atom and clause-atom boundaries.

Script
Samaritan
Source
Samaritanus project, Halle-Wittenberg (ed. Stefan Schorch): MS Dublin Chester Beatty 751 and MS Garizim 1
אAlpha

The ETCBC Targum Corpus

Pseudo-Jonathan, the Fragment Targums, the Cairo Genizah fragments and Neofiti — with scribal variants kept as first-order data rather than flattened away.

Script
Aramaic square
Pipeline
SQL to XML to Text-Fabric, included in the repository
𐎀In progress

The Copenhagen Ugaritic Corpus

278 tablets of Die keilalphabetischen Texte aus Ugarit, annotated down to tablet, column, line and side, with certainty marked on individual signs and alternative readings recorded.

Script
Ugaritic alphabetic cuneiform
With
Tania Notarius (University of the Free State), Alex Sosnovshchenko and Kseniia Yerofeieva
ΑIn development

The Septuagint

A Text-Fabric edition of the Greek text, under construction at CenterBLC. The repository holds the conversion pipeline and the dataset as it grows.

Script
Greek
Status
No release yet — follow the repository

Method

One annotation system, adapted per language

We follow the BHSA

The Biblia Hebraica Stuttgartensia Amstelodamensis — the annotated Masoretic Text maintained by the ETCBC — sets the conventions. CACCHT follows them and adapts them where a language or a text demands it, so that features keep the same meaning across corpora and a query written for one dataset largely survives the move to another.

g_conslexglossspvsvtpsgnnuvostlstrailer

Word-level features in the Syriac corpus. Each dataset documents its own set.

Read it with Text-Fabric

Nothing needs downloading by hand. Point Text-Fabric at a dataset and it fetches the data, then gives you nodes, features and a search language over them — in a Jupyter notebook, or in the browser it ships with.

The same datasets can be read with Context-Fabric, an API-compatible engine that loads them with a far smaller memory footprint and exposes them to AI agents over MCP.

from tf.app import use
A = use('dt-ucph/sp')
tf dt-ucph/sp

The Samaritan Pentateuch, in a notebook and in the Text-Fabric browser.

Open by default

Every dataset lives in a public repository, and each release is archived with a DOI at Zenodo. All of them are free to use for research and education. Licenses differ per corpus, so check the card above and cite the release you actually used.

Tool MT–SP Parallels Read the Masoretic Text and the Samaritan Pentateuch side by side, verse by verse, built on the BHSA and the CACCHT SP data.

Publications

How the annotations are made, written up

  1. 2026

    Identifying Phrase Boundaries in the Samaritan Pentateuch with Machine Learning

    Cantanhêde, S. d. O., Naaijer, M., Højgaard, C. C., & Glanz, O. — Religions 17(2), 192.

    10.3390/rel17020192
  2. 2024

    Text-Fabric Dataset of the Samaritan Pentateuch

    Naaijer, M., Højgaard, C. C., Schorch, S., & Ehrensvärd, M. — Research Data Journal for the Humanities and Social Sciences 9(1), 1–13.

    10.1163/24523666-bja10051
  3. 2023

    A Transformer-based parser for Syriac morphology

    Naaijer, M., Sikkel, C., Coeckelbergs, M., Attema, J., & Van Peursen, W. Th. — Proceedings of the Ancient Language Processing Workshop, Varna, 23–29.

    ACL Anthology

Team

A collaboration of six

CACCHT works together with specialists in each field to build the datasets.

  • Martijn NaaijerUniversity of Zurich
  • Willem van PeursenVrije Universiteit Amsterdam
  • Oliver GlanzAndrews University
  • Christian Canu HøjgaardFjellhaug International University College
  • Martin EhrensvärdUniversity of Copenhagen
  • Robert RezetkoUniversity of Copenhagen

With Stefan Schorch and the Samaritanus project, Martin Abegg, Tania Notarius, Alex Sosnovshchenko, Kseniia Yerofeieva, and colleagues at the ETCBC and DANS.

Use the data. Cite the release.

The datasets are free for research and education. Each repository states how to cite it.

Open the GitHub organization