CLLG Freed Corpus

Corpus Liberatum Linguae Graecae

Overview

The CLLG Freed Corpus is a free XML-TEI corpus of Ancient Greek texts derived from the Thesaurus Linguae Graecae (TLG) E CD-ROM. It was first converted to TEI/XML by Diogenes (Peter Heslin); the CLLG then did the extra work that makes it usable with Dapytains: citation structures, bibliographical metadata and a cleaner XML import.

The corpus is developed within the CLLG project, funded by the Programme Inria Quadrant at Inria and by the Agence Nationale de la Recherche (France 2030, «ANR-24-RRII-0002»). The Diogenes TEI/XML export it starts from was produced by Peter Heslin (Diogenes).

The corpus is tested with HookTest.


How this corpus was built

Every text states its own history in its teiHeader (titleStmt/respStmt and publicationStmt/availability); this is what they say, for example in tlg2062.tlg017.cllg-grc1.xml.

  1. First conversion: Diogenes (Peter Heslin). Diogenes development and TEI/XML export. This XML text (filename tlg2062017.xml) was originally generated by Diogenes (version 4.6.2) from the TLG corpus of classical texts (author number 2062, work number 017), which was once distributed on CD-ROM in a very different file format. Diogenes and Peter Heslin's TEI/XML export are the foundation of the whole corpus: they turned the TLG CD-ROM format into XML.
  2. Refinement: Corpus Liberatum Linguae Graecae (Thibault Clérice, Nicolas Angleraud, Antonia Karamolegkou, Benoît Sagot). Refined conversion to XML-TEI. This file has then been released by the Corpus Liberatum Linguae Graecae (dir. Thibault Clérice) which added citeStructure information, bibliographical metadata and reworked, where applicable, the citation structure to follow edition and physical citation systems. On top of the Diogenes export, our work is what makes the texts usable with dapytains: declared citeStructure trees, a cleaner XML import (normalised structure, the Leiden sigla encoded as TEI) and bibliographical metadata.

Features


Data and Processing

  1. Extraction from TLG E CD-ROM and first conversion to TEI XML by Diogenes (Peter Heslin)
  2. Cleaner XML import by the CLLG (normalized structure, Leiden sigla encoded as TEI)
  3. Metadata normalization and bibliographical metadata
  4. Integration and rework of citation structures (citeStructure, for Dapytains)
  5. Validation and quality control

Texts are distributed under French law, which holds that the texts of critical editions are not subject to copyright protection, as they do not constitute an expression of the editor’s personality (CA Paris, Pôle 5, 2ème Ch., 9 June 2017, no. 16/00005).


Site: overview, citeStructure and readable views

scripts/build_site.py builds a static site (public/, published by the pages CI job) with:

The views are produced by xslt/tei-to-html.xsl (XSLT 3.0, Saxon) from dapytains passages. The stylesheet renders any excerpt, and shows the Leiden sigla that scripts/leiden_convert.py encoded as TEI: [ ] for lost text (adjacent lacunae share one pair), dots for lost or illegible characters, a dotted underline for uncertain letters, ⟨ ⟩ for omitted text, { } for surplus, ⟦ ⟧ for deletions. The reader has a switch to hide them (and the page marks). xslt/site.css documents every class.

pip install dapytains lxml          # saxonche comes with dapytains
python scripts/build_site.py --out public              # everything (warnings go to build.log)
python scripts/build_site.py --only 'tlg0007.*' --out /tmp/site   # a few texts
python scripts/build_site.py --views none              # structure pages only
python scripts/build_site.py --reuse-from https://<pages-url>   # reuse unchanged texts from a previous build (--no-reuse to disable)
python scripts/build_site.py --pack shell              # reading pages as shell + gzip payload only (smallest)

Repository Structure

data/
  workgroup/
    metadata.xml
    work/
        edition.xml

License

Each file states this in its teiHeader (publicationStmt/availability), with each paragraph linked (@resp) to the group that produced it (#diogenes-transformation, #cllg-transformation), and records the funders in titleStmt/funder. scripts/header_licence.py applied that header block to every text.


Citation

If you use this corpus, please cite:

Corpus Liberatum Linguae Graecae (CLLG), Freed Corpus.
@misc{clerice2026cllg,
  author = {Clérice, Thibault and Angleraud, Nicolas and Karamolegkou, Antonia and Sagot, Benoît},
  title = {Corpus Liberatum Linguae Graecae},
  year = {2026},
  howpublished = {\url{https://gitlab.inria.fr/almanach/cllg/freed-corpus}},
  note = {Inria ALMAnaCH}
}

Contact & Contributions

Issues, corrections, and contributions are welcome via the project repository.