CLLG Freed Corpus
Corpus Liberatum Linguae Graecae
Overview
The CLLG Freed Corpus is a free XML-TEI corpus of Ancient Greek texts derived from the Thesaurus Linguae Graecae (TLG) E CD-ROM. It was first converted to TEI/XML by Diogenes (Peter Heslin); the CLLG then did the extra work that makes it usable with Dapytains: citation structures, bibliographical metadata and a cleaner XML import.
The corpus is developed within the CLLG project, funded by the Programme Inria Quadrant at Inria and by the Agence Nationale de la Recherche (France 2030, «ANR-24-RRII-0002»). The Diogenes TEI/XML export it starts from was produced by Peter Heslin (Diogenes).
The corpus is tested with HookTest.
How this corpus was built
Every text states its own history in its teiHeader (titleStmt/respStmt and
publicationStmt/availability); this is what they say, for example in tlg2062.tlg017.cllg-grc1.xml.
- First conversion: Diogenes (Peter Heslin). Diogenes development and TEI/XML export. This XML text (filename tlg2062017.xml) was originally generated by Diogenes (version 4.6.2) from the TLG corpus of classical texts (author number 2062, work number 017), which was once distributed on CD-ROM in a very different file format. Diogenes and Peter Heslin's TEI/XML export are the foundation of the whole corpus: they turned the TLG CD-ROM format into XML.
- Refinement: Corpus Liberatum Linguae Graecae (Thibault Clérice, Nicolas Angleraud, Antonia Karamolegkou, Benoît Sagot). Refined conversion to XML-TEI. This file has then been released by the Corpus Liberatum Linguae Graecae (dir. Thibault Clérice) which added citeStructure information, bibliographical metadata and reworked, where applicable, the citation structure to follow edition and physical citation systems.
On top of the Diogenes export, our work is what makes the texts usable with dapytains:
declared
citeStructuretrees, a cleaner XML import (normalised structure, the Leiden sigla encoded as TEI) and bibliographical metadata.
Features
- TEI P5–compliant XML encoding
- Open and redistributable texts
- Normalized structure and metadata
- Canonical citation (
citestructure) support - Usable with Dapytains (https://github.com/distributed-text-services/MyDapytains)
Data and Processing
- Extraction from TLG E CD-ROM and first conversion to TEI XML by Diogenes (Peter Heslin)
- Cleaner XML import by the CLLG (normalized structure, Leiden sigla encoded as TEI)
- Metadata normalization and bibliographical metadata
- Integration and rework of citation structures (
citeStructure, for Dapytains) - Validation and quality control
Texts are distributed under French law, which holds that the texts of critical editions are not subject to copyright protection, as they do not constitute an expression of the editor’s personality (CA Paris, Pôle 5, 2ème Ch., 9 June 2017, no. 16/00005).
Site: overview, citeStructure and readable views
scripts/build_site.py builds a static site (public/, published by the pages CI job) with:
index.html: every author, work and text with its citation structure (all trees, units and counts), tokens and Leiden figures. The page is driven by an embedded JSON and renders only the authors being looked at; it filters by text, structure, status, Leiden sigla and by kind (dubia, spuria, fragmenta, pseudo-authors);structures.html: statistics on the distinct citation structures;about.html: how the corpus was built (read from theteiHeader) and this README;cache.json: per text, a sha of its TEI +xslt/tei-to-html.xsland the files built; the next build downloads it from the live site and only renders texts whose sha changed;t/<file stem>/index.html: one text, with its declaredciteStructure, the reference tree and the Leiden counts;t/<file stem>/read-<k>.html: a readable view, cut at top-level citation units.
The views are produced by xslt/tei-to-html.xsl (XSLT 3.0, Saxon) from dapytains passages. The stylesheet renders any excerpt, and shows the Leiden sigla that scripts/leiden_convert.py encoded as TEI: [ ] for lost text (adjacent lacunae share one pair), dots for lost or illegible characters, a dotted underline for uncertain letters, ⟨ ⟩ for omitted text, { } for surplus, ⟦ ⟧ for deletions. The reader has a switch to hide them (and the page marks). xslt/site.css documents every class.
pip install dapytains lxml # saxonche comes with dapytains
python scripts/build_site.py --out public # everything (warnings go to build.log)
python scripts/build_site.py --only 'tlg0007.*' --out /tmp/site # a few texts
python scripts/build_site.py --views none # structure pages only
python scripts/build_site.py --reuse-from https://<pages-url> # reuse unchanged texts from a previous build (--no-reuse to disable)
python scripts/build_site.py --pack shell # reading pages as shell + gzip payload only (smallest)
Repository Structure
data/
workgroup/
metadata.xml
work/
edition.xml
License
- Texts: CC0 1.0 (https://creativecommons.org/publicdomain/zero/1.0/). Texts of critical editions are not subject to copyright protection under French law, as they do not constitute an expression of the editor’s personality (CA Paris, Pôle 5, 2ème Ch., 9 June 2017, no. 16/00005).
- Citation structure and metadata: CC-BY 4.0 (https://creativecommons.org/licenses/by/4.0/), i.e. the
citeStructureinformation, the bibliographical metadata and the reworked citation structure added by the CLLG.
Each file states this in its teiHeader (publicationStmt/availability), with each paragraph linked (@resp) to the group that produced it (#diogenes-transformation, #cllg-transformation), and records the funders in titleStmt/funder. scripts/header_licence.py applied that header block to every text.
Citation
If you use this corpus, please cite:
Corpus Liberatum Linguae Graecae (CLLG), Freed Corpus.
@misc{clerice2026cllg,
author = {Clérice, Thibault and Angleraud, Nicolas and Karamolegkou, Antonia and Sagot, Benoît},
title = {Corpus Liberatum Linguae Graecae},
year = {2026},
howpublished = {\url{https://gitlab.inria.fr/almanach/cllg/freed-corpus}},
note = {Inria ALMAnaCH}
}
Contact & Contributions
Issues, corrections, and contributions are welcome via the project repository.