AI & Data-Mining Licensing
What's open, what's reserved, and how to license training use
We want these texts read, cited, and built upon — including by AI. This page states plainly how that works, so there’s no ambiguity about what is freely permitted and what requires a separate license.
Effective date: July 4, 2026 · Operated by the Embassy of the Free Mind, Amsterdam, the Netherlands. · Companion to our Terms & Licensing.
Freely permitted — no license needed
- Original texts & page images are in the public domain. Use them freely.
- Our AI-generated translations and editorial content are licensed CC BY-SA 4.0 — free for individuals, researchers, and organizations, with attribution and ShareAlike.
- Search indexing. Search engines and AI search crawlers (Googlebot, Bingbot, Claude-SearchBot, OAI-SearchBot, …) are welcome to index our pages so readers can discover and cite them.
- AI assistants reading on a user’s behalf. When a person asks an assistant to read or quote a specific Source Library page, that’s welcome — please show the quote with its page number and the page’s sourcelibrary.org link.
- The public API & MCP server, within posted rate limits, for research and individual use. See /developers.
Reserved — a separate license is required
Using our content to train, fine-tune, or build AI models, and bulk text-and-data-mining (TDM) for those purposes, is expressly reserved and requires a separate license from us. That reservation rests on several independent grounds:
- TDM reservation. We reserve text-and-data-mining rights within the meaning of Article 4 of EU Directive 2019/790, expressed in machine-readable form via
/.well-known/tdmrep.json(TDM Reservation Protocol), theTDM-Reservation: 1HTTP header on every response, and the training-crawler rules in/robots.txt. General-purpose AI providers serving the EU market are required to identify and respect this reservation. - Database right. The Source Library corpus — the curated, verified, and structured collection of texts, transcriptions, translations, and metadata — is a database within the meaning of EU Directive 96/9/EC, reflecting substantial investment by the Embassy of the Free Mind. Extraction or re-utilization of a substantial part of it (including the OCR and metadata layers) requires our authorization, independent of the copyright status of any individual item.
- First publication of unpublished works. Where we are the first to lawfully publish a previously unpublished public-domain work — as with a number of the manuscripts we digitize — we hold the exclusive economic rights granted by Article 4 of EU Directive 2006/116/EC for 25 years from publication.
- Copyright. Human-authored editorial content (collection essays, blog posts, curatorial descriptions) is protected by copyright. Our AI-generated translations and annotations are offered under CC BY-SA 4.0 — and that license’s ShareAlike term would require a model built on this material to be released under CC BY-SA, which proprietary models do not do. The free license therefore does not authorize proprietary AI training.
Bulk or training access is available — through the API or a dataset license, under the standard terms below or a partnership. We’d genuinely like these texts in the models that shape how people learn; we just ask for a conversation and attribution. Please don’t scrape the full corpus around the controls.
Standard training license — rate card
A training license is available to anyone, on standard terms:
- Per book: at least $250 per book (per edition, as listed in our catalog), non-exclusive.
- Full corpus: $200,000 per year, non-exclusive, with quarterly refresh, delivered as structured data via the API or dataset export — no crawling required.
The license covers training, fine-tuning, and embedding-index use of our translations, transcriptions, and editorial content, with attribution. What you’re licensing is unique: billions of words of aligned parallel text — original and English translation, page by page — across more than a hundred languages, including scarce low-resource ones (Latin, classical Chinese, Sanskrit, Tibetan, Syriac), none of it in Common Crawl. Corpus partners additionally receive priority input on what gets translated next through our sponsorship program.
Unlicensed training use is unauthorized. Where we identify training use of our content without a license, we will invoice at the standard rates above and pursue payment under the reserved rights described in the previous section.
Licensing & partnerships
For an AI-training or bulk-dataset license, or a research partnership, contact derek@sourcelibrary.org with the subject “AI Licensing Inquiry.” We respond within a day.
This page states our position and our express reservation of rights; it is not legal advice. Nothing here waives any right not expressly granted.