Open Science in the Age of Artificial Intelligence

European Experiences in the Digital Humanities



Christof Schöch
Trier Center for Digital Humanities
Trier University, Germany


International Digital Humanities Symposium:
Humanities Data Infrastructure in the Age of Generative AI

Tokyo, Japan

18 Jul 2026

Introduction







Slides: tokyo.dh-trier.de

Abstract

Over the last two decades, the Digital Humanities have flourished and spread across Europe. In the process, Digital Humanities researchers have embraced Open Science principles as a driving force and profoundly transformed their scholarly practices. National and European research infrastructures are being developed to support accessibility, reproducibility and sustainability for research data, methods and training.

Now, Artificial Intelligence is reshaping the landscape, offering both profound challenges and transformative opportunities. Closed datasets and AI models, inequalities in access to resources, and linguistic biases threaten to undermine Open Science’s core principles. Copyright issues, privacy risks, and environmental costs demand urgent attention. At the same time, Artificial Intelligence enhances research in the Digital Humanities in multiple ways, whether by automating rich annotations on large datasets, enabling insights into multilingual resources or democratizing reseachers’ ability to design algorithmic pipelines.

This keynote will explore how Europe’s Digital Humanities are navigating this complex intersection, leveraging the opportunities of Artificial Intelligence to advance research while safeguarding its ethical and inclusive foundations. The future of open knowledge production in the Humanities depends on our ability to bridge these worlds by fostering collaborative infrastructures, developing creative solutions, and maintaining scholarly best practices.

Overview

  1. Introduction
  2. Open Science
  3. Artificial Intelligence
  4. Research Infrastructures
  5. Conclusion

Open Science
(and the Digital Humanities)

Aspects of Open Science

Why practice Open Science?

Building on the essential principles of academic freedom, research integrity and scientific excellence, open science sets a new paradigm that integrates into the scientific enterprise practices for reproducibility, transparency, sharing and collaboration resulting from the increased opening of scientific contents, tools and processes.

– UNESCO, Recommendations on Open Science, 2021.

The Open Science Staircase

Open Humanities, or how does Open Science apply to the (Digital) Humanities?

  • Open Access publications (85% of journals)
  • Open Data: TextGrid Digital Library, ELTeC, DraCor, PoeTree
  • Openness as scholar-led standards (XML-TEI)
  • Open software: stylo, Gephi, Voyant, …
  • Openness as a community value

Case study: “Mining and Modeling Text”

  • Open Access, e.g. Hinzmann et al. (2026)
  • Open Dataset: Corpus of French Novels (in XML-TEI)
  • Open Standards: Linked Open Data (RDF, Dublin Core, etc.)
  • Open Infrastructure: Wikibase, Wikidata, OpenRefine
  • Open Licences: Creative Commons
  • Open Educational Resources: Tutorial

The Advent of Artificial Intelligence in Digital Humanities

A timeline of recent developments in AI

  • 2003: Topic Modeling
  • ~2006: Deep Learning / Neural Networks become competitite
  • 2012: Static word embeddings (word2vec)
  • 2017: Contextual word embeddings (BERT)
  • 2018: Encoder-decoder models (GPTs)
  • ~2023: Multimodal models become popular
  • 2024: Agentic models (MCP)

Applications of AI in Digital Humanities

Case study: Historical Wine Labels

Case study: Historical Wine Labels

Multimodal models to generate image descriptions

  • Small, local, open-weights model (Gemma4-12B)
  • Low hardware requirements (12GB GPU, runtime 60s)
  • Description sufficient for search / retrieval

Controlled vocabulary + structured results

  • Multilingual taxonomy with authority data
  • Output in pre-defined JSON structure
{
  "identifier": "SMW-D-E124",
  "designation": "1918er Maringer Rosenberg",
  "objects": [
    {"wid":"Q941", "eng":"man", "deu":"Mann"},
    {"wid":"Q939", "eng":"woman", "deu":"Frau"},
    {"wid":"Q938", "eng":"wine glass", "deu":"Weinglas"},
    {"wid":"Q881", "eng":"grape", "deu":"Weintraube"},
    {"wid":"Q879", "eng":"wine leaf", "deu":"Weinblätter"},
    {"wid":"Q875", "eng":"flower", "deu":"Blume"},
    {"wid":"Q966", "eng":"abstract form", "deu":"Abstrakte Form"},
    [...]
  ]
}

Result: Map and database of historical wine labels

Challenges and Lessons Learned (especially for Open Science)

  • BERT-models and generative LLMs are a game changer for research in DH
  • They should only be used for tasks that can be evaluated
  • However: LLMs challenge Open Science
    • High requirements for access (financial / hardware)
    • Models are not truly open: limited reproducibility
    • Models are driven by commercial interests, not scholarly needs
  • Additional concerns
    • Environmental impacts
    • Respect of copyright law
    • Protection of personal data
    • Bias towards English language and culture


Consequence: Use small, local, open-weight models whenever possible

Research infrastructures for Digital Humanities (in the context of AI and Open Science)

National Research Infrastructures: Germany

European-Level Resarch Infrastructures for SSH

  • Research Infrastructures
    • DARIAH-EU (Digital Research Infrastructure for the Arts and Humanities)
    • CLARIN-EU (Common Language Resources and Technology Infrastructure)
    • OPERAS (Open Scholarly Communication in the European Research Are for HSS)
    • OpenAIRE (Open Scholarly Communication Infrastructure)
  • Specific Services
    • ORE (Open Research Europe): mega-journal platform)
    • Zenodo: long-term data deposit
    • EOSC (European Open Science Cloud): collaboration platform

Case Study: European Open Science Cloud (EOSC)

How RI support Open Science and Artificial Intelligence

  • Promote FAIR principles (NFDI, DARIAH)
  • Enable and support Open Access Publishing (OER, OPERAS)
  • Offer digital collaboration tools: filesharing, notebooks, CPU/GPU resources (EOSC)
  • Offer dedicated high-performance computing resources (NHR)
  • Enrich, improve, and make available data resources (DARIAH, CLARIN)
  • Offer training materials and opportunities (DARIAH, NFDIs)
  • Advocate for the needs of the SSH
  • Serve as community hubs and DH incubators
  • Perform basic research for infrastructure development

Solution with ‘Derived Text Formats’

Derived Text Formats (DTFs) are the result of a strategic transformation of textual materials that are protected by copyright in their original form, such that the resulting data is useful for computational analyses and can be openly shared following best practices of Open Science without infringing copyright law.

  • Four steps to create a DTF
    • Legal access to originals
    • Information enrichment
    • Information reduction
    • Open sharing possible

Example from Sherlock Holmes

original

Mr. Sherlock Holmes, who was usually very late in the mornings, save upon those not infrequent occasions when he was up all night, was seated at the breakfast table. […]

(Arthur Conan Doyle, The Hound of the Baskervilles, 1902)

annotated

Mr._Mr._NNP
Sherlock_Sherlock_NNP
Holmes_Holmes_NNP
,_,_PUNC
who_who_WP
was_be_VBD
usually_usually_RB
very_very_RB
late_late_JJ
in_in_IN
the_the_DT
mornings_morning_NNS ,_,_PUNC

randomized

save_save_IN
,_,_PUNC
who_who_WP
infrequent_infrequent_JJ
table_table_NN
very_very_RB
occasions_occasion_NNS
he_he_PRP
the_the_DT
up_up_RP
was_be_VBD
usually_usually_RB
Mr._Mr._NNP

Current research on DTFs

Conclusion

Open Science, Large Language Models, and Research Infrastructures

  • Open Science remains a cornerstone for excellent research
  • LLMs are a game-changer for research in the Digital Humanities
  • Very large, web-based LLMs are largely incompatible with Open Science
  • Research infrastructures are already doing essential work for Open Science
  • Now they also need to train open models and provide computing resources

  • Research infrastructures can bring Open Science and Artificial Intelligence together

Thank you!

References

Arnold, Taylor B., and Lauren Tilton. 2024. “Explainable Search and Discovery of Visual Cultural Heritage Collections with Multimodal Large Language Models.” CHR2024. https://ceur-ws.org/Vol-3834/paper1.pdf.
Brunner, Annelen, Ngoc Duyen Tanja Tu, Lukas Weimer, and Fotis Jannidis. 2020. “To BERT or Not to BERT - Comparing Contextual Embeddings in a Deep Learning Architecture for the Automatic Recognition of Four Types of Speech, Thought and Writing Representation.” In Proceedings of the 5th Swiss Text Analytics Conference (SwissText) & 16th Conference on Natural Language Processing (KONVENS), Zurich, Switzerland, June 23-25, 2020. https://ceur-ws.org/Vol-2624/paper5.pdf.
Da, Nan Z. 2019. “The Computational Case Against Computational Literary Studies.” Critical Inquiry 45 (3): 601–39. https://doi.org/10.1086/702594.
Du, Keli. 2023. “Understanding the Impact of Three Derived Text Formats on Authorship Classification with Delta.” In Jahreskonferenz der Digital Humanities im deutschsprachigen Raum. DHd-Verband. https://doi.org/10.5281/zenodo.7715298.
Du, Keli, Sarah Ackerschewski, Uygar Navruz, Nazan Sinir, Julian Valline, and Christof Schöch. 2025. “Reconstructing Shuffled Text: Bad Results for NLP, but Good News for Using in-Copyright Texts.” Journal of Computational Literary Studies 4. https://doi.org/10.48694/jcls.4163.
Du, Keli, and Christof Schöch. 2024. “Shifting Sentiments. What Happens to BERT-Based Sentiment Classification When Derived Text Formats Are Used for Fine-Tuning.” In Digital Humanities 2024: Book of Abstracts. ADHO.
Fraser, Nicholas, Fakhri Momeni, Philipp Mayr, and Isabella Peters. 2020. “The Relationship Between bioRxiv Preprints, Citations and Altmetrics.” Quantitative Science Studies 1 (2): 618–38. https://doi.org/10.1162/qss_a_00043.
Fu, Darwin Y., and Jacob J. Hughey. 2019. “Releasing a Preprint Is Associated with More Attention and Citations for the Peer-Reviewed Article.” eLife 8: e52646. https://doi.org/10.7554/eLife.52646.
Hicke, Rebecca M. M., and David Mimno. 2023. “T5 Meets Tybalt: Author Attribution in Early Modern English Drama Using Large Language Models.” https://arxiv.org/abs/2310.18454.
Hinzmann, Maria, Matthias Bremm, Tinghui Duan, Anne Klee, Johanna Konstanciak, Julia Röttgermann, Moritz Steffes, Christof Schöch, and Joëlle Weis. 2026. “Patterns in Modeling and Querying a Knowledge Graph for Literary History.” In Patterns in Language and Communication, edited by Sabine Arndt-Lappe, Sören Stumpf, Milena Belosevic, Peter Maurer, Claudine Moulin, and Achim Rettinger, 135–76. De Gruyter. https://doi.org/10.1515/9783111199740-006.
Huang, Chun-Kai, Cameron Neylon, Lucy Montgomery, Richard Hosking, James P. Diprose, Rebecca N. Handcock, and Katie Wilson. 2024. “Open Access Research Outputs Receive More Diverse Citations.” Scientometrics 129 (2): 825–45. https://doi.org/10.1007/s11192-023-04894-0.
Keith, Allison, Antonio Rojas Castro, Kerstin Jung, Hanno Ehrlicher, and Sebastian Padó. 2026. “A Computational Analysis of Character Archetypes in the Works of Calderón de La Barca.” Journal of Computational Literary Studies 5. https://doi.org/10.48694/jcls.4167.
Kocula, Martin. 2021. “Volltext vs. abgeleitetes Textformat: Systematische Evaluation der Performanz von Topic Modeling bei unterschiedlichen Textformaten mit Python.” Trier University. https://doi.org/10.5281/zenodo.5552486.
Kugler, Kai, Simon Münker, Johannes Höhmann, and Achim Rettinger. 2024. InvBERT: Reconstructing Text from Contextualized Word Embeddings by Inverting the BERT Pipeline.” Journal of Computational Literary Studies 2 (1). https://doi.org/10.48694/jcls.3572.
Langham-Putrow, Allison, Caitlin Bakker, and Amy Riegelman. 2021. “Is the Open Access Citation Advantage Real? A Systematic Review of the Citation of Open Access and Subscription-Based Articles.” Edited by Sergi Lozano. PLOS ONE 16 (6): e0253129. https://doi.org/10.1371/journal.pone.0253129.
Schöch, Christof. 2023. “Repetitive Research: A Conceptual Space and Terminology of Replication, Reproduction, Revision, Reanalysis, Reinvestigation and Reuse in Digital Humanities.” International Journal of Digital Humanities. https://doi.org/10.1007/s42803-023-00073-y.
Schöch, Christof. 2026. “Derived Text Formats as Strategic Transformations of in-Copyright Materials to Support Open Science: A Survey.” In Leveraging Derived Text Formats to Unlock Copyrighted Collections for Open Science: Book of Abstracts, edited by Trippel, Thorsten, Barth, Florian, Tello, Jose Calvo, Genêt, Phillipe, Lendvai, Piroska, and Schöch, Christof. https://api.zotero.org/users/228821/publications/items/3RZJJCIK/file/view.
Schöch, Christof, Maria Hinzmann, Julia Röttgermann, Katharina Dietz, and Anne Klee. 2022. “Smart Modelling for Literary History.” International Journal of Humanities and Arts Computing 16 (1): 78–93. https://doi.org/10.3366/ijhac.2022.0278.
Scholger, Martina, S. Strutz, and Christopher Pollin. 2024. “Empowering Text Encoding with Large Language Models: Benefits and Challenges.” In TEI Members’ Meeting and Conference 2024. https://doi.org/10.5281/zenodo.13969082.
Trippel, Thorsten, Florian Barth, José Calvo Tello, Phillipe Genêt, Lendvai, Piroska, and Christof Schöch, eds. 2026. Leveraging Derived Text Formats to Unlock Copyrighted Collections for Open Science. LREC.
Trippel, Thorsten, Florian Barth, Jose Calvo Tello, Keli Du, Phillipe Genêt, Daniel Kurzawe, Peter Leinen, Piroska Lendvai, Christof Schöch, Andreas Witt, and Arden Zimmermann. 2026. DIN 19461: A National Standard for Derived Text Formats.” In Leveraging Derived Text Formats to Unlock Copyrighted Collections for Open Science (Workshop at LREC 2026), edited by Trippel, Thorsten, Barth, Florian, Tello, Jose Calvo, Genêt, Phillipe, Lendvai, Piroska, and Schöch, Christof.
UNESCO. 2021. UNESCO Recommendation on Open Science.” UNESCO. https://doi.org/10.54677/MNMH8546.
Vaucher, Romain, and Camille Thomas. 2026. “Diamond Is the New Green — Why Green Open Access Is Not a Sustainable Long-Term Model for Scientific Publishing.” Sedimentologika 4 (1). https://doi.org/10.57035/journals/sdk.2026.e41.2397.
Weis, Joëlle, and Christof Schöch. 2024. “Vom Perler Hasenberg zur Lehmener WürzlayWeinetiketten digital Erschließen.” In Digital ist besser? Sammlungsforschung im digitalen Zeitalter, edited by Katharina Günther and Stefan Alschner. Göttingen: Wallstein. https://api.zotero.org/users/228821/publications/items/KL289SR5/file/view.

Bonus slides

The Open Access Citation Advantage

What is ‘Repetitive Research’?

  • A mode of research that repeats and/or is repeatable
  • A conceptual space with three dimensions: research question, dataset, and method/code
  • Several forms of repetitive research
    • Follow-up research: similar question, but new or similar data and method
    • Re-use & re-analysis (of data): same data
    • Revision (of method): same question and data
    • Replication (of research): same question, data and method