Lab Expertise

Digital Editions

The Elijah Lab specializes in producing accurate and innovative digital editions based on open-source principles, standardized formats, and current scholarly standards. We support the entire publication process for historical texts from its earliest stages – from scanning and OCR or HTR processing, through textual correction and editing, tagging and syntactic annotation, to the design of the edition and its publication on web-based platforms.
Our editions are typically based on the TEI (Text Encoding Initiative) format, which enables precise representation of textual structures, documentation of textual variants, linguistic annotation, and both internal and external linking. We also develop multi-layered annotation systems that support different levels of interpretation, conceptual tagging, entity linking, and dynamic presentation of the edition—for example, in geographical, visual, or conceptual contexts.
In many projects, we produce multilingual editions, such as German–English and Hebrew–Judeo-Arabic editions, structured to facilitate philological, translational, and cross-cultural comparison.

Databases for Qualitative Research

The Lab specializes in building complex databases from diverse sources - including documentary materials, testimonies, and social, media, and cultural sources - tailored to question-driven qualitative research.

Each database is tailored to the specific needs of an individual researcher or research team and is designed from the outset to integrate with the Lab’s other capabilities, including language processing, image analysis, linked texts, geographic mapping, and visual analysis.
The emphasis is not only on storing information, but also on developing flexible conceptual models that support visualization, visual exploration, and the interpretation of complex relationships. We combine standard tools with customized interfaces, recognizing that qualitative research requires not only data, but also ways of accessing and engaging with it that foster interpretation, understanding, and reflection.

AI in the Humanities

The Elijah Lab seeks to deeply integrate artificial intelligence into humanities research, adapting advanced tools to the distinctive needs of researchers across a wide range of fields, including history, literature, philology, sociology, law, and education.
Our work includes automated document processing, large-scale corpus analysis, and the development of advanced natural language processing (NLP) models tailored to historical languages, multilingual sources, and texts with complex structures.
We also develop and adapt machine-learning tools for researchers without programming expertise, with an emphasis on user-centered design, extracting insights from diverse source materials, and maintaining transparency and openness to multiple interpretations of the results. For us, artificial intelligence is not a substitute for critical thinking, but a means of deepening it: a tool for identifying patterns, uncovering hidden connections, and generating new questions from the depth and complexity of the data.

Historical Mapping and GIS

Development and application of tools and methods for the geographic analysis of historical sources, combining spatial analysis, GIS technologies, and textual and visual analysis.

  • Image processing and spatial positioning of visual sources – including collaborations with projects in cartography, historical geography, and the visual analysis of photographs and maps.
    Development of interactive interfaces for visualizing complex bodies of knowledge – dynamic historical atlases that integrate textual, geographic, and temporal information.
    Manuscript mapping, using both manuscript metadata and the extraction and processing of geographic information embedded within the texts themselves. Combining these capabilities enables researchers to trace patterns of dissemination, traditions of writing and copying, and the spread of texts and ideas across time and space.

From Document to Database

Classification, extraction, linking, and the construction of knowledge networks from multilingual archival materials, with an emphasis on book history and demographic data.

We engage in the in-depth processing of archival sources through what we call “deep cataloging.” This includes named entity recognition (NER), linking entities to relevant ontologies, and automatically establishing connections between records from diverse sources, even when they are written in different languages or use inconsistent formats.

For each project, we develop a customized data model based on the understanding that an archive is a living network of information rather than simply a static collection of documents.

This distinctive expertise enables us to extract historical and social insights in a variety of ways: through the analysis of place lists, as in the Prenumeranten project; through the processing of reader lists and catalogs from historical libraries—from scanning and text recognition to interactive presentation, as in the Strashun Library Project; and through the extraction of demographic information from sources such as Ottoman population registers.

By combining technology, historical research, and domain-specific design, we enable rich, multidimensional, and meaningful exploration of archival sources.

Identifying Connections Between Texts

Development of tools and methods for identifying intertextual relationships across texts in multiple languages, including Hebrew, Aramaic, Arabic, and Ancient Greek, using ACT, an innovative environment for detecting intertextuality developed at the Lab.

These capabilities are based on advanced natural language processing technologies and make it possible to identify not only exact quotations, but also paraphrases, reworkings, and stylistic or thematic variations. One of the key innovations of this work is the ability to detect relationships across different dialects and linguistic registers, and even across languages – for example, between Judeo-Arabic and Classical Arabic.
These tools also support large-scale stylometric research, the identification of hidden relationships among groups of manuscripts, and the internal mapping of manuscripts based on their constituent literary and linguistic units.

Making Hebrew Manuscripts Accessible

The Elijah Lab specializes in advanced solutions for Hebrew handwritten text recognition (HTR), using cutting-edge tools such as eScriptorium and Transkribus.

We support researchers throughout the process of digitizing, deciphering, and converting manuscripts into searchable and editable formats. Beyond transcription itself, we integrate HTR outputs with natural language processing (NLP) tools, enabling linguistic, semantic, and structural analysis of historical texts.
This integration makes it possible to produce digital editions in TEI format, as well as to link text and image through advanced computational methods as part of a multilayered system for making both visual and textual sources more accessible.

Distant Reading of Rabbinic Midrash

The Elijah Lab develops computational tools for the analysis of rabbinic literature, with a particular focus on late midrash, aggadic compilations, and related literary corpora. We combine techniques of distant reading with philological and syntactic expertise in order to uncover underlying patterns in the relationships between midrashic units, identify genres and subgenres, and analyze processes of editing, incorporation, and reuse within large-scale textual collections.
A central focus of our work is understanding the dynamics of midrashic traditions as documented in manuscripts, particularly fragments from the Cairo Genizah. We explore new ways of mapping and analyzing networks of textual variants, sources, and textual layers. Our approach combines computational technology with literary sensitivity, making it possible to study midrashic literature not only as a textual tradition, but also as a living and evolving phenomenon.