Claude at work. (Screenshot of in-progress convo in Claude Code local coding agent for Windows.)
I have started publishing an experimental digital edition of Reuven Tzarfati’s Commentary on the Great Parchment: a medieval Hebrew kabbalistic commentary preserved in Jewish Theological Seminary MS 2367, folios 76b–97b. The project brings together manuscript transcription, English translation, annotation, and comparison with a surviving diagrammatic witness. You can read the edition online or explore the files and code on GitHub.
The feature I find especially exciting is Claude Opus 5.5’s role as an orchestrator of the transcription pipeline. In this project, an AI agent can work across the sequence from obtaining manuscript images to producing a corrected, line-numbered transcription. It can run specialized software, inspect the resulting images and text, write corrections into files, and return to a doubtful passage. That ability changes the practical scale of what I can attempt with a difficult Hebrew manuscript.
I wrote about Ha-Yeri‘ah ha-Gedolah in a 2019 Seforim Blog essay. The composition’s compressed language and unusual combinations of biblical images had interested me for years. This project gives me a way to work through the commentary systematically and make the accumulating results available for others to inspect.
The work belongs to early fourteenth-century Italian Kabbalah; the JTS copy was made in 1392. Tzarfati comments on Iggeret Sippurim, the “Epistle of Narratives,” a composition transmitted in the Great Parchment tradition. Its biblical phrases, names, and images are arranged within a kabbalistic diagram, an ilan, or tree. Rivers, patriarchs, animals, bodily parts, and commandments acquire positions and associations within the sefirotic structure. The commentary explains how to read those associations.
This makes the manuscript particularly interesting for Jewish digital humanities. Understanding it involves handwriting, biblical and rabbinic allusion, textual transmission, and visual arrangement. The edition puts the commentary into conversation with the Oxford Great Parchment, Bodleian Library MS Hunt. Add. E, accessible through the Ilanot Portal. That parchment is another witness to the underlying composition; the JTS manuscript supplies the commentary I am transcribing and translating.
What an AI orchestrator does
A transcription begins well before the first Hebrew word is typed. I need a sufficiently clear photograph, a reliable way to identify the page, and manageable views of its lines; I also need a way to preserve partial work and know which reading belongs to which part of the manuscript. These tasks can consume much of the time available for reading.
The project uses high-resolution images obtained through the National Library of Israel’s IIIF service. IIIF is a standard way for libraries to expose digitized images and their metadata. The download code assembles small image tiles into a full page. That gives the reading workflow access to detail that can be lost in a lower-resolution PDF export.
Next comes kraken, specialized software for recognizing handwritten text. A model for Italian Hebrew produces an initial machine reading; a layout model identifies text regions and line baselines; while yet another script uses that information to make enlarged, deskewed crops of individual lines. Long lines are split into overlapping right and left halves, so the reading can follow Hebrew’s direction while retaining enough detail to distinguish letters.
Claude then works through small batches of those crops alongside the machine reading: It corrects the Hebrew, retains scribal abbreviations, marks uncertain words, and saves numbered lines to a Markdown file for the folio. The instructions specify concrete conventions: [?] for a doubtful reading, <..> for illegible text, and bold for the phrases that the commentator quotes and explains. The next pass concentrates on unresolved readings and discrepancies with comparative evidence.
The impressive part is the coordination: The agent can connect an image-download script, a handwriting recognizer, a crop generator, a reading task, and a persistent transcription file. Each stage leaves something the next stage can use: When the reading raises a problem, the workflow can move back to the image or another witness. An agent operating in a working directory can help manage that entire sequence.
The repository records a division of labor: Claude Fable assembled the initial image and recognition pipeline; Claude Opus 5.5 continued the transcription and reviewed translations against manuscript images; Codex developed the translations, annotations, comparative apparatus, research survey, and reading site. The current experiment has grown through that collaboration, with instructions and corrections carried forward in the repository.
For me, this is the exciting possibility: I can give the system a sustained research task whose intermediate steps include both programming and close reading. The files preserve progress between sessions; the scripts make recurring operations repeatable. The folio instructions describe the scribe’s habits and the expected output. Together, these let the work advance page by page without rebuilding the procedure each time.
Corrections
One early lesson is that a convincing English translation can hide a bad Hebrew reading: Once a mistaken word has been rendered into grammatical prose, it becomes easier to miss. The workflow therefore includes comparison before translation and a return to the manuscript when the witnesses disagree.
Folio 80b provides a concrete example. An initial transcription read a phrase as lo qibbalti, “I did not receive.” A subsequent image recheck identified natnu lanu qabbalah, “they gave us a tradition.” That changes the apparent statement about the commentator’s access to inherited knowledge. The correction is recorded in the folio’s notes, alongside other rechecked phrases. The Oxford parallel helped identify where to look; the JTS photograph supplied the evidence for changing the JTS transcription.
The page’s red marks supplied another practical lesson. A red bar generally sits in the space above a quoted phrase. In a tightly cropped image, a bar at the bottom can belong to the following line. Misassigning it changes which words are treated as a lemma and which as commentary. The project instructions now explain that relationship, and the crop script looks for red ink above the baseline. Even then, the reader must inspect the image: a color flag does not establish a quotation’s boundaries.
These are useful examples of what orchestration enables. A difficulty encountered during reading can lead to a change in instructions or software, followed by another check of the text. The next folio benefits from that accumulated experience. The pipeline includes a feedback loop between image evidence, editorial decisions, and the tools that prepare the evidence.
What readers can use now
As of this draft, the repository’s publication registry lists sixteen translated folios, from 76b through 84a. The combined reader extends through 96b, and separate working transcription files reach 97b, the end of the commentary. These are different stages of completion. The transcriptions remain drafts; several translated pages explicitly record unresolved readings and marginal material still awaiting treatment.
The reading editions pair Hebrew and English by line, with notes and Oxford parallels where supplied. Their line numbers let a reader locate the passage being discussed. The Hebrew remains visible beside the translation, and uncertainty is carried into the English. This is useful for readers who want help entering the work and for specialists who want to question a proposed reading.
I am treating the project as an experimental edition of one manuscript witness. Claims about a complete critical text would require a fuller comparison of the commentary’s surviving witnesses. Busi’s 2004 edition of the underlying Iggeret Sippurim and the scholarship on kabbalistic diagrams provide essential context, but this project has not checked its readings against a supplied copy of Busi’s volume. The repository’s research survey explains the documentary layers and bibliography.
What I can offer now is a substantial, inspectable working edition and a record of how it is being made. The combination of a specialist handwriting recognizer, an agent that coordinates tools and visual reading, and a second system that develops and checks translations is already productive. Claude Opus 5.5’s orchestration makes this especially compelling as an experiment: the model helps carry a manuscript through a sequence of technical and scholarly tasks, with intermediate results that can be revisited.
For Jewish digital humanities, I see considerable scope in that arrangement. Many manuscripts require this combination of image preparation, difficult script, abbreviated Hebrew, recognizable sources, and unfamiliar argument. I hope the project will attract corrections, comparisons, and suggestions from readers who know these materials. The online edition is the easiest place to begin; the repository makes the method and its working results available alongside it.
Appendix 1 - Technical: from image to reading edition
Source and scope
The base witness is JTS MS 2367, fols. 76b–97b: 43 manuscript sides. The local repository contains 43 canonical per-folio transcription files. At the drafting snapshot, its translation manifest marks sixteen folios as published, ending at 84a; the combined reader ends at 96b. Separate transcription files reach 97b. These counts describe coverage, not validated reading accuracy. The Oxford Great Parchment supplies comparative lemmata (דיבור המתחיל) from the underlying composition and remains distinct from the JTS commentary.
Pipeline and review
Identify and retrieve.
folio_fl_map.tsvmaps folios to NLI image identifiers.iiif_tiles.pyassembles 333-pixel tiles at the service’s native dimensions.download_all.pyanddownload_from.pycoordinate retrieval. Cached tiles support reuse; server rate limits may require reduced concurrency.Recognize and segment.
run_all_htr.pyinvokes kraken withItalian_01recognition and the SoferMahir layout model. It writes machine readings and a logical-order derivative for this workflow. ALTO XML records line geometry.Prepare visual evidence.
alto_lines.pyextracts enlarged, deskewed crops, separating main text and other regions and splitting long lines into overlapping halves.lines.tsvrecords geometry and red-ink flags. Crop identifiers anchor the transcription; machine-text numbering can differ and needs an explicit alignment note.Read and correct. Claude reads crops alongside the machine output, following
FOLIO_TASK.md: at most eight images per batch, immediate saving of corrected lines, preserved abbreviations, and explicit uncertainty. A second pass compares doubtful readings and lemmata.fetch_oxford.pysupplies zone-based comparative text; discrepancies prompt JTS image rechecks.Translate and present. Canonical Hebrew is stored in
transcription/NNNx.md; generator-managed English, notes, and apparatus are stored intranslations/Nx.md.build_translation.pyimports the Hebrew from the transcription and creates staged reading pages.translation_manifest.jsoncontrols draft and published navigation.combine_transcriptions.pyandbuild_site.pyassemble the reader.
Manuscript credit: Courtesy of the Library of the Jewish Theological Seminary, The National Library of Israel. “Ktiv” Project, The National Library of Israel. The public project publishes an independent transcription and does not reproduce the manuscript photographs.
Appendix 2 - A second manuscript witness: collating Munich 311
The JTS manuscript is not the only copy of Tzarfati’s commentary. Munich, Bayerische Staatsbibliothek, Cod. hebr. 311 preserves it on fols. 1a–33b, and the National Library of Israel serves its images through the same IIIF interface. After the JTS transcription reached the end of the commentary, Claude ran an experiment: could a second copy resolve the passages the JTS transcription had left uncertain?
The machine reading of Munich. Claude downloaded the relevant Munich folios and ran them through the same kraken pipeline. Layout detection was generally good, but recognition was weak. The Munich photographs are grey-scale microfilm scans at about a third of the JTS resolution. Aligned letter by letter against the JTS text of one section, the machine reading matched only 49% of the letters. On several pages, the line segmentation also broke into fragments. The conclusion was that the machine text can help locate the Munich passage that corresponds to a JTS line, but it cannot supply readings. Claude therefore read the Munich pages directly, as enlarged horizontal bands of about six lines, split into right and left halves, with close-ups where needed.
Claude did not transcribe Munich in full: It read Munich only where JTS was doubtful: at each word marked [?] and at each lemma. The first trial was the thirteenth narrative (JTS 93b–94b). The main effort then went to 90a–92a, the folios with the most uncertain readings (about 240 marked words). The pace was about 20–30 minutes per JTS folio. Munich’s scribe writes a clear, regular hand, and at many points where JTS is faint, blotted, or cramped, Munich is easy to read.
The Munich readings did not go straight into the edition; the project treats Munich as an independent witness. For each proposed change, Claude reopened the JTS line image and asked whether JTS itself supports the Munich reading. It applied about forty corrections this way; each one is recorded in a dated note with the old and new reading. About twenty-five points stayed unchanged, because the JTS image does not decide them. There the note records the Munich variant beside the retained JTS reading.
Results. Most of the corrections turned out to be errors in our first-pass draft rather than in the JTS scribe’s text. Munich showed what to look for; the JTS photograph then confirmed it. Several results change the sense:
92a:18–20. The draft read menorah (”lampstand”) three times, including in the lemma “Blessed are You, O Lord, who graciously grants knowledge…”. Munich reads Moshe, as does the Oxford parchment. On re-examination, the JTS scribe also wrote Moshe: our draft had misread a dotted abbreviation. The passage concerns Moses, whose name “is like his master’s name”.
90a:23. An unintelligible
ובינים שמאתis u-ve-yom simhat libbo (”on the day of the gladness of his heart”). The commentary continues Song 3:11 after “the crown with which his mother crowned him”.90b:4. The lemma is the talmudic principle “there is authority to the written text and authority to the transmitted text” (B. Sanhedrin 4a). This is exactly what the following lines explain.
92a:47–49. Readings taken as “body” and “soul” are the abbreviations for the attributes of mercy and judgment. In this reading, qanqan (”jug”, from “do not look at the jug”) splits into qan qan, so the sense changes.
Munich is not a better text in every respect: Its scribe skipped whole phrases twice in this stretch, jumping from one repeated word to its next occurrence. Because the Munich photographs show no red ink, only JTS marks where a quoted phrase begins and ends.
The experiment shows how the pipeline extends to a second witness: The same agent that transcribed JTS located the parallel passages, read a different hand from poorer images, and fed the results back into the base transcription. It also kept a written record of what each witness says. The next step would be a recognition model trained on a few hundred corrected Munich lines, which would make fuller collation cheaper. London, British Library Add. 27078 (dated 1573) is a third copy awaiting the same treatment.
The collation notes are in each folio’s transcription file, and the method is described in witnesses/munich311/PILOT.md in the repository.


