Based on a transcript of my recent 20-minute presentation at the ISF workshop “Artificial Intelligence in the Service of Talmudic and Rabbinic Literature Research” at Hemdat Academic College on 8-Oct-2026.1
Thanks so much to Menachem and all the organizers for this workshop. As others have mentioned, it’s special to have it here in the south. Being in Sderot on October 7 was especially meaningful. Thanks for putting this all together.
We’ve heard a variety of perspectives on AI, including its use in education. I work in industry, and I’d like to focus on the cutting edge: recent developments beyond chatbots, and how I’ve been using them. In a sense, chatbots are already almost “traditional.”
A bit about me: I work in tech and have a graduate background in Talmud. My academic interests, which I pursue on the side, include onomastics, Talmudic Aggadah, literary structure, and digital editions. I have a blog with a few hundred posts, which sounds more impressive than it is because I use AI assistance and have developed a whole pipeline. Of course, that also involves curation, checking, and annotation.
I’m also developing a ‘vibe-coded’ website, using natural language to build a site that draws on texts from Sefaria and presents them in my own opinionated way. I’m trying to make these materials accessible to a wider audience. I also have academic papers and videos available online.
So, what has changed over the last year? I recently saw a comment on Hacker News, a major technology forum, arguing that AI, at least LLMs, isn’t really coming up with anything new. This was even after the release of GPT-6 and its impressive problem-solving results. I think there’s still something to that. This idea was also discussed at yesterday’s roundtable with Lynn, Nachum, and others. Nachum pointed out that other forms of AI can do certain things that LLMs cannot (such as manuscript joins, and dating and provenance of medieval scripts). But what LLMs are doing is making existing kinds of work much faster. They allow individuals and much smaller teams to work more quickly and iteratively. It’s not magic, but it’s a tremendous acceleration.
In a previous workshop I discussed translation, summarization, tagging, outlines, tables, and similar applications. I’d like to focus instead on the more recent capabilities.
Someone mentioned sub-agents earlier, and I know others here, including Hillel Gershuni and Josh Waxman, are also working at the cutting edge. I’ll discuss subagents, different platforms, where chatbots remain relevant, and a few recent case studies.
The division of labor is still basically the same: you have a question, you prompt the AI, and then you validate the result. But there’s another important step, especially with coding agents: using scripts. I don’t think we’ve talked about that enough at this workshop.
First, some terminology. These terms can be context-specific and sometimes complicated, but they’re important. An LLM, or “large language model,” is technically a subset of AI, although it’s what most people mean by ‘AI’ nowadays. Well-known examples include GPT and Claude. At the cutting edge, we now have GPT-6.1 and Claude Opus 5.5. For this talk, when I say ‘AI’, I generally mean LLMs.
A chatbot is the familiar interface where you ask a question and receive an answer, usually in a browser. What I want to discuss is the coding agent. Instead of just answering questions, it actually does things. It can work on your computer, write files, and produce outputs. It has vision capabilities, can correct its own errors, and can change course along the way. That greatly expands what’s possible.
Then there are orchestrators and sub-agents. The main agent, often called an orchestrator, can spin up sub-agents to handle specific tasks. I’ll return to that shortly.
There are other relevant terms, including context, tokens, and skills. I won’t go through each one now.
What have been the biggest changes? First, much better models. The current state of the art includes GPT-6.1 and Claude 5.5. Many people are still using free or weaker models, and that shapes their impression of AI. Hallucinations are still a problem, but grounding has improved substantially. There is a real gap between free and paid models, weaker and stronger models, and frontier models more generally.
Back to coding agents: tools that write files, run scripts, produce materials, and access the internet. Anthropic has Claude Code, and OpenAI has Codex. I subscribe to both, at around $20 a month each. These are among the main tools for agentic work.
What can they do? They write and execute scripts. A year ago, I would ask a chatbot something like, ‘Write me a script to count the words in this Talmud file.’ It would produce code, which I’d copy and paste somewhere else, troubleshoot, and run. Now I can give the task to a coding agent. It can download the file, write and run the script, find bugs, and fix them. It’s very powerful and very fast.
And it’s not limited to scripting. An agent can produce PDFs, Word documents, and slides like these. For the original draft of this presentation, I told Claude Opus 5.5 to look at my notes and produce a PowerPoint. It did (followed by many rounds of iteration). Earlier tools had some of these capabilities too, but they made far more mistakes. Now agents can use “vision” to inspect their output, iterate, and correct errors.
Even in industry, many people are still using chatbots, which is understandable: they’re much more accessible. But for the most powerful, cutting-edge work, I think agents are where it’s at.
Search and discovery have also improved. Hallucinations remain a problem, and all the usual caveats apply. But models are much better grounded and can search the internet and prepare research documents. I used to say that AI wasn’t very good for searching and discovering sources. At the time, I think that was true. Today it’s considerably better. It still makes mistakes and misunderstands things, but it can produce a very useful first draft.
Agents can access the internet, install external libraries, and even use a computer or browser to perform tasks. They can run code, research, retrieve data, and deploy projects. Over the last few months, I’ve built websites with GitHub Pages and published them on GitHub, with the agent doing much of the building and deployment autonomously. You can say, ‘Here’s what I want. Build this website,’ and it can do a remarkable amount of the work.
There’s also a category of third-party tools designed especially for building user-facing websites. For a new, ‘greenfield’ project, I personally use Replit. There’s also Base44, owned by the Israeli company Wix. These tools let you describe a website in natural language and have the platform build and deploy it.
They can be expensive, increasingly so as models grow more powerful and consume more tokens. But for people with the budget and interest, I recommend exploring them. They’ve continued to improve, with features like buying domains, deployment, and code review becoming much more accessible.
A word about chatbots. I still use them all the time for simple questions. But a year ago, there was considerable excitement, including on my part, about building custom chatbots into websites. I called my own website “ChavrutAI” because I envisioned including a chatbot that would act as a chavruta. Yesterday’s roundtable also discussed virtual rabbis or poskim, and Dicta has its Rav Dicta.
I’ve become less enthusiastic about that idea, and I think the industry has moved somewhat in the same direction. Custom website chatbots haven’t become nearly as widespread as many expected. I’ve pivoted away from that approach myself. Sefaria recently launched an AI feature that I think points in a better direction: using AI for search, discovery, and exploring sources, rather than asking it to act as your rabbi and tell you what to do. It’s less normative and more descriptive. Dicta also has a tool called Iluy, which I think is more in that direction.
Before the case studies, a few best practices for coding agents. I’ve been using them almost every day for the last year, so these observations come from experience.
Start with a sample. If you’re working with a huge text, like the Talmud or a manuscript, start small. By default, an agent might try to process the entire corpus, consuming all your tokens and budget. Tell it to begin with a few lines or a page, and set explicit limits.
Use Git commits. Git lets you take snapshots of a project, and GitHub lets you store them remotely. This not only preserves your work but also helps the agent manage a project over time.
Maintain documentation and skills. The concept resembles the ‘insight files’ Asaf mentioned. Essentially, you’re maintaining text files that explain what the project is, what the agent should know, and what has gone wrong previously. When a mistake occurs, you can tell the agent to update the documentation so that it can (hopefully) avoid repeating it.
Keep reusable scripts, and know when to restart. Especially for recurring tasks or long-term projects, scripts should be part of the project. And if an agent starts making serious mistakes, it can still be useful to begin a fresh session, perhaps with a stronger model. Ask for a handoff summary to give the next agent.
Monitor long-running tasks. Some people leave agents entirely alone, but I prefer to keep an eye on them. They’re extremely powerful and know far more than I do in many areas, but I can still contribute to managing the process. Agents don’t necessarily have a good sense of elapsed time. A script might run for ten minutes, and I’ll notice that something seems wrong. I can ask whether it’s processing correctly or whether there’s a more efficient approach. In principle, an agent can keep going for hours without getting tired.
Watch costs. The more powerful models can be expensive, and it’s easy to spend substantial amounts, especially when running large jobs. Most of us don’t have limitless budgets; I certainly don’t.
Review the whole project periodically. It’s worth asking the agent to step back, optimize the workflow, and clean things up.
Some older practices are becoming less important. I think formal ‘plan mode’ is a lot less relevant than it used to be. The same goes for prompt-engineering tricks. Clear instructions and good project documentation are definitely still important, but you no longer need as many special formulas to get useful results.
Here’s an example of a recent prompt: I tell the agent to start a new project in a folder and initialize Git. I also give it my current ideas while explicitly inviting it to propose better methods. Agents won’t always push back on your suggested approach unless you ask them to. Often, when given that freedom, they find methods I hadn’t considered.
Now, orchestrators and sub-agents. Think of an orchestrator as a team lead or chief of staff, with sub-agents doing the specific “tedious” tasks.
There are several advantages. First, cost: you can have a powerful model coordinate the work while sending some tasks, particularly vision tasks, to cheaper models such as Gemini or other alternatives. Second, speed: sub-agents can work in parallel. Third, cleaner context.
Context is somewhat like working memory. There’s only so much information a model can keep in its active context at once. If you distribute work among ten or twenty sub-agents, the main orchestrator has less detail to manage. Different agents can also specialize in different tasks.
The orchestrator, being a powerful model working at a higher level, can also check feasibility, test samples, evaluate accuracy, and decide whether the full project is worth undertaking.
Let me turn to a very recent project, from about two weeks ago, involving OCR and, more specifically, handwriting text recognition (HTR).
The source is a fourteenth-century Italian cursive Kabbalistic manuscript. I’ve been interested in it since graduate school. It’s an unpublished commentary on the Yeriah Gedolah, itself a difficult and obscure text. The original Yeriah Gedolah was published about twenty years ago, but this commentary has never been published.
For years I tried reading it myself. Later, with the development of AI, I tried Gemini and ChatGPT, but they struggled. Two weeks ago, I made by far the most progress I’ve ever made. The results may be approaching human-level transcription, though I’m not entirely sure yet.
I gave the manuscript to Claude Opus 5.5, and it discovered tools I hadn’t even known about. For example, Ktiv provides access to high-resolution manuscript images through IIIF, a standard for serving and structuring digital images. The agent could retrieve those images directly.
Claude then used Kraken and open models, including, I think, one called Sofer Mahir. It divided each page into smaller regions, deskewed the images, structured the material, and attempted line-by-line proofreading and correction.
It also used vision and distributed tasks among its own sub-agents, including less powerful agents within the Anthropic family. The work proceeded remarkably quickly.
But it was expensive. Multiple agents running in parallel, especially with vision, can consume a lot of tokens. I estimate I spent more than $150 in API costs, so I paused the project.
The system was also collaborating with Codex. That’s another advantage of Git: different agents can work on different aspects of the same project. Claude and its sub-agents handled transcription and the processing pipeline, while Codex and other models worked on translation and annotation. It’s a powerful workflow, but again, potentially expensive.
Another case study involves curation on my website: Jastrow and the Brown-Driver-Briggs lexicon (BDB), available on Sefaria. These are dense older reference works, full of abbreviations, archaic terminology, and scholarly citations. The risk is relatively low because I’m not changing the underlying text. Instead, I’m expanding abbreviations, mapping names, and making the material easier to navigate. I’ve blogged about this. BDB, in particular, is extremely hard for me to read in its original form because so much is abbreviated or unfamiliar. I built this interface with third-party tools and have been iterating on it continuously.
I also implemented transliteration. BDB contains Greek, Arabic, and other languages. In the nineteenth century, scholars often assumed readers could handle Greek, Arabic, Ge’ez, Syriac, and more. Today automated transliteration is relatively straightforward. I know a little Greek but no Arabic or Ge’ez, so this makes the material much more readable for me personally.
To conclude, as I said yesterday, I hope LLMs will eventually produce genuinely new chidushim. I don’t think they’re quite there yet. But they already allow an individual scholar like me to do work that would previously have required an entire team or would have taken much longer.
Q&A
Question: When you use Git to keep versions so you can return to earlier stages, what exactly are you saving? As someone who uses graphical interfaces, what would that look like?
EB: Most developers use GitHub. That’s what I use, and it’s free. I’m not a professional developer; I suppose I’m a vibe coder. But GitHub is the standard platform for this.
Git runs locally on your computer and takes snapshots of your work. On GitHub, you create what’s called a repository, which can hold a project. For my Yeriah Gedolah project, for example, I created a new repository, and I tell the agent to push snapshots containing the code. I don’t push everything: manuscript images are large and copyrighted, and I can retrieve them again when needed.
My code is public. These are open-access projects, and I’m not trying to monetize them. Publishing the code helps me, helps the agent, and makes the project accessible to others.
Audience member: The repository doesn’t have to be public. GitHub also offers private repositories, including free options, so you can keep some work private and make other parts public.
Audience member: Some people here understand all this technical language, but others, like me, don’t. You can put the presentation into a chatbot and ask it, step by step, what to do. It can guide you through opening a GitHub account and using these tools. You can use an agent as your personal computer adviser.
And whenever you don’t understand, stop it and say, ‘I didn’t understand a word of that. Please explain in plain English or Hebrew, without all the buzzwords.’ I do this all the time.
Another recommendation is code review. If Claude writes your code, ask a different agent, such as Codex, to review it. It will often find mistakes, and you can go back and forth between them.
Question: Sub-agents are new to me. How do you actually set them up? Do you tell the orchestrator how many to create? Do you specify their tasks? What’s the best strategy?
EB: You can simply say, ‘Spin up sub-agents.’ In my manuscript project, the agent did that on its own. For bigger projects, it will sometimes take the initiative. I’ve also seen this in third-party tools like Replit.
Audience member: You can just ask?
EB: Yes. You can ask, but it will often do it on its own.
Audience member: One major distinction is whether you’re using a tool in the browser or running it on your own computer. It may involve the same model provider, but it’s a different working environment.
EB: Right. For the full coding-agent experience, you generally need a coding-agent interface rather than a conventional chatbot, at least as far as I know.
Audience member: How would you recommend using these tools with a research team? I’m a scholar, and I don’t have the time or inclination to work with agents myself. I may now need to spend less on programmers and can put more resources toward humanist research assistants. What’s the most effective way to incorporate AI into a team?
EB: That’s a good question. I’ve generally worked independently, so I have less experience managing collaborative research projects.
I don’t know that I have special insight beyond the usual division of labor. Git and GitHub are designed for collaboration, and they can be used privately as well as publicly.
That collaborative model is increasingly relevant now that agents are becoming central to software workflows. The tools are improving at coordination not only among people, but between people and agents and among agents themselves. It’s worth exploring shared repositories, versioning, and clear divisions of responsibility.
Includes audience Q&A at the end.
Slides used in the talk are here: https://www.academia.edu/178527494/Automating_Assistance_AI_as_a_Versatile_Multifunctional_Tool_for_Talmudic_Research
I was glad to receive a lot of positive feedback on my talk. In general, I learned about many very interesting projects happening in the space. It was an amazing experience to spend substantial time with a group of scholars and researchers working on topics so closely aligned with my own interests. I plan to write a separate post about the workshop as a whole, sharing my impressions and some photos.

