Context-Based Vocabulary Learning — What the CALLIDUS Project Shows

“You shall know a word by the company it keeps.” — With this sentence, linguist J. R. Firth summed it up (A Synopsis of Linguistic Theory, 1957): A word only reveals its meaning through the words around it.
What if we learned Latin vocabulary exactly that way—not as a word-to-word pairing on a flashcard, but in real sentences? A research project at Humboldt University in Berlin systematically investigated precisely this question: CALLIDUS. It not only advocated for context-based vocabulary learning in Latin but also turned it into a functional digital tutor—and in the process produced honest, instructive results. This article introduces the project: its underlying premise, what makes the approach unique, the study’s findings, how students responded—and what we’ve built upon it for effective Latin learning.
What was the CALLIDUS project?
Context-based vocabulary learning was at the heart of CALLIDUS, a DFG research project at Humboldt University of Berlin (Beyer & Schulz). It developed the e-tutor Machina Callida, which automatically generates contextualized vocabulary exercises from real Latin text corpora. The basic premise: You don’t learn a word in isolation, but through the sentences in which it appears.
The key: the mental lexicon
This basic assumption stems from cognitive science. Our brain does not store words like a dictionary—neat entries, sorted alphabetically. It stores them as a network. Seidenberg and colleagues put it this way as early as 1989 (cited in Beyer/Schulz 2020):
“Lexical memory does not consist of entries for individual words … Knowledge of words is embedded in a set of weights on connections …” (Seidenberg et al. 1989)
A word is thus the sum of its connections—Firth’s “company it keeps.” From this follows the key insight: Unconnected learning does not allow for transfer. Anyone who knows “res” only as a “thing” will be at a loss in the next sentence by Caesar where it says “state.”
“A Latin word has no German meaning; nor does it have a Latin meaning—it simply has a meaning.” (Hermann Steinthal 1971, cited in Beyer/Schulz 2020)
Key point: A flashcard stores a word. A sentence stores its meaning.
Added to this is the cognitive load (Sweller 1988): isolated lists overload working memory, while context reduces the load. For those who want to dive deeper: We’ve dedicated a separate article to cognitive load in Latin learning.
What makes this approach special: corpus-based and automated
What made it special was implementing context in a corpus-based and automated way. Machina Callida extracts authentic sentences from text corpora (PROIEL, Perseus, AGLDT) and uses them to generate “comprehensible input”—real Latin, automatically processed, with three exercise types:
| Exercise Type | Principle |
|---|---|
| Mark words | Mark word forms in the text (Noticing) |
| Matching | Matching collocations |
| Cloze | Cloze exercise: Infer the word from the context |
Behind this lies a second observation: Meaning is often found in the combination of words—facere rarely simply means “to do”:
| Latin Phrase | Literal | Actual meaning |
|---|---|---|
| pontem facere | “to make a bridge” | to build a bridge |
| fidem facere | “to establish trust” | to build trust |
| aliquem reum facere | “to make someone guilty” | to accuse someone |
All three are from Caesar, *Bellum Gallicum* I. And here’s a tough statistic: To truly understand a text, you need to know at least 95% of the words in a paragraph (Nation 2006)—a goal achievable only through a well-established vocabulary that can be recalled in context.
Understand Latin instead of cramming — try it right now.

From Research to Practice
CALLIDUS has demonstrated what context-based practice can look like—focusing on sentences, collocations, and meaning within context. The implications for teaching and personal learning are detailed in the practical article “Learning Latin Vocabulary in Context.”

The Vision: A Word That Remains Retrievable in a New Context
“The goal of this new perspective on Latin lexical acquisition is to ensure that information about a word can still be retrieved from the mental lexicon even when the word appears in a new context.” (Beyer/Schulz 2020)
A high, precise standard: not “more vocabulary,” but retrievable, transferable vocabulary.
Honest results—science that acknowledges its limitations
This is where the project deserves special respect: The researchers reflected on their pilot study (2020, four weeks) and assessed it openly. They do not hide the fact that the post-test initially showed no measurable increase in vocabulary. Instead of glossing over it, they clearly identified the cause—it lay not in the principle itself, but in the study design. This is precisely where the lasting value lies: science that doesn’t oversell itself.
What the participants demonstrated—and what lessons the project drew
The key lay in the students’ behavior: too many unfamiliar formats at once; the unfamiliar format captured their attention—leading to cognitive overload. The lesson that Beyer and Schulz draw: focus on one or two formats so that learners can get used to the methodology. It’s remarkable what still worked: teachers rated polysemy and collocation formats as “very helpful and useful” (Beyer/Schulz 2020); the students worked through them successfully—so intensively, in fact, that they took ten minutes longer than expected.
From Sentence to Story—Taking the Research Further
There is one thing CALLIDUS deliberately does not address: that sentences, in turn, belong to a coherent story. The project establishes context at the level of the sentence and collocation (“comprehensible input”; meaning “only at the text level,” Beyer/Schulz 2020)—and it would be disingenuous to attribute more to it.
However, learning psychology carries this same central theme—the reduction of the burden on working memory—to the next level, from word to text. Anne Friedrich (Dyslexia and Latin Instruction, 2017): Younger learners process information primarily in an event-based manner; disconnected individual sentences force working memory to clarify anew in each sentence who is currently acting—whereas recurring characters relieve it of precisely this burden. This creates a stable situational model that can be “more easily anchored in long-term memory” than isolated microstructures.
Read in this light, Friedrich does not contradict CALLIDUS; she reinforces and expands upon this line of thought. It is precisely at this point—vocabulary in a sentence, a sentence in a story—that our own approach begins; you can read about what it looks like in the practical article: Learning Latin Vocabulary in Context.
Building on CALLIDUS’s work—what we’ve developed from it
This is exactly where we begin—cautiously, without claiming more than is substantiated.

For students, we apply the context framework in two ways. First, in measured doses: A box system based on Leitner’s method introduces new formats gradually over several weeks, rather than bundling them together as in the four-week pilot. Second, in a coherent manner: vocabulary, verb forms, and sentence elements are practiced using the same sentence material—the learner encounters the same sentence again in the verb form trainer, the sentence quiz, and the reading passage, embedded in a continuous story with recurring characters.

For teachers—and with their support—we’re delivering what CALLIDUS had hoped for but couldn’t test in the field: diagnostics. While students practice in context, a research dashboard and a teacher dashboard (Officina) quietly collect data in the background—which meaning, which form, which sentence component is causing trouble for which student—and use this information to provide targeted support. The learner is unaware of the assessment; the teacher sees where to focus their efforts.
We, too, aren’t “proving” anything overnight with this. What we’re doing is making this widely accepted didactic principle continuously applicable—and finally measurable.
Frequently Asked Questions About the CALLIDUS Project
What is the CALLIDUS Project?
CALLIDUS was a DFG research project at Humboldt University of Berlin focused on corpus-based, context-oriented vocabulary work in Latin instruction. It developed the e-tutor Machina Callida, which automatically generates contextualized vocabulary exercises from authentic Latin texts.
Has CALLIDUS proven that context-based vocabulary learning is more effective?
No. The 2020 pilot study showed no measurable increase in vocabulary in the post-test—primarily because the four-week study was too short and the students were overwhelmed by too many new exercise formats. There is broad pedagogical consensus on the approach, but no hard evidence of its effectiveness.
Why isn’t it enough to memorize the German translation?
Because a Latin word usually has multiple meanings, and it’s the sentence that makes the meaning clear. res can mean “thing,” “situation,” “state,” or “possession.” If you only learn one translation, you’ll struggle with every sentence that uses a different meaning.
What is Machina Callida?
The digital tutor from the CALLIDUS project. It extracts authentic sentences from text corpora and uses them to generate exercises such as “Mark Words” (highlighting in the text), “Matching” (collocations), and “Cloze” (fill-in-the-blank exercises).
How does Lapicida implement the CALLIDUS concept?
Through a comprehensive system rather than a four-week pilot: a box system that introduces new formats in measured doses and helps you become familiar with them over several weeks, as well as context-based modules like Sensus (the correct meaning in context). You can read about how this works in practice in the article “Learning Latin Vocabulary in Context.”
What remains: a method you can build on
CALLIDUS has researched a method, formulated a precise goal, and honestly demonstrated what matters—context, pacing, and retrievable vocabulary. That is precisely what makes the project valuable: as a foundation, not as a ready-made recipe. How we build on this—and how you can learn in context yourself—is explained in detail in the practical article: “Learning Latin Vocabulary in Context.”
Sources
More on the project: Publications of the CALLIDUS Project (Humboldt University of Berlin).
Sources (selection): Beyer/Schulz 2020 (CALLIDUS—Corpus-Based, Digital Vocabulary Work in Latin Instruction) · Kuehnast/Schulz/Lüdeling 2024 · Friedrich 2017 (Dyslexia and Latin Instruction) · Firth 1957 · Seidenberg et al. 1989 · Steinthal 1971 · Nation 2006 · Sweller 1988. (Seidenberg, Steinthal, Nation cited in Beyer/Schulz 2020.)
Understand Latin instead of cramming — try it right now.







