Roxxem Logo
Research,  Engineering

Inside the Roxxem Adaptive Test

Jingjing Ren, PhD

Previous post showed how Roxxem estimates each learner's level from the games they already play. That estimate needs history, and a new learner has none. A teacher still needs a placement on day one, and a district needs a baseline before the term starts. The adaptive test fills that gap in about 10 minutes, on the same IRT scale as practice.

The Roxxem proficiency test has four sections: Words, Sentences, Passages and Listening. This post explains what learners do in each, how their answers shape the test, and how we build, validate and refresh the bank of about 10,000 Spanish questions behind it. Other languages are in development and validation.

Why adaptive testing needs fewer questions

A fixed test spends most of its questions where they say little: a weaker learner guesses on items far above their level, a stronger one breezes through items far below it. As Part 1 showed, an item is most informative when its difficulty b is close to the learner's ability θ, where the outcome is least predictable. An adaptive test chooses every question to sit there, so it reaches the same precision with far fewer items. Roxxem adapts within each of four sections, described next.

What the four sections measure

The four sections move from individual words to sentences and passages, then to listening. They run in that fixed order. The section design draws on DET's published architecture (Nydick & Lockwood, 2024), adapted for low-stakes classroom placement.

The four test sections as learners see them: a Yes/No word (encuentro), a sentence with a blank and four options, a short passage with four candidate titles, and a dictation prompt with an audio button and a text box. Each screen offers "I don't know".

Figure 1. The four sections, in test order.

  • Words (Yes/No vocabulary). See a word, say whether it is real. Measures vocabulary breadth, after Meara's Yes/No test.
  • Sentences (vocabulary in context). A word is blanked from a real sentence; choose it from four. Measures vocabulary knowledge in context, in a multiple-choice gap-fill design.
  • Passages (title the passage). Read a short passage; choose the best title from four. Measures reading comprehension at the gist level, in a standard reading-comprehension design.
  • Listening (dictation). Hear a sentence, type it. Measures listening plus written accuracy.

Here an item is a single question, scored once: one word to judge, one gap to fill, one passage to title or one sentence to transcribe. In Part 1, an item was a piece of content played through a practice task and scored over a session; both kinds sit on the same scale.

The answer choices and scoring rules matter too:

  • Pseudowords catch overclaiming. 30% of Yes/No items are plausible non-words that follow the language's spelling and sound patterns. A learner who says yes to everything is caught at once.
  • Distractors are tempting for a reason. Wrong options match the answer in part of speech and agreement, so grammar alone cannot eliminate them. They are false cognates, words from the same semantic field, or common learner confusions.
  • "I don't know" beats guessing. A four-option item has a 25% guessing floor, which inflates estimates, most of all at low levels. Multiple-choice items offer "I don't know", scored incorrect but logged separately. In simulation, learners who used it instead of guessing were placed within one level of their true level 92% of the time, against 75% for guessers. Real usage is logged, so we can check whether learners take the honest exit.
  • Dictation is graded on closeness. Case, accents and punctuation are ignored; the rest is compared letter by letter with the reference, giving a grade from 0 to 1. A near-miss moves the estimate more than a blank answer, and a missing word costs more than a typo. The audio plays once and can be replayed twice, within one minute.

How the test adapts to each learner

Those four sections define what the test measures. Within each section, question difficulty adapts to the learner's answers:

  1. Start. The first section begins at the learner's current estimated level, θ, if they have one, otherwise at a random point near the middle of the scale (level 5 to 6), so first items vary between learners.
  2. Pick. The engine ranks unseen items by closeness to θ, shortlists the nearest 20 and picks one at random (see Keeping the bank secure below).
  3. Update. After each answer, θ moves in proportion to the residual: the observed score minus the model's predicted probability of a correct answer (an Elo-style update). Steps shrink as answers accumulate, and the standard error (SE) falls.
  4. Stop. A section ends when SE drops below a threshold after a minimum number of items, when the learner is clearly beyond the top or bottom of the bank, or at an item cap.

Each section's final θ supplies the starting estimate for the next, so a strong vocabulary result starts the reading section at a harder item. Selection and updates then continue with the new task.

Estimated level over the course of the test for two learners who start at the same level: one climbs to level 7, the other settles at level 4, and the confidence band around each narrows until the test stops after 18 and 17 items.

Figure 2. Two learners, one starting point. Within a few items their paths split; by the end they share almost no items, and each has a level with a confidence value.

For the reported result, reading and listening are estimated separately and combined by precision: the track measured more precisely weighs more. Teachers see both track levels and the combined level.

Building and validating the item bank

Choosing a question near a learner's estimated level depends on having suitable questions across the scale, each with a credible difficulty estimate. Because learners see different questions, that quality has to hold across the bank.

Writing thousands of items by hand across ten levels and several languages takes years; an LLM drafts them in hours, but each draft still needs validation. Quality control determines the bank's size: we keep about 10,000 items, a number we can still vet closely, and will expand as field data shows where more are needed. We evaluate generation with an explicit rubric, automated judges checked against expert labels, a versioned golden set, and monitoring once items are live.

Item pipeline. Top row: sources, LLM generation, exact gates, LLM judge, live bank, retired; items that pass the judge are served and calibrated, and items that are flagged are retired. Bottom row: a golden set sampled from the live bank and labeled by two providers goes to expert review, whose decisions tune the judge, its thresholds and the generation prompts.

Figure 3. The item pipeline. Every item passes cheap exact checks, then an LLM judge; a human-ratified golden set keeps the judge and the generator honest.

Choosing source material and estimating difficulty

Items start from authentic content: openly licensed text corpora, the material a native Spanish speaker actually reads. Every item keeps its source.

Before anyone answers it, each word-based item gets a prior difficulty from two independent signals: the word's frequency in a large corpus of film and TV subtitles, and its level in a national curriculum inventory (for Spanish, the Instituto Cervantes Plan Curricular). Within a level, rarer words start harder. When the two signals disagree, the word is re-leveled by frequency and flagged.

Generating and screening test items

The LLM writes only what the source cannot supply: pseudowords, distractors and passage titles. Deterministic checks run first, because they are cheap and exact:

  • Only content words are tested; function words and very short words are never blanked.
  • Passages below a quality threshold are dropped.
  • Any punctuation cleanup is compared with the original by embedding similarity, so no words are added or changed.

How AI checks item quality

Every item that survives the gates is scored by an LLM judge on five dimensions, each from 0 to 1:

  • Safety. Suitable for an all-ages classroom.
  • Language match. All learner-facing text is in the target language.
  • Pedagogical soundness. Fits its claimed CEFR level; exactly one defensible answer.
  • Linguistic quality. Correct, natural grammar and spelling; no broken fragments.
  • Construct validity. Tests what it claims: plausible distractors, a genuinely best answer.

Safety and language match are hard gates: a low score on either rejects the item outright. The scores map to a verdict of pass, review or reject. Only passing items are served; the rest are held back. Yes/No words have no stem or options, so they get a safety-only judge, which also catches pseudowords that happen to read as offensive in another language.

How experts keep the AI judge accurate

A judge is only useful if it agrees with experts. We keep a golden set: a sample stratified by task and level, labeled independently by judges from two model providers, then accepted or rejected by a language expert. The expert decisions are frozen as ground truth and versioned. Two providers matter because a model tends to favor text from its own family (Panickssery et al., 2024), and a panel of diverse judges is less biased than a single one (Verga et al., 2024). On the hard gates the panel takes the lower score, so either judge can veto.

The golden set does three jobs:

  • Measures the judge. Agreement with expert verdicts shows whether a new judge, prompt or threshold is an improvement or a regression, before it touches the bank.
  • Tunes the rubric. Dimension weights and verdict thresholds are set against expert decisions, not intuition.
  • Improves generation. Generation prompts are optimized against expert-accepted items only, never the judge's own scores, so the generator learns to satisfy experts rather than the judge.

We iterate the golden set. Each version adds the hard cases, items where the judges disagreed or the expert overturned them, so it keeps testing the judge where it is weakest (Shankar et al., 2024). The same expert decisions let us compare newer judge models before adopting them.

Refining item difficulty with learner responses

Expert review checks whether an item is suitable. Learner responses then show how difficult it is and how well it distinguishes proficiency levels.

Every test response feeds the same periodic IRT fit as the practice data in Part 1. Measured difficulty replaces the prior, and each item earns a discrimination score. Items that stop separating learners are flagged, then retired. Retired items are archived before they leave the serving pool, so every item ever served has an audit trail.

Keeping the bank secure

Calibration depends on responses that reflect language ability. If learners can memorize the bank, correct answers become less useful evidence. Three controls limit repeated exposure:

  • No repeats. An item is never shown twice in one test.
  • Randomized selection. Always picking the single nearest item made tests predictable: in simulation, 54% of learners at one ability saw the same Yes/No item. Picking at random among the 20 nearest cut peak exposure to about 20%, with no loss in accuracy.
  • Periodic refresh. We generate new items and rotate the bank on a regular schedule, and refresh the golden set with it, so a leaked item loses value quickly and the evaluation always tracks the live bank.

How to read a Roxxem level

Teachers see a learner's current level with a reading and listening breakdown, its direction over the term, and its confidence. A district sees levels that are comparable across classes and schools, because every item, from a game or a test, sits on one calibrated scale. The test and ongoing practice update the same estimate, and any level can be traced to the items behind it, and every item to the evidence that admitted it.

A level is an estimate, and should be read as one:

  • It is not a certification. It is aligned to ACTFL and CEFR, not an official rating under either.
  • Its confidence matters. A low-confidence level is provisional; it firms up as the learner practices.
  • It covers reading and listening only. Speaking and writing are not yet measured.
  • It is not yet externally validated. Agreement with established placement results is still to be shown.

Our roadmap addresses these: field studies with teachers and students in real classrooms, speaking and writing sections, and validation against external placement results.

A level matters only if it changes what happens next: what each student practices, what the class works on, what the teacher does on Monday. Our future post picks up there, starting with From Behavior to a Personalized Path: how Roxxem combines a level with what students and teachers do to build profiles, recommendations and, step by step, a path for each student.


How Roxxem Measures Proficiency

Further reading

Adaptive testing

  • Nydick, S. W., & Lockwood, J. R. (2024). An Overview of Duolingo English Test Administration and Scoring (DRR-24-03). https://go.duolingo.com/dettechnicalmanual
  • Way, W. D. (1998). Protecting the integrity of computerized testing item pools. Educational Measurement: Issues and Practice, 17(4), 17–27.

Vocabulary measurement

Evaluating LLM