Roxxem Logo
Engineering,  Research

The Engine Behind Personalized Language Learning

Jingjing Ren, PhD

Students in one class differ in level, interests and goals, and a teacher cannot choose the next piece of practice for each of thirty. Software can, if it knows each student well and aims at the right target. Most recommenders optimize for engagement; ours is held to whether students learn.

This post covers the engine that makes it possible, the Pulse intelligence engine. It has two parts: profiles of every student, teacher and class, built from what they do, and a recommender that turns what the profiles reveal into action: what each student should learn next and what each teacher may teach. Both build on how we measure each student's level and place new students.

Profiles: what behavior says about a person

By behavior we mean what people do in Roxxem, logged as events: a student's graded answers, plays, replays, completions and saves; a teacher's classes, assignments and lessons. Explicit input is what people tell us: student goals, teacher notes, declared level, onboarding topics. The engine reads both and keeps them apart, because what someone says and what they do can disagree, and both are worth knowing.

The engine keeps one set of profiles, and every feature reads it. No feature computes a level or a taste of its own. When Pulse, a shelf (a titled row of recommended content in the app) and Roxxy describe a student, they describe the same student.

Students, teachers and classes are first-class entities in a daily feature store; a class profile aggregates its students' profiles alongside its teacher's.

Data flow: behavior (graded responses, plays, saves, assignments) and explicit input (student goals, teacher notes, declared level, onboarding topics) feed the Pulse intelligence engine, which runs IRT calibration and a daily profile recompute. It writes three profiles: ability (θ + confidence per skill), interests and teaching. Pulse insights, content shelves (ranked by the recommender) and Roxxy each read all three, and the activity they lead to becomes new behavior.

Figure 1. Behavior in, profiles out. Every feature reads the same profiles, and the practice they lead to feeds the next update.

Ability: progress, with its uncertainty

How well can a student read and listen, and are they improving? The ability profile holds a level and its confidence per language and skill, estimated from graded answers in games/tasks, and the adaptive test with Item Response Theory (how it works).

  • Unmeasured is a state, not a zero. Without enough graded answers, a student shows as warming up and stays out of class statistics, rather than reading as a beginner.
  • Only measured skills get a number. Today that is listening and reading; speaking and writing show as not yet measured.

Interests: taste, weighted by depth

What does a student choose to spend time on? Interests come from recent activity, so they track what a student likes now, not last year. Each signal is weighted by depth: a long study session counts more than a single play, and a play more than an open. The profile keeps top topics, creators and grammar points per language.

A new student has no history, so onboarding asks for at least three topics, and shelves use them at once. They step aside as behavior accumulates. More generally, the profile keeps what a student says (declared level, goals, onboarding topics) next to what they do, and does not merge the two blindly. The level that ranking uses is the measured one when it exists and the declared one otherwise, and the profile records which.

Teaching and class: beyond the student

What does a teacher teach, and to whom? Teachers get a profile of their own, per language taught. Today it records what they teach and to whom: their classes and students, the grade and proficiency levels of those classes, how much they assign and how recently. Next, it learns how they teach from what they assign and build: whether a teacher leans on music or news, stretch or review, games or full lessons.

The teaching profile drives the teacher's own shelves ("For teaching…", "Popular with classes like yours"). As it learns a teacher's approach, class suggestions in Pulse will fit the teacher as well as the class.

Pipeline: rebuilt daily from activity

  • Role follows activity. Studying creates a student profile; running a class creates a teaching profile. A teacher who also studies gets both.
  • Rebuilt daily from scratch. Behavioral profiles are recomputed every day, so there is no incremental state to drift. Explicit input, such as goals, is read live.
  • Daily snapshots. The ranker trains on each profile as it was when a card was shown, so the future does not leak into training.

Recommender: personalization that aims at learning

The recommender is where learning becomes the goal. Engagement still counts, since a student who stops opening the app cannot learn, but it serves growth, not the other way round.

Every shelf, for students and teachers, runs on the standard two-stage design for recommenders: candidate generation narrows the library to a short list of plausible items, and ranking orders that list for one person. Four sources propose candidates, ranking balances engagement with learning, and each shelf says why it shows what it shows.

Two-stage recommender: the library feeds candidate generation (content similarity, collaborative filtering, topic tags, trending), whose sources are merged and then ranked (a learned ranker predicting graded engagement, level fit, and a fixed formula blending the two, tuned per shelf) into shelves that state their reasons. Impressions and graded engagement flow back as the ranker's training labels.

Figure 2. The recommender. Multiple sources propose candidates; ranking blends graded engagement with level fit, and learns from how students engaged.

Candidates: four sources, merged

  • Content similarity. Content is embedded with a multilingual sentence-embedding model (Reimers & Gurevych, 2020), so items match by meaning across languages. A student's history pulls in its nearest neighbors from a vector index.
  • Collaborative filtering. Students who engaged with the same videos tend to engage with the same next ones. We fit implicit-feedback matrix factorization, with confidence built from watch time, progress, games played and saves, decayed so that last month counts more than last year.
  • Topic tags. Curators tag content by topic, genre and grammar point. A student's top topics in the interest profile pull in matching items.
  • Trending. Usage trends across the platform, computed separately for students and teachers: what students are playing and what teachers are assigning right now.

Similarity and filtering run as nightly batch jobs and cover our content library: videos, lessons, playlists, vocabulary lists and question sets.

Ranking: engagement and learning, balanced

One ranking stack serves every content shelf, from the home page and For You feed to teacher shelves and learning paths; each shelf tunes how much each part counts.

Not all engagement is equal. Every card a student sees is logged with its shelf and position. What follows is graded: studying or playing a game counts more than watching, and watching more than a click. A click says a title caught an eye; a finished game says the student learned with it. Cards shown and never opened are negatives, and logging position lets the ranker correct for top cards being clicked whatever their relevance (Joachims et al., 2017). A learned ranker trains on these graded signals to predict what a student will engage with deeply, and is tested on a time split: trained on the past, evaluated on the future.

A hand-tuned scoring function balances engagement and growth. Today each card's score combines that prediction with level fit: how close the content's difficulty is to the student's level. The target sits just above the student's level, by a tunable margin of about 15%: input they mostly understand, with a little that stretches them (Krashen, 1985; Vygotsky, 1978). Where enough students have played an item, its difficulty is measured on the same scale as student ability; otherwise its editorial rating stands in. The less sure we are of a student's level, the easier it aims.

Shelves: reasons that help people choose

Ranking sets the order; the shelf says why. As with the rows on a Netflix home page (Gomez-Uribe & Hunt, 2015), every shelf states its reason in plain words, such as "More reggaeton for you" or "Because you recently studied…", and teachers see the same on theirs ("Popular with classes like yours"). A reason helps people choose: a student can pick the shelf that matches what they want right now and skip one that misreads them, and a teacher can judge a suggestion before assigning it. Reasons come from topic tags and history, facts a person can check, not from an obscure model score. Shelves also interleave content types, so music, the largest category, does not crowd out TV, news and podcasts.

Path: what comes next

The goal is a path toward a goal the student sets, judged by whether they progress. Three steps lead there.

  1. Learning as the top metric. Today engagement judges each ranking change. Next, the judge is whether a student's level climbs over months, and that signal sets the scoring function's weights. Optimizing for long-term learning, not the next click, is an active research field (Doroudi et al., 2019), and one we are working in.
  2. Learner agency. Learners are agents who set goals and steer their own learning, not passive recipients of it (Bandura, 2001). Students set goals and answer three short check-ins: what they want this week, what they credit after a level-up or streak, and how an activity went. Today they feed the profile as explicit input. Next, they carry weight at every step, from which candidates are found to how they are ranked and explained: a student who writes "I have a trip to Madrid in March" should see it in what comes next.
  3. Better content, curated on evidence. Learning gain shows which content actually helps students learn and teachers teach. That evidence guides what we add to the library, and finds the most effective content for a school's curriculum or a textbook unit.

What does not change: Roxxem proposes, the teacher decides, and every claim about a student stays one the teacher can inspect and overrule. The aim is better student outcomes, and teachers freed for the high-value moments with students that no recommender can replace.

Getting the most from it

The engine learns from what students and teachers do, so a few habits make it work better for a class:

  • Assign games/quizzes, not only videos. Graded answers are what measure a level. A class that only watches stays "warming up".
  • Read a level with its confidence. A new or low-confidence level is provisional; lean on it lightly until it firms up.
  • Ask students to set a goal. Goals and check-ins sit next to measured ability and are next in line to shape what each student sees.
  • Use the shelf's reason. Every suggestion says why it is there. If the reason misreads your class, skip it.

The engine does the sorting, so a teacher's time goes to the students who need it most.

Further reading

Recommender systems

Learning science

  • Bandura, A. (2001). Social cognitive theory: An agentic perspective. Annual Review of Psychology, 52, 1–26. https://doi.org/10.1146/annurev.psych.52.1.1
  • Doroudi, S., Aleven, V., & Brunskill, E. (2019). Where's the reward? A review of reinforcement learning for instructional sequencing. International Journal of Artificial Intelligence in Education, 29(4), 568–620. https://doi.org/10.1007/s40593-019-00187-x
  • Krashen, S. D. (1985). The Input Hypothesis: Issues and Implications. Longman.
  • Vygotsky, L. S. (1978). Mind in Society: The Development of Higher Psychological Processes. Harvard University Press.
Estimated level over the course of the test for two learners who start at the same level: one climbs to level 7, the other settles at level 4, and the confidence band around each narrows until the test stops after 18 and 17 items.
Research,  Engineering

See how Roxxem's adaptive proficiency test places learners in about 10 minutes, and how we generate and validate its 10k-item bank with LLM evals.