A spaced-repetition app makes one promise: show me this card again just before I'd forget it. If the scheduler is subtly wrong, nobody notices for weeks. Cards come back a bit early or a bit late, reviews pile up or memories fade, and by the time it shows, months of history were scheduled by the wrong maths. So we treat Chamber's scheduler as the part of the app that most needs evidence.
This is a companion to the Chamber architecture post.
The model in one paragraph
FSRS (the Free Spaced Repetition Scheduler) describes each card with three numbers. Difficulty D runs from 1 to 10. Stability S is the number of days until your chance of recalling the card falls to 90%. Retrievability R is the chance you'd recall it right now. In FSRS-6, R follows a power-law forgetting curve, R(t, S) = (1 + f·t/S)^(−w₂₀), where the factor f is chosen so that R(S, S) = 0.9 exactly. Every review updates D and S using 21 weights, and the next interval is simply the time until R drops to your target retention.
applyReview); the maths is pure and deterministic. Tap to zoom.All of this lives in lib/scheduler/fsrs.ts, about 140 lines of pure functions with no I/O, no clock and no randomness. The state machine that wraps it (learning steps, graduation, lapses, leeches) lives separately in scheduler.ts. Purity matters here because Chamber is local-first: the same review log must produce the same schedule on every device and on every replay.
A worked example
With the canonical FSRS-6 default weights and 90% target retention, here is one card's life, step by step. These exact numbers are what the test suite asserts:
| Event | R at review | Stability after | Difficulty after |
|---|---|---|---|
| First review, “Good” | — | 3.26 days | 4.88 |
| 10 days later, “Good” | 0.796 | 21.79 days | 4.87 |
| 20 days later, “Again” (a lapse) | 0.906 | 3.18 days | 7.23 |
| An hour later, “Good” | — | 3.46 days | — |
Two things are worth noticing. A successful review at lower retrievability earns a bigger stability boost: you remembered something you were close to forgetting. And a lapse resets stability well below its previous value, but the formula is clamped so that forgetting can never raise stability.
Why “it says FSRS-6” isn't enough
Our spec has a line we take seriously: do not claim compatibility based only on naming. A scheduler can carry the right name, use the right variable names and still compute something different. Testing an implementation against values it produced itself proves nothing beyond "it hasn't changed".
So the formulas were implemented twice: once in Chamber's TypeScript, and once independently in Python from the published equations. The Python version generated golden values at the canonical default weights, and those values are committed as fixtures in fsrsReference.test.ts.
The suite checks values to four decimal places (the reference computed six, and four absorbs floating-point ordering differences). It tests at two levels on purpose:
- The pure functions: initial stability and difficulty, the forgetting curve, recall and lapse updates, the same-day update, and the interval for a given retention.
- The shipped state machine:
applyReviewis driven through the same review sequence on a real clock, so the parity claim binds the code the app actually runs, not just the helpers it calls.
One detail shows the care the tests need. The interval test probes at 80% retention, not 90%. At exactly 0.9 the forgetting-curve factor cancels and the interval equals S for any decay weight, so a test there would pass even if the decay weight were wrong.
Three bugs it caught
This wasn't theoretical. Bringing in the reference suite exposed three real defects, all since fixed:
- The “default” weights weren't the default. The constant labelled as the FSRS-6 defaults was actually one person's optimised vector. The giveaway was several weights sitting exactly on the optimiser's bounds, the fingerprint of an overfitted result rather than an averaged default. It differed from every reference implementation at every index. The fix replaced it with the canonical vector. Because weight vectors are versioned (
paramsVersion), every past interval can still be traced to the exact weights that produced it, and stored review states weren't rewritten. - Update order. In FSRS-6, the new stability is computed from the difficulty before this review, and difficulty then updates in parallel. One code path updated difficulty first and fed the new value into the stability formula. The scheduler and the optimiser now share the canonical order, so the optimiser fits the same model the app schedules with.
- The same-day floor. For reviews less than a day apart, FSRS-6 floors the stability multiplier at 1 for a correct answer. Without the floor, a same-day “Good” on a very stable card could lower its stability. The test pins both sides: “Good” at S = 100 stays at 100, and “Again” still drops it.
The vectors also guard the edges: a weight vector with the wrong length or a NaN in it throws immediately, instead of silently turning every schedule into NaN.
A personal optimiser, kept simple
The default weights describe an average memory. After enough reviews, Chamber can fit weights to yours.
The optimiser in optimizer.ts is deliberately small. It groups the append-only review log into per-card histories and replays each one under a candidate weight vector. It scores the vector by the mean log-loss of predicted recall against what actually happened. Only reviews at least a day apart count toward the loss, because same-day repeats say little about long-term memory, but every review still updates the simulated state.
The search is coordinate descent: nudge one weight at a time by ×0.7, ×0.85, ×1.15 and ×1.4, clamp it to the same per-weight bounds the maintained fsrs-rs optimiser uses, keep any improvement, and sweep three times. The bounds matter. With one uniform range for every weight, the search wandered into regions the reference forbids (a decay weight outside 0.1–0.8 distorts the whole curve), and the result couldn't be reproduced elsewhere.
It's gated at 1,000 reviews, because fitting 21 parameters to a sparse history mostly fits noise. And it's deterministic: the same log produces the same weights on every device, which matters when the log syncs.
Honest limits: this optimiser isn't the official gradient-based FSRS optimiser, and we don't claim it matches its output. It's a simpler search over the same model and the same bounds. Parity is pinned at the default weights by committed fixtures. Cross-checking against a live reference library is an occasional offline step, not part of CI.
The log makes it possible
None of this works without trustworthy history, which is why reviewLogs has been append-only since day one. Undoing an answer doesn't delete it. It records a withdrawal and restores the state the log captured before that answer, and only if the card hasn't changed since; otherwise it refuses rather than guessing. Every reader of the log filters out withdrawn answers, and a test sweep checks that each one does. The optimiser trains on exactly what you did, minus exactly what you took back.
See also: Chamber · Chamber's architecture · How the arcade fits together