How the Algorithm Works
Note: this write-up was drafted with Claude, working directly from the source code.
The recommender itself was designed and built by yours truly.
In one sentence: the site learns from everyone's ratings, predicts how you'd rate each title, drops whatever your filters exclude, and shows you the best matches first.
Your ratings are the starting point
Everyone's ratings go into one shared model, which looks for things people tend to like together. If people who loved Arrival and Blade Runner 2049 usually rate Ex Machina highly too, and you loved those two, Ex Machina rises for you.
Ratings are the only input. Clicks, viewing activity and watchlists aren't used.
Your taste profile updates as soon as you rate
Your ratings, from Hated to Loved, tell the model where you sit compared to everyone else. That's your taste profile. It turns on after 8 ratings and is recalculated every time you rate something, so your queue changes right away. The shared model itself retrains every three hours, which is when your ratings start to affect other people's predictions as well. ReelsGraph doesn't have enough ratings of its own yet, so for films the model is also trained on the MovieLens research dataset. MovieLens has no TV series.
Predictions lean on the sitewide average
Every title has a sitewide average, built from ratings here plus votes on TMDB, the public movie and TV database. Your prediction is a mix of that average and the model's personal guess, which comes from your taste profile. The more ratings a title had when the model was trained, the more weight the personal guess gets.
The average is cautious about small numbers, so a handful of fans can't lift a title above a widely loved classic.
Two personal adjustments are added on top. The first is your level: if you tend to rate generously or harshly, the average is moved to match. The second is a genre nudge: genres you consistently rate above or below your usual get a small bump up or down. Until you reach 8 ratings, every score you see is the sitewide average, and Discovery shows popular titles you probably know, so you have something to rate.
The prediction slider: your taste vs. the crowd
Power Users can change this balance with the slider labelled using K=... in the Discovery and playlist sidebars. K is roughly how much evidence the site wants before it trusts a personal guess over the sitewide average.
- Left, smaller K: trusts your taste more, even when a movie has few ratings. You'll probably find more unexpected favourites, and more misses as well. Your favourite genres also get more say.
- Middle: the default, tuned against real ratings. Leave it here if you're not sure.
- Right, larger K: leans more on the crowd and less on your own taste. Well-reviewed movies rise more easily, but the list gets less personal.
So a movie with few ratings that looks perfect for you will climb as you move left, and a widely loved one will tend to win as you move right. The slider only changes the order of Discovery and your playlist, and only once your taste profile is on. It doesn't touch your ratings or your filters.
The technical explanation lists exactly what each step of the slider changes.
Your filters decide what can appear
Discovery leaves out titles you've already rated, saved or dismissed, plus anything outside your settings: release age, popularity, language and streaming availability, and content preferences. These are hard cutoffs. A filtered-out title stays out no matter how well it scores.
Your queue shows the best matches first
Whatever is left gets ranked by prediction. There are a few ways to steer it:
A low rating teaches the model as much as a high one.
It's fine to be generous or harsh, but a history of nothing but 8s doesn't tell the model much.
They don't affect your taste profile, and you can undo them from the skipped page.
Too old, too obscure, not streaming where you are? A slider fixes that straight away, and rating more titles won't.
Reading the scores
On a title's page, Averages is the sitewide average in the form the rankings use. It's a bit more skeptical of titles with few ratings than the average your prediction starts from, so the two can differ slightly. The ML score is your full prediction on a 0 to 100 scale, where 67 is roughly Good, 78 Excellent and 89 Amazing. Rated is the model's personal guess by itself, before it's blended with the average. Despite the name, it isn't something you rated. For a title with few ratings the ML score and Rated can be far apart, because the blend is holding the personal guess back.
How we check it
A fixed set of real ratings is kept hidden from the model, which then has to predict them. As of August 2026 its guesses are off by about one point on the 0 to 9 scale, 20% closer than the sitewide average alone. Every change is also checked by actually reading the queues it produces, because an error score won't notice a queue that shows everyone the same obscure title.
The formulas the live system runs on, and why each one was chosen.
1. Signal and training data
The model is trained on explicit feedback only: integer ratings \( r \in \{0,\dots,9\} \), one per (user, item) pair. Implicit signals such as viewing activity, watchlists or clicks are not used. Before training, users with fewer than 8 ratings and items with fewer than 5 ratings are dropped.
ReelsGraph does not have enough ratings of its own yet to train the model, so films are bootstrapped from the MovieLens 32M research dataset. Those ratings are re-anchored so that each film's baseline matches its TMDB community score, once that score has been mapped onto our 0 to 9 scale by the provider curve in §4. MovieLens has no TV series, so a series is trained on native ratings where there are enough of them, and otherwise goes through the cold-start path in §5.
2. Matrix factorization
The model is plain matrix factorization with no bias terms, trained through ML.NET's wrapper around LIBMF. Each user \( u \) and item \( i \) gets a latent vector in \( \mathbb{R}^{48} \), and the predicted rating is their dot product:
Your predicted rating is how well your taste vector lines up with the title's vector.
The dimensions are learned from rating patterns and carry no labels such as genre. One might loosely correspond to crowd-pleaser versus arthouse, and another to nothing you could put a name on. The vectors are fitted by SGD on squared loss with L2 regularization:
Find the vectors that reproduce everyone's real ratings as closely as possible, while keeping them small so that a few ratings cannot push anything to an extreme.
The training settings, chosen against a held-out split of real ratings:
| Rank | Iterations | \( \lambda \) | Training rows |
|---|---|---|---|
| 48 | 30 | 0.075 | MovieLens 32M plus every native rating |
There is no bias term: no global mean and no user or item offsets, so how high a user rates in general is encoded in their vector. Two later decisions follow from that: the fold-in can't be a closed-form solve (§3), and the crowd prior needs a per-user shift (§5). Ratings are not weighted by age either, so an old rating counts as much as a new one.
The model retrains every 3 hours. After each run the serving snapshot is swapped atomically, every user is re-folded, and the user-similarity snapshot is reloaded.
3. Serving: AdaGrad fold-in
Between retrains the item factors \( Q \) are frozen and every user vector is computed by fold-in: initialize \( p_u = 0 \) and replay LIBMF's own SGD update over the user's rated items against the frozen \( Q \), using the trainer's learning rate (\( \eta = 0.1 \)), iteration count and \( \lambda \), plus AdaGrad's per-coordinate step scaling:
Replay your ratings to find your vector, taking smaller steps in any direction that has already moved a lot. Here \( g_f \) is the current gradient for coordinate \( f \), \( g_{t,f} \) the same gradient at an earlier step \( t \), and \( \varepsilon = 10^{-8} \) only keeps the division finite.
The fold-in runs synchronously on every rating, so your own queue reacts within seconds. The 3-hour retrain is what carries your ratings into \( Q \), and through it to other users. Fold-in is also the only serving path: trained and folded vectors agree to within a few tenths, so everyone is served a folded vector and there is one code path to maintain instead of two.
The AdaGrad term is what makes the result independent of the order in which the ratings are replayed. At a constant learning rate the per-item steps are near-critical ( \( \eta \lVert q_i \rVert^2 \approx 0.5 \) ), so the vector ends up mostly fitting whichever ratings the loop visited last.
Below 8 ratings no vector is folded and predictions fall back to the crowd prior.
4. The crowd prior: calibrated pooled Bayesian average
Every item has a Bayesian average built from two rating populations: our own users and the TMDB community. The internal 0 to 9 axis is our users' scale. Provider scores \( v \) are mapped onto it with a power curve, and our own ratings are left as they are:
TMDB scores are bent down to match how people rate here: an 8.0 on TMDB counts as about 6.9 for a film and 6.7 for a series on our 0 to 9 scale.
The movie curve is fitted on the titles MovieLens and TMDB have in common. MovieLens and our own users both put the average title at around 5.2 to 5.4 on the 0 to 9 scale, and TMDB sits higher. The curve corrects for that level and leaves title order and vote counts alone. TV gets a steeper exponent because the movie curve over-credited series. It was fitted by comparing how the same people rate films and series.
The pooled average is then shrunk toward a global base \( b \), which counts as \( S \) pseudo-votes:
Average every rating here and every translated provider vote, plus \( S \) imaginary votes at the base, so a title needs real votes to move far from it.
Here \( S = 50 \), and the base \( b \) is derived from the vote-weighted global mean \( \mu \) by stretching its distance from the top of the scale:
The base sits a little below the average vote, near the typical title rather than the typical popular one.
The stretch is needed because \( \mu \) is vote-weighted: it describes the typical vote, and most votes are cast on popular, well-liked titles. That is a mismatch between a weighted and an unweighted mean, not a selection-bias correction. The measured weighted and unweighted averages were 5.81 and 5.41 over about 445,000 titles. A prior for an item with no votes should sit near the typical item, so the base is pulled toward the lower figure.
Further safeguards:
- Leaderboards use a stricter \( S \). Ranked pages recompute the prior as \( \bar r_i^{\text{rank}} \) with \( S = 250 \), because provider averages tend to be highest where vote counts are lowest. Before the change, a film with 143 TMDB votes outranked The Shawshank Redemption with more than 31,000. The cost is that a few real classics with thin vote counts drop some places. Most Hated keeps the smaller \( S \) on purpose: shrinking toward the mean makes a title look less hated, which would pull out the very titles that list is for.
- New releases are trusted gradually. Compared with critic scores, TMDB's vote average starts about 1.2 points too high and settles with a time constant of roughly 4.5 months. So a new title's provider votes count at 0.4 of their full weight at release, ramping up to 1.0 over 9 months. The floor is there so that a well-known newcomer doesn't collapse all the way onto the global base.
5. Prediction: evidence-gated shrinkage plus a user-scale shift
The served prediction is a blend of the MF score and the prior, weighted by how much training evidence the item has:
The more ratings a title was trained on, the more the model's guess counts. The rest comes from the crowd score, moved to your level.
- \( m \) is the number of ratings the item's position was actually trained on: native and MovieLens ratings. TMDB votes are left out on purpose, because the factorization never sees them. A title can have thousands of provider votes and still be shrunk hard if its trained position rests on a handful of ratings.
- The default is \( K = 25 \), chosen so that a typical title still leans on the crowd prior and a well-rated one leans on the model.
\( K \) is also one of the three settings Power Users can change from Discovery's sidebar. See the prediction slider in §7.
\( \delta_u \) is the user shift, and it exists because of the missing bias term. MF scores are on the user's personal scale (someone who rates everything 9 gets MF scores near 9), but the prior is on the crowd's scale. Without a correction, items with little evidence sank for generous raters and rose for harsh ones. The shift is applied to the prior only:
How much more generous or harsh you are than the typical rater, trusted more as your rating count grows.
Here \( \text{offset}_u \) is the user's mean residual against the Bayesian averages of the titles they rated, taken relative to the typical rater's. \( K_g = 5 \), the same constant that shrinks the genre tags in §6, stops a short rating history from producing an extreme shift. \( w = 1.25 \) sets how much of the measured difference is carried into predictions.
Since the shift is measured relative to the typical rater, most users move by less than a point and only the extremes move noticeably. \( w \) is not set higher because some of what looks like generosity may just be a user picking titles well, and that doesn't carry over to titles they didn't pick.
The no-position path. Fresh films and series without enough native ratings may have no trained position at all. Serving their raw Bayesian average would flood every queue with them, because popular newcomers have generous priors, so the average is faded toward the global base:
With no trained position, start from the base, go only part of the way toward the crowd score, then add the personal terms in full.
A fade of 0.5 halves how far such a title can stand out from the typical one, which stops the same popular newcomer from topping every account's queue. The personal terms are added after the fade, so they are not halved. The fade is a step and not a ramp: it applies only to titles with no trained position at all, and a title just past the training floor is blended by \( \alpha \) against the unfaded prior.
6. The genre-affinity correction
MF picks up genre taste only where it has evidence, so a separate layer measures the user's residual preference per tag and nudges predictions with it. For each eligible tag, the user's mean residual against the items' Bayesian averages is shrunk by \( n/(n+K_g) \), with the same \( K_g = 5 \) as the user shift in §5. Without the shrinkage, one bad rating of a rare tag would become that tag's whole estimate and be applied to every item carrying it. For each item, the deltas of its eligible tags are averaged into \( \Delta_{ui} \), capped at ±1, and then applied with a weight that depends on evidence:
Add your average leaning across the title's tags, capped, then weighted more heavily the less the model knows the title.
The weight can exceed 1, so the cap limits \( \Delta_{ui} \) and not the final nudge. With the values below, a capped delta moves a prediction by up to ±1.25 points on the thinnest trained title and ±1.25 on one with no position. The cap itself has not been fitted yet.
The fitted weights show how much genre signal the factor model leaves behind:
| \( w_{\text{floor}} \) | \( w_{\text{cold}} \) | \( w_{\text{no-position}} \) |
|---|---|---|
| 0 | 1.25 | 1.25 |
For heavily rated films backed by MovieLens, the factor model already captures all the genre signal we can measure, so the floor is zero. The correction only comes in as training evidence gets thinner, and titles with no trained position have a weight of their own.
Movies and series share one set of tags. TMDB's few TV-only compound genres are mapped to their film equivalents where possible, so that a series and a film count toward the same preference instead of two unrelated ones.
7. Discovery: filters, then rank
The Discovery queue is built in a single pass over the in-memory catalog. About a dozen eligibility checks run first: collection, watchlist and dismissed exclusions, language, streaming availability and content preferences, plus the year and popularity sliders. Those sliders are hard cutoffs and not weights. Each notch maps to a threshold (the top year notch means released within the last 2 years, the middle one within 40), so anything a slider already filters on is not scored a second time. Whatever passes is ranked by the prediction \( \hat r \) from §5 and §6, using a streaming top-k selection instead of a full sort.
A few queues work differently:
- Accounts with fewer than 8 ratings are ranked on popularity alone. Predictions would just be the same crowd score for everyone, and what matters for the first few cards is that the user can actually rate them: a famous disappointment is more useful there than an acclaimed title nobody has heard of.
- Upcoming titles have no ratings at all, so they are ranked on hype, meaning the provider's anticipation signal (TMDB popularity), multiplied by \( 1 + 0.35 \, \Delta_{ui} \) to tilt it toward the viewer's genres.
The prediction slider. Power Users get one more slider at the bottom of the Discovery sidebar. It decides who wins when your taste and the crowd disagree about a title: moving it left favours your personal signals, moving it right favours the crowd. It also orders your playlist, and in Party Mode each member's own setting applies to their share. Each step changes three parameters together, so that they stay consistent with each other:
- Shrinkage \( K \) from §5. A lower \( K \) lets a prediction built on a handful of ratings stand on its own. That is where the surprising picks come from, and the misfires too.
- The no-position fade from §5. A lower fade flattens the crowd score of titles the model can't place, so they end up ordered by genre fit alone. A higher one lets well-reviewed unknowns rise, at which point every account's queue starts with the same few.
- The genre weights from §6, all scaled by a single multiplier so the calibrated ratio between them is kept. A lower multiplier means your genre leanings count for less.
| Setting | \( K \) | Fade | Genre weights |
|---|---|---|---|
| Wild, no safety net | 0 | 0.2 | ×1.4 |
| Bold, leans on your genres | 5 | 0.3 | ×1.25 |
| Slightly bolder than default | 12 | 0.4 | ×1.1 |
| Calibrated for everyone | 25 | 0.5 | ×1 |
| Slightly safer than default | 60 | 0.6 | ×0.85 |
| Safe, just follows the crowd | 150 | 0.7 | ×0.7 |
| Boring, ML almost off | 500 | 0.8 | ×0.5 |
The middle row, in bold, is the calibrated default that everyone else gets, and it tracks the constants whenever they are refitted. The rest of the pipeline, user shift included, is identical at every step. The slider only reorders the released queue, and only for accounts past the 8-rating mark. Newer accounts, the upcoming queue, backlog sorting and the numbers on a title page are not affected. It works well together with the popularity slider: the Wild setting plus a low popularity floor is how you go looking for obscure gems.
8. The numbers on a title page
- Averages is the leaderboard version of the crowd prior, \( \bar r_i^{\text{rank}} \) at \( S = 250 \), so it matches the rankings it appears in. It is not the \( \bar r_i \) the prediction starts from, which is why the two can differ slightly.
- Rated is the raw factorization score \( p_u^\top q_i \), clamped to 0 to 9, with no shrinkage, shift or genre term. It is what people with your taste would give the title, regardless of how much evidence that estimate rests on.
- ML score is \( 100 \cdot \hat r / 9 \), the full served prediction. For a title the model knows well, \( \alpha \to 1 \) and it tracks Rated. For one it doesn't, it stays close to the crowd prior, which is why the two can disagree.
9. Social taste compatibility
Compatibility between two users is the mean-centered cosine of their ratings on the titles they have both rated, mapped to [0,1] and then shrunk by overlap as \( \text{sim} \cdot \frac{o}{o+5} \). With fewer than 5 co-ratings there is no score. A user's closest matches are found on request with a single scan over everyone else's ratings, which are held in memory and reloaded after each training run (§2). That avoids maintaining a table of every pair, which would grow as \( O(n^2) \).
10. How the constants are chosen
Most of the numbers above were measured, not picked by hand. A few are still judgment calls that haven't been fitted yet: the genre cap, the cold fade, the leaderboard prior and the hype tilt. The rest are tuned against a held-out check: a fixed slice of real ratings is hidden from the model, predicted, and scored by mean absolute error. As of August 2026 the full pipeline is off by about 1.09 rating points on average, against 1.36 for the crowd prior alone.
The pipeline is a chain with no feedback loop, so the constants can be refitted in one pass, in order:
The trainer's settings are fitted separately, since they depend on ratings and not on priors. There is one dependency between the two: the imported films are re-anchored through the provider curve (§1), so the curve has to be fitted before the trainer. Genre weights always come last, because they are fitted on whatever error the rest of the pipeline leaves.
The check has one big blind spot: it can only grade ratings people chose to give. A change that shows every account the same obscure title can still score well, because a title almost nobody has rated barely appears in the hidden set. So every change is also checked by reading the queues it produces, and if the queues look wrong the change is dropped, whatever the number says.
What was tried and didn't work
- Discounting old ratings. In practice it acted like extra regularization with an age bias. Simply raising \( \lambda \) scored better, and the decay penalized people who really do prefer older titles.
- A closed-form fold-in. On paper the exact ridge solve \( p_u = (Q_u^\top Q_u + \lambda I)^{-1} Q_u^\top r_u \) is more correct than SGD, but with no bias term it latches onto the popularity axis and predicts 9 for every popular title.
- A fold-in without AdaGrad. Scores on a 341-rating account swung by up to 10 points between reloads, for no reason other than the database returning the ratings in a different order.
- A weaker prior and no cold fade. The held-out check preferred both. Each of them put the same few untrained titles at the top of almost every queue, so both were reverted.