Deep diveTechnology, Media & Telecommunications
How a streaming recommendation engine decides what each viewer sees next
A streaming recommendation engine narrows a catalogue of thousands of titles to a few rows per viewer in stages: eligibility filters for rights and age ratings, candidate generation, a ranking model and a re-ranking step for diversity and freshness. What it should optimize, how it handles new titles and how you prove it works are the hard parts, and in the EU and UK recommendations now carry transparency and profiling duties too.
On this page
- Recommender vocabulary for streaming product and data teams
- From a full catalogue to one personalized row
- Choosing what a streaming recommender should optimize
- Offline metrics, online tests and long-term holdouts compared
- Cold start for new titles and new subscribers
- Transparency and profiling rules that reach a recommender
- Feedback loops that quietly narrow a catalogue
- A hypothetical regional catalogue launch with sparse data
- Questions and answers
- Sources
Recommender vocabulary for streaming product and data teams
- Candidate generation
- A fast first stage that pulls a few hundred plausible titles from the whole catalogue, using collaborative filtering, embeddings, co-watch graphs or simple rules such as continue watching.
- Two-tower model
- A neural model with one tower encoding the viewer and context and another encoding the title. Their embeddings can be compared with nearest-neighbor search, which makes candidate retrieval cheap at catalogue scale.
- Ranking model
- A heavier model that scores each candidate with richer features (recency, device, time of day, past completion) to predict an outcome such as starting and finishing a title.
- Re-ranking
- A final pass that adjusts the ranked list for goals the ranking model does not see: diversity across genres, freshness, editorial priorities and avoiding near-duplicate rows.
- Implicit feedback
- Signals inferred from behavior, such as plays, completion, skips, rewinds and abandonment. Plentiful but noisy; explicit ratings are clearer but rare.
- Rights window
- The territories, dates and platforms where a title may be offered under its license. A recommender that ignores it promotes titles viewers cannot play.
- NDCG and recall@k
- Offline ranking metrics. Recall@k asks whether the titles a viewer later watched appeared in the top k; NDCG also rewards placing them higher.
- Popularity bias
- The tendency of models trained on past behavior to keep recommending what is already popular, starving the long tail of exposure and data.
From a full catalogue to one personalized row
- Eligibility filter
Removes titles outside the viewer's territory, license window, plan or parental settings.
- Candidate generation
Several retrievers each contribute a few hundred titles: embeddings, co-watch, trending, continue watching.
- Ranking model
Scores every candidate for the predicted outcome using viewer, title and context features.
- Re-ranking
Applies diversity, freshness and editorial rules and removes near duplicates.
- Page assembly
Chooses which rows appear, in what order, with which artwork.
- Feedback logging
Records what was shown, where, and what the viewer did, including what they ignored.
Choosing what a streaming recommender should optimize
The objective is a business decision disguised as a modeling choice. Optimizing for plays rewards clickable artwork and short titles; optimizing for watch time favors long series and autoplay; neither says whether a viewer was glad they watched. Most mature teams train on a blend of signals, such as completion, return within a week and explicit thumbs, and treat raw engagement as a guardrail rather than the target.
Catalogue goals matter too. A service that paid for a library wants it discovered, a news or sports service needs recency, and a service with expiring licenses may want to surface titles before they leave. Write these goals down as re-ranking rules with owners, so editorial and commercial input is visible instead of being hidden in feature weights.
Metadata quality sets the ceiling. Genre, cast, language, mood and maturity tags feed both cold start and diversity rules. Poor or inconsistent tags make every model worse, and cleaning them is often the cheapest improvement available.
Offline metrics, online tests and long-term holdouts compared
| Question | Offline replay metrics | Online A/B test | Long-term holdout |
|---|---|---|---|
| What it tells you | Whether a model ranks past choices well | Whether viewers behave differently now | Whether retention and satisfaction change over months |
| Speed | Hours | Weeks | A quarter or longer |
| Risk to viewers | None | Limited to the test cell | A small group gets a weaker or older experience |
| Main blind spot | Only sees titles the old system chose to show | Novelty effects and short-term proxies | Slow, and confounded by catalogue changes |
| Best use | Shortlisting models to test | Deciding which model ships | Checking the objective itself is right |
Offline metrics such as NDCG inherit the bias of the system that produced the logs, so a model that recommends genuinely new titles can score worse offline and better online.
Cold start for new titles and new subscribers
A new title has no viewing history, so collaborative signals cannot place it. Content embeddings fill the gap: encode synopsis, metadata, trailer frames or audio so a new title lands near similar ones, then reserve some exploration slots so it collects real feedback quickly. Without that exploration budget, popularity bias keeps new titles invisible indefinitely.
A new subscriber is the mirror problem. Use what is known at sign-up (territory, device, time, the title or campaign that brought them) and a short onboarding choice of genres or titles, then shift weight to behavior as it accumulates. Keep onboarding optional; a forced questionnaire costs more sign-ups than it saves in relevance.
Transparency and profiling rules that reach a recommender
Digital Services Act (Regulation (EU) 2022/2065): recommender provisions
European UnionApplies whenThe service is an online platform, such as a video-sharing or social service that disseminates user content to the public, and uses a recommender system1.
- Set out the main parameters of each recommender system in the terms and conditions, including the most significant criteria and why they matter, and any options users have to change them (Article 27)1.
- Where several options exist, let users select and change their preferred option directly from the part of the interface where content is prioritized (Article 27)1.
- Very large online platforms must offer at least one option for each recommender system that is not based on profiling (Article 38)1.
Directive 2002/58/EC (ePrivacy Directive), Article 5(3)
European Union, through national lawApplies whenThe service stores or reads identifiers on a viewer's device, for example for cross-device tracking or measurement2.
- Provide clear information and obtain consent unless the storage or access is strictly necessary for the service the viewer asked for2.
General Data Protection Regulation (EU) 2016/679 and UK GDPR
European Union and United KingdomApplies whenViewing histories and inferred preferences relate to an identifiable person3.
- Identify a lawful basis for profiling, tell viewers about it and honor objection and access rights3.
Age appropriate design: a code of practice for online services (ICO)
United KingdomApplies whenThe service is likely to be accessed by children4.
- Switch options that use profiling off by default, and allow profiling only with measures that protect children from harmful content4.
Feedback loops that quietly narrow a catalogue
Self-reinforcing popularity
Early signalShare of plays from the top of the catalogue rises every quarter while catalogue size grows.
MitigationReserve exploration slots and track exposure across the long tail as a reported metric.
Training only on what was shown
Early signalNew models keep agreeing with old ones and offline gains vanish online.
MitigationLog impressions with position and use propensity weighting or small randomized slots to correct for exposure.
Engagement proxies drifting from satisfaction
Early signalPlays rise while completion, ratings or retention flatten.
MitigationShip on satisfaction metrics and treat engagement as a guardrail.
A hypothetical regional catalogue launch with sparse data
Questions and answers
Does a streaming service need deep learning for recommendations?
Not at first. Matrix factorization, co-watch rules and good metadata go a long way for a catalogue of a few thousand titles. Deep models such as two-tower retrievers and neural rankers pay off when you have many signals, many contexts (devices, times, profiles) and enough traffic to test small improvements reliably.
How can a platform explain recommendations to viewers?
Two levels work together. At the system level, describe the main parameters in plain language, as Article 27 of the Digital Services Act requires of online platforms. At the item level, use simple reasons tied to real signals, such as because you watched a named title, and avoid explanations that the model did not actually use.
How do we test a new recommender without hurting engagement?
Start with offline replay to discard weak candidates, then run online tests on a small share of traffic with guardrail metrics and automatic stop rules. Interleaving, which mixes results from two rankers in one list, detects preference differences with less traffic than a split test and limits exposure to a worse model.
Must every platform offer a feed that is not personalized?
Under the Digital Services Act, the duty to offer at least one recommender option not based on profiling applies to very large online platforms and search engines. Other online platforms must disclose main parameters and any options they offer. In the UK, services likely to be accessed by children should switch profiling off by default under the ICO's code.
Sources
- Regulation (EU) 2022/2065 on a Single Market For Digital Services (Digital Services Act) — EUR-Lex · checked 10 October 2026
- Directive 2002/58/EC concerning the processing of personal data and the protection of privacy in the electronic communications sector — EUR-Lex · checked 10 October 2026
- Regulation (EU) 2016/679 (General Data Protection Regulation) — EUR-Lex · checked 10 October 2026
- Age appropriate design: a code of practice for online services, standard 12: Profiling — Information Commissioner's Office · checked 10 October 2026