How Does Netflix Use Artificial Intelligence? Four Layers, One Lesson
How does Netflix use artificial intelligence? In four layers: a recommendation foundation model decides what each member sees, bandit algorithms pick which artwork represents each title, machine learning grades and shapes video quality during encoding, and generative AI now works inside production, search, and ads. Every layer is tuned against one thing: whether members keep watching.
Most write-ups on this topic recycle the same 2016 statistics from secondhand sources. I went to Netflix's own tech blog, its research papers, and its latest SEC filing instead, because the system has changed a lot in the last eighteen months. Then I pulled out the part that matters for anyone building a product: how they decide where AI earns its place.
TL;DR
- One model, not hundreds. In March 2025, Netflix described a recommendation foundation model trained like an LLM on hundreds of billions of interactions from over 300 million users.
- Recommendations are worth a measurable chunk of viewing. A November 2025 Netflix paper estimates engagement would fall about 4% with a basic matrix factorization recommender and about 12% with popularity ranking.
- Thumbnails are personalized. Netflix's artwork system picks an image per member per title, at a peak of over 20 million requests per second.
- Encoding is quality-scored by ML. AV1 now carries about 30% of Netflix viewing and uses one-third less bandwidth than AVC or HEVC, with 45% fewer buffering interruptions, per its December 2025 post.
- GenAI is in production. Netflix's Q2 2026 shareholder letter says GenAI workflows were used in roughly 300 titles in 2026.
The map: where AI sits in the Netflix stack
Netflix does not have "an AI feature." It has a stack where every layer answers a different question about a member's evening: what should they see, how should it look, how should it arrive, and how was it made. Search sits across the top as the newest entry point.
Let me take them in order of how much they matter.
Layer 1: the recommendation foundation model
For years Netflix ran a fleet of specialized models, one for Continue Watching, one for Today's Top Picks for You, and many more. In its March 21, 2025 post, the personalization team explained why that stopped scaling: maintenance was costly, and an improvement in one model rarely transferred to the others because each was trained independently on the same underlying data.
Their fix borrowed directly from large language models. Treat each member's history as a sequence, tokenize it, and train one big model to predict the next interaction. A few details are worth stealing:
- Tokenize behavior, not clicks. Raw actions get merged into meaningful events, the way byte pair encoding merges characters. A five-minute trailer play and a two-hour film are not equal tokens, so they weight targets differently.
- Latency forces short context. Serving needs millisecond responses, so inference context runs to hundreds of events, not the thousands a heavy viewer generates. Sparse attention and sliding-window sampling during training let the model still learn from full histories.
- Cold start is designed in. A brand-new title has no viewing data, so each title embedding blends a learned ID embedding with a metadata embedding (genre, storyline, tone), and the mix leans on metadata when a title is young.
- One model, many consumers. Downstream teams use it directly, fine-tune it for their surface, or just pull the embeddings it produces for members and titles.
Netflix also reported that scaling laws hold: more data and more parameters kept improving results. That is the same bet the LLM labs made, applied to a home page.
What recommendations are actually worth
The famous claim that most Netflix viewing comes from recommendations is a decade old, and the original paper's own abstract is mostly about method: the 2015 ACM paper by Carlos Gomez-Uribe and Neil Hunt describes A/B tests focused on member retention and treats search as another recommendation problem. The newer number is better.
In The Value of Personalized Recommendations: Evidence from Netflix (submitted November 10, 2025), Netflix researchers modeled real viewing data and ran counterfactuals. Swap the current system for plain matrix factorization and engagement drops about 4%. Swap it for a popularity list and it drops about 12%.
Do the arithmetic on that. The gap between "show everyone the top 10" and a modern recommender is roughly 12 points of engagement. The gap between a competent classic recommender and Netflix's current one is about 4 points. Most of the value comes from doing personalization at all, and the expensive frontier work buys the last few points. At Netflix's scale those points are worth the research budget. At yours, the first 8 are the ones to chase.
The paper adds a second finding I find more useful: the biggest gains land on mid-popularity titles. Blockbusters get watched anyway and niche titles have small audiences. Personalization pays most on the good item nobody would otherwise find.
Layer 2: artwork that changes per member
Picking the title is half the job. Netflix also picks the picture. In its December 2017 artwork post, the team described moving from finding one best image per title (via multi-armed bandits) to choosing the best image per member using contextual bandits.
Their own examples: a romance viewer sees Good Will Hunting with Matt Damon and Minnie Driver, a comedy viewer sees Robin Williams. A Uma Thurman fan gets her on the Pulp Fiction tile; a John Travolta fan gets him.
The hard parts were not the model:
- Attribution. A member can only react to the one image shown, so the system has to separate "this image caused the play" from "they would have played it anyway."
- Honesty. The art pool must be engaging and representative, or the optimizer learns clickbait.
- Scale. The service handled a peak of over 20 million requests per second with low latency.
Layer 3: machine learning that grades every stream
The least glamorous layer saves the most money. Netflix built VMAF, an Emmy-winning perceptual video quality metric that fuses several measurements through a trained model to predict how a human would rate a frame. It is open source, ships a library to train custom models, and got a new generation of v1 models in June 2026.
A quality score a machine can compute is what lets an encoder optimize for what viewers actually see instead of raw bitrate. The payoff shows in Netflix's December 1, 2025 AV1 post:
- AV1 carries about 30% of all Netflix viewing, its second most-used codec.
- AV1 sessions score 4.3 VMAF points above AVC and 0.9 above HEVC.
- They use one-third less bandwidth than both, with 45% fewer buffering interruptions.
- Film grain synthesis, productized in July 2025, strips grain before encoding and rebuilds it on the device.
Quick math: if a hypothetical title streams at 6 GB per hour on AVC, one-third less is 4 GB. Multiply by a catalog's worth of hours and that is the CDN bill AI quality scoring helps shrink.
Layer 4: generative AI in production, search, and ads
This is the newest layer, and Netflix now discloses it to investors. Its Q2 2026 shareholder letter, filed in July 2026, says:
- GenAI workflows were used in roughly 300 Netflix titles in 2026, concentrated in post-production, and some productions would otherwise have dropped key shots.
- LLMs are being used to improve title discovery, with voice search and AI-powered natural language search for members.
- AI-powered tools now span the full ad lifecycle, from planning and creative production to optimization and reporting, with ads revenue on track for about $3 billion in 2026.
The pattern holds: generative AI went where a cost or a constraint already hurt (expensive shots, keyword search that cannot handle "something funny and short"), not where it made a good demo.
The operator lesson
Strip out the scale and Netflix's approach is a playbook any team can run:
- Pick one north-star metric. Netflix tests against retention and engagement, not model accuracy. Decide what your AI is supposed to move before you build it. My guide to measuring AI ROI walks through setting that baseline.
- Consolidate before you specialize. One shared model with clean data beat a fleet of bespoke ones. For most businesses that means one well-maintained context layer that every AI workflow reads from.
- Personalize the presentation, not just the pick. The artwork lesson applies to email subject lines, landing page headers, and onboarding.
- Measure quality the way users feel it. VMAF exists because bitrate was the wrong proxy. Find the proxy you are wrongly optimizing.
- Close the loop. Every layer learns from what members did next. If your AI never sees outcomes, it never improves. I wrote up how to build that kind of self-correcting cycle in Write Loops, Not Prompts.
For a wider view of how other companies are deploying this, see how companies are using AI in 2026.
The bottom line
Netflix uses artificial intelligence less as a feature and more as plumbing: one foundation model for what to show, bandits for how to show it, learned quality scores for how it streams, and generative AI where production and search had real constraints. The newest data says most of the value comes from doing personalization competently at all, and the biggest wins land on good items nobody would otherwise find. That part scales down to any business with a catalog, a list, or a funnel.
I break down one real AI system like this every week, with the parts you can copy. Join the newsletter, and grab Write Loops, Not Prompts to build your first feedback loop tonight.
How does Netflix use artificial intelligence?
Netflix uses artificial intelligence in four main places. First, a recommendation foundation model, described on the Netflix tech blog in March 2025, learns from members' full viewing histories and feeds the rows and rankings on the home page. Second, contextual bandit algorithms pick which artwork to show each member for each title, so two people see different thumbnails for the same film. Third, machine learning measures streaming quality through VMAF, Netflix's open source perceptual quality metric, which guides how every title is encoded. Fourth, generative AI now shows up in production: Netflix's Q2 2026 shareholder letter says GenAI workflows were used in roughly 300 of its titles in 2026, mostly in post-production. Search is the newest layer, with LLM-powered natural language and voice search rolling out to members.
What is the Netflix recommendation foundation model?
It is a single large model that replaced much of the work previously done by many separate, specialized recommendation models such as Continue Watching and Today's Top Picks for You. Netflix explained in a March 21, 2025 tech blog post that maintaining those independent models had become costly and that improvements in one rarely transferred to another. The foundation model borrows the large language model recipe: it treats a member's interaction history like a sequence of tokens and is trained to predict the next interaction, using hundreds of billions of interactions from over 300 million users. Other teams then use it directly, fine-tune it for a specific surface, or consume the member and title embeddings it produces. Netflix also reported that scaling laws held, with performance improving as data and model size grew.
Why does Netflix show different thumbnails to different people?
Because the artwork is a recommendation in its own right. In a December 2017 tech blog post, Netflix explained that it personalizes the image for each title per member using contextual bandits, a class of algorithm that balances trying new options against exploiting what already works. A member who watches a lot of romance might see Good Will Hunting represented by Matt Damon and Minnie Driver, while a comedy fan sees Robin Williams. The hard part is attribution: a member can only respond to the one image they were shown, so Netflix has to learn from that feedback loop carefully and avoid clickbait art that misrepresents the title. The system had to handle a peak of over 20 million requests per second at low latency when Netflix described it, which is why the engineering mattered as much as the model.
How much do Netflix recommendations actually matter?
A November 2025 paper by Netflix researchers, The Value of Personalized Recommendations: Evidence from Netflix, put a number on it with a structural model of real viewing data. Replacing Netflix's current recommender with a standard matrix factorization approach would cut engagement by about 4%, and switching to a popularity-based ranking would cut it by about 12%. The same study found that most of the lift comes from effective targeting rather than simply exposing titles more often, and that the largest gains go to mid-popularity titles, not blockbusters or very niche content. That is the useful takeaway for anyone running a catalog: personalization earns the most where a good item would otherwise be invisible, not where demand already exists.
Does Netflix use generative AI to make its shows?
Yes, and it says so in its investor materials. Netflix's Q2 2026 shareholder letter, filed with the SEC in July 2026, states that GenAI workflows have been used in roughly 300 of its titles in 2026, with the largest concentration of work in post-production. The letter frames the tools as a way to deliver higher quality output faster and at lower cost, and notes that some productions would otherwise have had to drop key shots and sequences. Netflix also says it is using large language models to improve title discovery, adding voice search and AI-powered natural language search, and has expanded AI-powered tools across the full advertising lifecycle, from planning and creative production to optimization and reporting. So generative AI touches production, discovery, and ads, on top of the older machine learning stack.
OpusJake is Jake Schincariol's operating system for building with AI: agents, workflows, prompts, and the free resources behind them. Get the next move every week.