AI

Parsing “Can” from “Should”: How Netflix's Product Team Got Thousands of Creatives to Trust an AI Overhaul

Inside Netflix's internal AI rollout, a process that included lots of good old-fashioned thinking, writing and talking.

Parsing “Can” from “Should”: How Netflix's Product Team Got Thousands of Creatives to Trust an AI Overhaul

Everyone says “prototypes are the new PRDs.” Mckenzie Lock doesn’t buy it. “Humans think through words,” she says. “The process of articulating in a way others can follow is the process of decision making.”

She’s helmed products and teams at Apartment List, Affirm, Pinterest and LinkedIn and most recently spent seven years at Netflix as a General Manager of Product, Operations & Creative. There, she oversaw a ~300-person multidisciplinary org with the small task of distributing everything Netflix produces — launching tens of thousands of titles a year and building everything ranging from the technology that dubs a Polish-language show with a natural English lip-sync to AI fan experiences that let Gen Zers put themselves in a scene from Outer Banks.

Netflix straddles the two worlds of creative and tech, each with profoundly different relationships to AI. At most tech companies, the central challenge is moving fast enough. In entertainment, the question is much more existential and fraught. Should this work be automated at all? Will automation dilute the core creative essence of the product? 

The company’s organizational setup reflects this tension: there are two co-CEOs, with Ted Sarandos directing the studio workflows on the creative side of the house, and Greg Peters leading the tech that scales it. Fittingly, Sarandos is based in Los Angeles, where Sunset Boulevard is lined with billboards for the latest movies and TV shows, while Peters is in Silicon Valley, where a drive down 101 to Netflix’s Los Gatos HQ is plastered with AI taglines.

So given the range of viewpoints at play in Netflix’s product decision making, it follows that Lock has some skepticism around AI trends in product development. “Just because an LLM can synthesize customer research doesn't mean a designer should stop doing user interviews,” she says. “The point isn’t to force AI into all your workflows, it’s to develop an instinct for what’s possible. The tools you learn today may be obsolete in two years, but your intuition for them and ability to adapt won't be.”

Netflix has over 17,000 titles available in nearly 200 countries and over 30 languages, serving over 300 million subscribers. AI, naturally, has revealed many opportunities to improve the promotion and distribution of such a massive and ever-expanding catalog. So Lock’s team set out to overhaul Netflix’s creative and operational workflows with a non-negotiable: Find out where workflows are painful, slow or broken, and implement AI to absorb the messiness without compromising the quality of the content, or undermining the work of the craftspeople producing it. “Most product decisions are really choices about where to absorb complexity: in a product, a process or a person, and AI is no different,” she says. “Good design doesn’t eliminate complexity, it puts it in the right place.”

There has been lots of focus lately on the technical menu of AI transformation: which harness to use, models to train, build this versus buy that. But the existential part of getting these changes right still requires the decidedly human processes of ideation, persuasion and judgment.

“Many AI strategies fail at the vision stage not because the technology isn't ready, but because they're built around what AI can do, rather than what humans should do in the context of the industry and company," says Lock. “At Netflix, the context was that we needed automation to massively scale the volume and diversity of our catalog, but in an industry that deeply values craft and creative control.”

The entertainment world may have differently shaped AI problems from the average tech company, but any team can learn from how Netflix tackles this central question: How do you figure out what you should do with AI without getting distracted by what you can do?

Forming a vision

At the outset of Netflix’s internal AI overhaul a few years ago, Lock’s team purposefully didn’t start with what the latest models could do. Instead, her team asked two very open-ended questions: What is the end experience we want to create for members? And where are today’s production workflows genuinely under-serving them?

Since then, this outcomes-first thinking has increasingly become the dominant approach for how forward-thinking teams are tackling AI adoption. But the line to toe for Netflix is especially tricky, given both the quality expectations of subscribers and the AI apprehension of creative partners.

Here’s how her team reimagined what members see on screen and the machinery behind it. 

Dream up your ideal customer scenarios 

Lock started with a hypothetical exercise: If technology were no object, what would we want our members to be able to do? Then the team would work backwards from those ideas. For example, what if someone returned to watching a show after a three-month hiatus and was greeted with a customized recap — which capabilities would they need to make that happen?

Here’s a scenario they dreamed up in vivid detail:

“Even if we never built those exact features, the level of nuance you get by writing out these scenarios is illustrative. The exercise made it obvious which models, services and capabilities we needed, and which were most extensible across content types and member experiences,” Lock says.

Survey the people actually doing the work

The second component of the vision was how these dream member experiences would get made. Everything you see when you log into the Netflix app is produced by an extensive bench of creative talent, including writers, editors, animators, translators, color scientists and dubbing artists. They’re the ones on the ground floor of production, and often the ones most hesitant about changing their workflows. So Lock’s team went straight to them to find out which workflows were the most painful.

She and her team shadowed the workflows of the people doing this work, then created a survey and ran structured debates with subject matter experts and leaders in every capability area. She made the survey open to anyone who wanted to fill it out, regardless of position, to invite more people into the conversation.

“The goal was honest accounting. What does today actually cost you? Where does the quality or speed of our service break down?” she says. 

These are some of the questions she asked in the survey: 

  • Describe the moment in your day when you wish you had information, but it was hard to get. If you had it, what would change in your work product or other people’s work product?
  • If you had to pick between automating tasks in the creation of the media you make or automating the operations around it, which would you choose and why? 
  • What would the best and worst outcome of applying AI to your workflow be? Be as specific as possible. The more examples the better.

Then she set up a series of debates about the results. She read responses ahead of time, picked out two people with opposing views and asked each one to present the bear and bull case to the room. “Everyone felt heard, and it aligned with Netflix’s cultural directive to farm for dissent,” Lock says.

It also brought possible areas of friction to light. One debate surfaced a tension between creative ambition and operational efficiency. A creative leader wanted AI to generate dozens of early visual directions, giving designers more room to experiment before committing to a concept. Another one thought building custom products for creative ideation wouldn’t meaningfully impact the creative outcome — or could lead to a reversion to the mean. And an operations leader saw the bigger opportunity in automatically adapting an approved design across formats and markets, eliminating repetitive production work without asking AI to make the creative decision.

“In a larger company, change is never just about the work, it's also about the often conflicting incentives surrounding it. Surfacing those early helped us manage the human side of this huge change ahead, not just the technical one,” Lock says.

Pain has to be named out loud before any vision can land.

Make the future specific enough to believe in

Lock didn’t want to dismiss any AI ambivalence from creative and editorial teams that surfaced in the survey. So she and her team wrote out a concrete plan for how their jobs might shift away from menial, time-intensive tasks and free up space to do more high-judgment work and cover more titles.

“Unlike tech teams, large creative teams don't just adopt AI for the sake of it,” says Lock. “They need to be able to see themselves in the future state. So I spoke in the language that this audience understood very well: storytelling.”

For every workflow Lock’s team surfaced during the survey phase, they put together a “today versus tomorrow” table. Each had two columns:

  • Today column: What the workflow is today. Budget constraints, quality variance, linear dependencies and time cost, warts and all.
  • Tomorrow column: A description of what someone’s job might look and feel like with an AI overhaul, not a tech roadmap.

Here’s an example of a side-by-side for one workflow in content launch management:

The stark contrast between the two columns does the heavy lifting of persuasion. “You don’t need to hard-sell the future if today is obviously painful,” Lock says.

She also wanted to get creative around how she presented these specific future states to the people who’d be living them. So the team created imaginary press and social media posts to drum up excitement for the process changes to come:

“Part of the vision stage of an AI strategy is literally visualizing the future. Our audience was creatives, so I wanted to use creative methods to visualize it,” she says. “The goal was to make everyone feel that they could be proud of this future work.”

So for a video editor, instead of saying, "AI will auto edit in 20 formats," they created illustrated ‘day in the life’ vignettes of future roles: "Daryl has tools to adjust lip sync, prosody and emotional range in real time. His job becomes exception-handling, quality judgment and craft, not stressing about production schedules."

Framing an AI transformation around human roles is what determines whether the vision survives contact with the people tasked with executing it.

Lock wanted to make both product and creative teams feel a sense of shared ownership in and accountability for achieving results with AI. So she set up an organizational structure, planning process and culture around shared ownership. Product and tech teams owned the ML and product roadmap, and the operational teams owned title launches, staffing and governance. But they not only shared the same annual goals, they co-wrote them and were both on the hook for results.

“This is easy to say and very hard to do. There’s an inherent tension between creative teams and technology teams at basically every company I’ve worked at,” she says. So Lock was very intentional around goal setting across teams. She took a cue from Patrick Lencioni’s “first team” model in The Five Dysfunctions of a Team, which Reed Hastings had everyone read when he was CEO: A leader’s top priority should be the collective success of the peer group they belong to (like fellow executives), not the function they manage. 

She took the idea a step further. She extended the shared accountability model across all functions: product, technology, operations and creative. Each function appointed a representative to a cross-functional “first team.” This group owned the core “system” decisions, domain planning and met regularly to honestly evaluate progress against shared goals. She also had every team add a standing agenda item for future or current misalignments. Calling something a “future” misalignment lowered the temperature and encouraged people to raise issues and discuss them openly before they slowed the work down.

The team then mapped out current and future workflows along with the steps to move from one to the other, including the underlying services they’d need to support them. A key part of this visualization was specifying the inputs and outputs required at each step. For example: 

  • What data or inputs does this step require today and in the future? Where do they come from? 
  • What are the outputs? What downstream canvases or steps consume them? 
  • Which decisions benefit from human discretion and craft? Where does intervention represent friction that the product should eliminate?
  • What metrics define success for this process (error rate, etc.)? How will they be measured/captured? 
  • What’s the highest value feedback loop for this, and how quickly can/should we close it?

“Having the blueprint in mind helped us figure out what to build now, versus later,” says Lock.

Devising a strategy with genuine tradeoffs

Next, the team had to translate the “today versus tomorrow” docs into strategic AI bets. Netflix has long maintained a company-level list of all the strategic bets it’s making, which Lock says is one of the more useful operating tools she’s encountered for creating alignment.

The format of Netflix’s list of bets is simple: Frame each bet as this versus that, with a crisp rationale. “Critically, both sides of the bet have to be genuinely desirable. You can’t just have a good idea versus a bad idea,” says Lock. “If one option is obviously wrong, you don’t have a bet, you have a decision you’re avoiding.”

A lot of company strategies don't acknowledge the inherent risk of a bet or create clarity on the order in which investments need to be made, and why, she says. “Most strategy documents are written to be hard to disagree with. They're full of ‘and’s, not ‘or’s. You nod along, feel good, and walk out of the room not really clear on what the real decision is. A bet should make the either/or explicit, which means someone actually has to choose a path,” she says. 

“For example, if you're driving from San Francisco to Washington, D.C., do you take the northern route through Chicago or the southern route through the Grand Canyon? Do you optimize for speed or for scenery? Do you choose the route with the cheapest hotels or the one with the best driving conditions? The answer depends entirely on what you're optimizing for at what point in time,” she says.

These are two key bets Lock’s team crystallized: 

  • Start where the stakes are lower and we can move fast vs. deploying AI first on content with the largest reach. The team knew they shouldn’t roll out new AI workflows on a tentpole title like Stranger Things, where the viewership is massive, member expectations are hefty, and the stakes for getting the quality right are high. But the tradeoff here was getting workflows up and running for fewer titles so they could safely scale up from there. “AI didn’t need to be perfect everywhere from day one,” says Lock. “We could deploy it where the value-to-risk ratio was highest, learn from real production, and progressively move up the tiers and use cases as the technology and our confidence matured.”
  • Invest in measuring incrementality when possible vs. optimizing for speed to scale without evidence of member impact. Lock’s team wanted to anchor around viewership increases, so they knew they needed to build out measurement capabilities. But the tradeoff was heavy upfront investment, and more A/B testing before ramping up production, instead of purely relying on subjective quality opinions (the typical path in movie production). “Tying AI investments to core company metrics, and not just new AI-specific ones, is what makes the vision understandable to the business,” she says. “We made a conscious decision not to anchor most AI investments on cost reduction. We aren’t making videos or marketing assets to reduce costs, we’re making them to personalize the experience & help members discover great content.”
If you can’t believably connect your AI strategy to the metrics your CEO tracks, you’ll always be fighting for resources.

Lock set ambitious and measurable 18-month automation outcomes, and asked her team to work backwards from them by setting concrete short-term milestones, like shortening turnaround times and lowering error rates, which created urgency.

This was a mindset shift for her team. “At most companies, the standard approach to prioritization is to work forward: Set a theme, generate a list of projects and rank them by ROI. This is fundamentally different,” she says. “You start with the ‘knowable’ outcome you want to achieve. Then you build a plan with a clear causal link between what you ship each month, what you expect to see as a result, where it gets you and what your alternatives are if the plan doesn't work.” 

However, this approach isn’t a one size fits all, Lock cautions. “It only works if leadership is willing to treat the automation as a real commitment and make the organizational, staffing and prioritization decisions it requires. If the strategic conviction isn’t there, or the product is too early to predict the path with confidence, set shorter-term hypotheses and let the evidence determine the next milestone.”

Another tactic that proved helpful at ruthlessly evaluating if a strategy was a specific bet: Get someone else to read it back to you. “I’d ask my leadership team to take turns verbally summarizing someone else’s strategy — something they didn’t write — and answering questions about it,” says Lock. “That had the additional benefit of encouraging systems thinking. It sounds intimidating, but there was actually a lot of laughter when someone heard their own strategy repeated back and realized, 'Oh, that's really not concrete enough.'"

From there, they got to work. Here’s how Netflix’s AI bets took shape across three creative workflows.

Lessons from the ground floor of production

Resist shiny object syndrome

Perhaps surprisingly to anyone watching Netflix in the US, most viewing on the app happens in a different language from the title's original one. That makes dubbing a central workflow for Netflix’s creative teams, but it’s costly and time-intensive to produce, and resources tend to cluster around the biggest titles. So Lock’s team wanted to find a way to use AI to help artists make good dubs for more titles across a broader range of content. 

To do this, the team mapped the current dubbing and subtitling workflow to pinpoint where automation could increase either speed or quality. They rebuilt the pipeline so most pieces of localized text artifacts now start with an AI draft with humans refining where necessary, instead of being built from scratch. And for some content — like titles that need to launch very quickly — they removed editing entirely. 

But the process surfaced a problem the coverage numbers alone didn't show: Viewers were dropping off sooner on dubbed titles than on non-dubbed ones, which meant dubbed titles were less "sticky" (well before AI was introduced). So the team did a deep dive with members, member data, linguists and dubbing specialists to find out why. It turns out one of the problems was lip-syncing: When on-screen lips don’t match the audio, it breaks the immersion in the watching experience. 

One of the first fixes they came up with was to “re-animate” actors' lips in post-production so that their audio matched the lip movement. It was a splashy solution: novel, technically hard, visibly impressive. But the team still tested it like any other bet. They found some pioneering filmmakers and consenting actors eager to try it out and ran A/B tests across these titles, tracking two core metrics: stickiness and viewership.

What they learned surprised them. Re-animating actors’ lips made the titles more sticky, but the AI introduced visual glitches that needed manual cleanup to meet their quality bar. And the benefit wasn't universal: It mattered most in dialogue-heavy content, like dramatic or romantic scenes, and most of the lift came from a few key moments within certain stretches of a film’s run time.

The team faced a choice: Invest heavily to scale the capability across more productions, or focus on other opportunities to hit their quality goals? As excited as they were about the lip re-animation technology, the ROI just wasn't there to justify scaling it broadly (although as the models continue to improve, that calculus could change). 

So they looked for simpler ways to operationalize them at scale. They built a model to read the visual rhythm of human mouths, grade how naturally a dubbed track matches the original screen performance and then highlight key moments to dubbers. They were able to customize machine translation of dub scripts to Netflix-level quality so that the final dub sounds extremely close to the original. If the original dialog looks more like a “hello” on the lips, but the literal translation is “good day,” they choose “hello” as the lip-synced translation.

The big takeaway for us was that the 'less shiny' solution often delivers the biggest punch, especially for the high-quality, long-form content we had on a big screen,” says Lock. 

Don’t get tunnel vision around one metric 

“The running internal joke was that if a title does well, it’s because of the content. If it does badly, it’s because of the artwork,” says Lock.

Most of what you first see in the Netflix app is promotional material — artwork, copy and trailers that help you choose what to watch. Netflix uses recommender systems to personalize this content: You and a friend might see an entirely different image for KPop Demon Hunters based on what each of you is most likely to respond to. That makes building a deep and varied library of creative assets across the catalog central to the strategy. But in Hollywood, more marketing dollars traditionally translated to more viewership, so directors and showrunners had long assumed the same logic applied to streaming: More spend and customization must mean more viewership. 

So Lock’s team decided to test that assumption, while still anchoring around viewership numbers. Sure enough, they found that stills and clips often performed as well as bespoke, conceptual artwork for many titles. That freed the team to automate clip and artwork production more decisively across the catalog — and reserve bespoke artwork and trailers for the popular titles that truly needed it.

To decide where to apply AI first, they broke down each promotional asset into its core components and mapped AI investments against them. For artwork, there were three: selecting and refining stills from the content, creating and applying a title treatment and making sure the artwork makes sense within Netflix’s UI. Still selection is important for personalization, so that’s where the team started.

They ran A/B tests that measured whether specific creative choices actually motivated a viewer to watch, to answer a question like, “Do close-ups beat ensemble shots for comedies?” The system had two halves: a retriever to comb through every still of a piece of content to pick the best one, and a quality control judge to see if it met Netflix’s standards and matched the title’s “vibe.”

The algorithm delivered on surfacing good stills — at the expense of the cohesion of product experience. It quickly started optimizing for a type of artwork that’s very effective at compelling a watch: portraits of good-looking actors. Unchecked, however, the Netflix homepage would have turned into a sea of Julia Roberts for some members. Another notorious example inside Netflix was a famous Indian actress with a small part in an American movie being selected by the artwork algorithm — which would be misleading for Indian members who watched the movie.

So they worked with creative experts to write new guidelines, and then retrained the models with a more nuanced set of metrics and auto-evals. The result of the initiative was large, statistically significant viewership gains and a homepage where members could quickly tell the difference between a dark Nordic psychological thriller and a sweeping nature documentary.

The lesson here was to remove the blinders around one specific metric. “Core metrics are the foundation of evidence-based decision making, but sometimes you need to balance them against what you just plain want the product experience to be,” says Lock. “Did we want a homepage full of Julia Roberts portraits, even if it was effective at increasing a viewership metric today?”

Shrink the problem before you solve it

Almost every system at Netflix runs on tags, or “semantic labels” — the words that describe a piece of content, like “steamy” for Bridgerton. They power everything from budgeting decisions to the collections of shows you see on your home screen. For years, creating this metadata was an extremely manual process: Content analysts watched a movie start to finish, wrote up long context documents and applied tags by hand. Tagging a single title could take more than twice the runtime of the full piece of content, and training a new analyst took two months before they could tag independently.

Lock's team wanted to explore replacing that manual pipeline with AI generated tags and embeddings. But the taxonomy had grown, tag by tag, over years, into thousands of governed concepts and values, each maintained with the same rigor. Teams across the company had quietly built workflows on top of individual tags, but not all of those tags carried equal weight. And the system was often very subjective. "Most of these are editorial labels, not factual ones. They are hard for two humans to agree on, let alone a machine," she says.

So the team shrank the problem first: They built a plan to cut the taxonomy down and create a self-serve prototype for the concepts that came off the list. For governed tags, they used a portfolio of models and interfaces that combine computer vision with both LLMs and traditional classifiers to automatically ingest scripts and third-party data and apply editorial tags. Next, an LLM-based auto-eval layer reviews the tagged video, scores confidence and routes anything below the threshold to a human for review.

So instead of a partner team filing a request and waiting weeks for the content understanding team to map it to an existing tag or build a new one, a merchandiser could type a concept like "cringeworthy awkward first date episodes in teen comedies" directly into a tool — with no one on Lock's team in the loop. Her team retired the long-form context documents analysts used to write for every title and rebuilt roles around quality governance, model prompting and training and evaluation.

But that meant asking people to give up work they were good at and liked. "Some of what we removed was work analysts genuinely enjoyed: watching something closely and writing about it," says Lock. “But it was important to me to be honest with them that two things can be true right now: a task can be enjoyable and add incremental value and also not be necessary for the business anymore. I would often say: our work is to fall in love with the job to be done, not how we do it.”

The results were promising: With the new tech and workflow, tagged titles were live within minutes of being ingested with no manual watching, down from weeks under the old process. This made it possible for Netflix to bring significantly more content onto the service, including time-sensitive and shorter form titles.

The real work wasn't training an agent on our existing workflows, it was admitting that a lot of the inherited processes didn’t need to exist in the first place.

AI is an amplifier, not a direction

Despite her rallying cries to usher Netflix toward an AI-native future, Lock is decidedly a tech neutralist.

Netflix’s internal tension around AI only affirmed that mindset. Perhaps balancing skepticism from creative teammates only strengthened the product team’s decision making. "AI isn’t inherently good or bad. Our decisions are,” she says. “Technology has always become increasingly capable of achieving our objectives. The harder question is whether we've chosen the right ones.”

You can train a model to optimize for almost anything. It’s people who decide what’s worth optimizing for.
Share on: