How Netflix Taught an LLM to Recommend Movies So That You Keep Watching
In this article, we will look at how GenRec was built.

Architecture, distributed systems and how large software holds up.
Updated
In this article, we will look at how GenRec was built.

Admission controls in software determine what does or doesn't happen next, and the most common implementation has ancient roots.

On October 11, 2026, the DNS root switches to a new key-signing key (KSK-2024). Learn what this means for you, and how RFC 8509 trust anchor sentinels allow you to test whether your DNS resolver is ready for the rollover.

When LLMs sometimes agree with incorrect claims, it is mostly because their training rewards such behaviour. This reward system is built on several things at once, such as accuracy, helpfulness, politeness, and responses that people like.


Use dedicated read replicas to isolate read-only workloads and distribute your data across data centers around the world.

In this article, we’ll look at why LLMs have this bias against middle information.

#187: Part 1 - Generative AI, RAG, AI Agents, and 14 others.

We celebrated our 16th birthday with 46 announcements across open source, post-quantum security, AI agents, and developer platform upgrades. Here’s a day-by-day roundup of everything we shipped.

A year after announcing our goal to hire 1,111 interns, more than 750 early-career builders have shipped real products across 48 teams at Cloudflare. From Birthday Week launches to post-quantum security, our interns prove that AI amplifies human potential instead of replacing it.

In response to my last fragments (probably the bit about us worrying if LLMs have consciousness when we when we should be wondering why they don’t have a conscience) “Metalanguage” replied: we shipped the id and forgot the superego. classic software lifecycle. I don’t know what…
Cheaper intelligence, faster interfaces, and the judgment needed to turn agent work into useful software.

Secure your spot in AI Evals in Practice. Enrollment closes in just 3 days! If you’ve been thinking about joining, now’s the time.

#186: Understand HNSW Graphs, Approximate Nearest Neighbor Search, and Key Parameters Engineers Tune in Production

Strands Decider: Why Not an Encoder? Expertise level: exhausted. As I admitted in my last post, I am very much not an expert model developer. What follows is likely to be, at least partially, inaccurate. I have tried to be careful and quantitative, to partially balance a lack of…
Secure shell is a method to access a remote machine securely over an unsecured network. It starts with a TCP connection from the SSH client to the remote machine (SSH server).

#185: Part 2 - JWT, WebSockets, OpenAPI, and 13 others.

Cloudflare is launching eight major updates that bring logs, traces, analytics, alerts, dashboards, querying, and telemetry export into one observability platform, with simpler and more predictable pricing.

Streamline demonstrates how to build long-running, continuous video processing pipelines by pairing Cloudflare Workers and Durable Objects with a containerized media engine.

#184: Part 1 - REST, GraphQL, Authentication, and 14 others.


How Neki only decodes the parts of a PostgreSQL result the router actually needs.
Not every presentation needs slides, but when they earn their place, they work as a visual channel that complements the speaker rather than replacing them. Sumeet Gayathri Moghe sets out the principles that follow from that idea — control over the rate of knowledge exposition…

A Sensible Default is a practice that, absent some overriding context, should be used when carrying out a certain kind of task. In software development such sensible defaults might include things like “use version control”, “separate UI logic from domain logic”, “automate…
Simon Willison: The more time I spend working with coding agents, the more convinced I am that they make software engineering even harder We can do amazing things with them, but unlocking their full potential requires extraordinary discipline and knowledge This has been a…
Small Decisions: Engineering a Leading Model Or trying to, at least. Update: An evolved version of the model described here is now available as part of Strands. Check out strands-decider on github and HuggingFace. My day job has been primarily in AI for three years now, but I’d…
What makes a good sharding strategy today might not work so well tomorrow. What happens when you have one tenant hitting your database orders of magnitude more than any other? Or product managers knocking down your door with new requirements? If you're already on Neki…
Cloudflare’s 100TB RAM saving, faster Postgres analytics, and keeping humans in the loop as agents take on more work.


ARM and x86 can give you different performance at the same vCPU count. Here's what differs, what to measure, and how to choose.
Rob Bowley is “flipping tables in his head” with anger at the current media coverage of the danger of AI killing us all The risk I’m worried about isn’t a future machine deciding to wipe us out. It’s today’s AI, being wired into everything, carelessly and fast. Cyber attacks…
General methods that scale with computation will inevitably displace hand-engineered domain knowledge! Sutton’s Bitter Lesson hangs like a Sword of Damocles over all of us. This paper, which just dropped, applies that lesson to data agents, and says that as LLMs improve, they…
A dialogue with Dr Cat Hicks on ethics, learning, and what we owe each other. Part one: start by listening.

My quest for a principled solution for metastable failures has taken me back to my roots on self-stabilization, as this recent paper related the problem to composition of self-stabilizing systems. But, my literature search for recent work on composing self-stabilizing systems…
One of the interesting challenges of the AI ecosystem in 2026 is that new, effective patterns emerge faster than I can adopt them. I’ll find a handful, get back to work, and realize a month later that I’d missed four or five more. The adoption cycle for Imprint this year has…
Coupling, not naming. First article on API design.

In early 2025, I started seeing people turn off their brain as they use LLMs. They would have an LLM take an action (summarize text, write some code, etc.), and just assume that it worked. This generally didn't work in early 2025 and the result was often quite silly. As LLMs…
AI doomerism is everywhere these days. Every field, and recently humanity as a whole, has had its "we're finished" post. In contrast to the run-of-the-mill hot take, Jason Potts has written an economics paper on why academia is doomed. So let's dive in. What does a university…
A framework for thinking about when AI involvement is additive or a violation

Aleksey and I are back to reading papers live. This paper, Jetpack(OSDI '26), attempts building a universal 1-RTT fast-path framework that bolts onto existing leader-based consensus protocols with minimal modification. Why would we want this? Classic consensus protocols like…
Where delays hide in the stack, how they add up, and how to measure them.

We previously noted that, while it's easier than ever to hit a particular quality bar by having coding agents use effective test techniques, software quality seems to be getting worse, indicating that whatever defaults developers are using may not work very well. Here, we test…
What latency really is. The physics and math that bound it, and why the average always lies.

Last week I wrote about modular verification of systems through open TLA+ specs and rely-guarantee discharge. In this post, I apply the same approach to study the metastability mechanics of a retry storm. Through this modeling I show that when the system is metastable, it is due…
Sprites are disposable cloud computers. They appear instantly, always include durable filesystems, and cost practically nothing when idle. They’re the best and safest place on the Internet to run agents and we want you to create dozens of them. Sprites are a place to run agents…
I started writing Logic for Programmers because there weren’t any good resources on logic for, uh, programmers. Now that the book’s out, the new problem is that there aren’t any good free resources on logic for programmers. So, to solve that problem (and maybe hype the book a…
I was curious how well the predictions of the most widely cited AI skeptic I've seen (Ed Zitron) have done, so I looked at how his predictions panned out. To disclose my own biases, I've never had a particularly strong pro or anti AI progress position. For example, in 2022, I…
I used to wonder why I see so many more bugs than most people. I easily observe hundreds to thousands of bugs per week and nothing seems to work, but most people I talk to don't see anything like this. For a long time, I thought this had something to do with how I use computers…
The other day, I saw a viral tweet saying that people talking about how LLMs are causing slow, bloated, code are going to eat crow once they re-write everything in super-optimized assembly. We're not quite at the point where we want to write everything in assembly, but some…
One thing that bothered me about Imprint’s product after joining was our lack of passkey support. Passkey support is a rare opportunity to increase resiliency to phishing attacks while simultaneously reducing login friction. If it’s good for our members, our partners, and our…
Six years ago, I wrote Tech Lead Management roles are a trap. My argument then was that TLM roles present themselves as easier than moving into a full management role, but the tension between doing the software engineering and engineering management aspects of the role made…
Lorenz and Little: How Much Does Your Tail Cost? Lorenz and Little sounds like hipster burger bar from 2015. It’s time for Marc’s Amateur Statistics Corner! Today: why I pay a lot of attention to tail latency when optimizing cost. I’ve written before on the importance of tail…
I am delighted to announce that my book, Logic for Programmers, is now available! You can check out the release site here or go directly to buy the ebook or print versions. If you bought any of the early access versions, you can get the 1.0 for free from leanpub. This has been a…
We’re Fly.io, a public cloud platform that is both our favorite way to put an app on the Internet and our favorite way to safely let a frontier agent coding harness cook. This is a post about our company, the future, and Sprites, which are computers for agents that you can check…
Aurora DSQL: Scalable, Multi-Region OLTP A paper! Our new paper, Aurora DSQL: Scalable, Multi-Region OLTP, is now available on Arxiv. I’m excited about this one: it’s a fully end-to-end look at how Aurora DSQL works, from query processing, to transactions, to replication, to the…

Eight years ago, I wrote about my theory of restoring struggling teams, which came down to four steps: A team is falling behind if each week their backlog is longer than the week before. Solve by hiring more. A team is treading water if they’re able to get their critical work…

I’ve recently been thinking a lot about the concept of “soil horizons”, which is the idea that there are many distinct layers of soil, from topsoil all the way down to bedrock, which all combine into a soil horizon. Translating this idea into software, the ideal codebase would…

Years ago I predicted that columnar storage would remake observability. What I didn't see coming: vendors would build it, nerf it, and sell a worse version back to you as "Datadog, but cheaper"

When you need to execute a coordinated change on a tight timeline, a mandate is the best and most honest way to fund it.

On externalities and engagement, the problem with purity politics, and why we need to make AI boring again

Meet Alice. Alice is impatient. What do you mean? Meet Alice. Alice uses your web service. Alice, like most humans, measures her time in seconds and minutes. Alice says your service is slow. You tell Alice that the mean request to your service completes in 100ms, but Alice says…
Building agents is fun. Rebuilding agents that break themselves… less so. A lot of Fly people are building agents with less of a penchant for self-destruction by teaching their agents to do anything risky in a Sprite. You get an agent that stays alive long enough to actually use…

It’s April Cools! It’s like April Fools, except instead of cringe comedy you make genuine content that’s different from what you usually do. For example, last year I talked about The best introductory video games for non-gamers. This year I’m picking a fight. This is “New York”…
This week, Amazon tightened review after a nearly six-hour retail outage and a 13-hour AWS incident tied to AI-written code, while Stripe's 11-task benchmark showed Claude Opus 4.5 hitting 92% on task

This week, why LLM-generated code can pass tests yet fail on performance, how a GitHub issue title triggered a supply-chain breach on 4,000 machines, and why verification hasn't caught up with AI code

As part of writing Logic for Programmers I produced a lot of “chaff”, code samples and sections I wrote up and then threw away. Sometimes I found a better example for the same topic, sometimes I threw the topic away entirely. It felt bad to let everything all rot on my hard…
I’m Ben Johnson, and I work on Litestream at Fly.io. Litestream is the missing backup/restore system for SQLite. It’s free, open-source software that should run anywhere, and you can read more about it here. Each time we write about it, we get a little bit better at golfing down…

We’re Fly.io, and this is the place in the post where we’d normally tell you that our job is to take your containers and run them on our own hardware all around the world. But last week, we launched Sprites, and they don’t work that way at all. Sprites are something new: Docker…

This week, Bloom filters slash API latency 16×, MAKER cracks a million-step task with zero errors using agent voting, and incremental architecture shows why small shifts beat big-bang rewrites always.
