How Netflix Taught an LLM to Recommend Movies So That You Keep Watching
In this article, we will look at how GenRec was built.


Newsletter
Explain complex systems with simple terms, from the authors of the best-selling system design book series.
20 posts in the last 90 days · newest 7 Oct
In this article, we will look at how GenRec was built.

When LLMs sometimes agree with incorrect claims, it is mostly because their training rewards such behaviour. This reward system is built on several things at once, such as accuracy, helpfulness, politeness, and responses that people like.

In this article, we’ll look at why LLMs have this bias against middle information.

Secure your spot in AI Evals in Practice. Enrollment closes in just 3 days! If you’ve been thinking about joining, now’s the time.

Secure shell is a method to access a remote machine securely over an unsecured network. It starts with a TCP connection from the SSH client to the remote machine (SSH server).

In this article, we will look at why state is the hardest thing in software development and the multiple strategies that developers can use to manage it.

In this article, we will look at how the DoorDash engineering team built this gateway and the decisions they made.

In this article, we will look at why this problem of hallucinations happens with LLMs and the techniques that can help make LLMs more dependable for answering.

I attended a MPP event at Stripe HQ, where I spoke with Emily Sands and Matt Schulman from Stripe, along with Brendan Ryan from Tempo.

Jev is TypeSafe AI’s first System One Model. It is 100x faster and cheaper than frontier LLMs.

Rebuild YouTube with AI, taught by a former YouTube engineer, kicks off on Saturday, September 26. Enrollment closes in 24 hours.

In this article, we are going to look at the entire data lifecycle from creation to deletion and the decisions that need to be taken at each step.

In this article, we are going to look at the various strategies to customize and fine-tune a model.

To understand how it all works end to end, we met with engineers on the GPT Voice team, Zahan Malkani and Justin Uberti (who created WebRTC).

A large AI model can run on modest hardware only by reducing the memory it occupies, reducing the calculations it performs, or moving some work to slower hardware.

Sending a request and reading JSON is one thing. Designing an API that other people can rely on is something where things get complicated.

In this article, we will look at how migrations work at scale and the key strategies that can help make it as efficient as possible.

In this article, we are going to look at how LLMs can find a needle in a haystack.

We’re relaunching Build with Claude Code, a 2-day intensive cohort-based course taught by John Kim, who has trained hundreds of engineers at Meta to use Claude Code in real production workflows.

In this article, we will learn how LLMs handle memory so that they are useful to end users in performing complex tasks that require conversation and holding context.
