Can a Cloud-Native Harness Make Agents Reliable Beyond the Desktop?
Kubernetes co-creators Craig McLuckie and Joe Beda aim to bring agent harnesses fully into the cloud.

Building with models, by people who ship with them.
Updated
Kubernetes co-creators Craig McLuckie and Joe Beda aim to bring agent harnesses fully into the cloud.



Here's a product feature which the world is going to need a whole lot more of over the coming months and years: default hard budget caps. I'm talking about the feature of pay-by-usage services and APIs that lets you say "after $X/month, cut this thing off and return errors"…

After leading Meta’s Llama models, Ahmad Al-Dahle is now transforming Airbnb with AI — from how its teams develop products to how it serves guests.

I'm at OpenAI DevDay today, in Fort Mason, San Francisco. Same as last year I'll be live blogging the keynote and some other notes during the day. OpenAI gave me a free ticket and a seat in the "creator" area for the keynote. You are only seeing the long-form articles from my…
A Visual Guide to RNNs, CNNs, Transformers, and Calibration, with Hands-On Experiments on Accuracy and Efficiency

Line-by-line review is going away. Whatever replaces it has to earn the trust reading used to provide.

On Friday I gave the closing keynote at the WeAreDevelopers World Congress North America in San Jose. I tied together the key trends from the past year into a chronological exploration of everything that happened in 2026. The video is on YouTube; here are my annotated slides and…


Yesterday was Grok 4.7 (pelicans) and MiMo v2.6 Flash/Pro (more pelicans). Today Anthropic released Claude Opus 5.5, and around an hour later OpenAI released GPT-6 Sol and GPT-6 Luna. It's going to take a while to get a good read on all of these new models, but here are my…

Last week TypeSafe AI unveiled Jev, their first example of a new category of model that they are calling "System One models" (I'm with Maggie Appleton, I think "decision models" is a better name for these). Jev is an interesting variant on the usual LLM format: it still accepts…

What it takes to run agents in a codebase older than the team


A Look at Recurrent Depth, Hidden Chains of Thought, and Recent Research on Looping Transformer Blocks


Agents can finish the task without teaching you anything. Building expertise now has to be deliberate.


A 48-minute video walkthrough of token sampling, watermark detection, and removal

A field guide to building a software factory that still has an owner.

An End-to-End Project With Dataset Construction, Model Training, Local Deployment, and RLVR

How LLMs Learn Low-, Medium-, and High-Effort Reasoning Modes
