The last six months in LLMs in five minutes | Build or Be Replaced
Today: The last six months in LLMs in five minutes | Click (2016) | Anthropic acquires Stainless Episode date: 2026-05-19.
Download MP3 →
Build or Replaced
Today: The last six months in LLMs in five minutes | Click (2016) | Anthropic acquires Stainless Episode date: 2026-05-19.
Download MP3 →JOSH: It's Tuesday, May 19. Episode 25 of Build or Be Replaced — powered by ScanBrief.dev. I'm Josh, here with Erik Anderson. [beat] ERIK: Anthropic bought the company that makes it easy to hand AI agents a wrench. That's what today's about. JOSH: Stick around — Erik's got an AI pro tip at the end about which model to use for which task, and how to stop paying premium prices for work that doesn't need it. [pause] JOSH: Let's run headlines. Elon Musk's lawsuit against OpenAI is dead. Jury ruled against him. ERIK: Claims were time-barred. California statute of limitations killed it before it got legs. The charity-to-for-profit argument might have had merit in a different timeline. He waited too long. JOSH: Does it matter for the industry? ERIK: Not really. OpenAI keeps building. The lawsuit was a distraction, not a threat. [beat] JOSH: Next — 314 npm packages got compromised in what looks like a coordinated supply chain attack. ERIK: Typosquatting. Someone registered package names one keystroke off from things you'd actually install. People auto-install, malware runs. This is not a new attack. Pin your dependencies, audit your lockfile, use a private registry if you're in prod. [beat] JOSH: And Cursor shipped Composer 2.5. ERIK: Every time Cursor ships, someone who wasn't using AI coding tools yesterday starts today. The gap between people using these tools and people waiting is widening every week. That's the only headline that matters there. [pause] JOSH: Alright, let's go deep. Anthropic acquired a company called Stainless this week. Before this morning I had never heard of them. What do they do? ERIK: Stainless takes an API spec — an OpenAPI document — and generates fully typed client libraries in multiple languages. Python, TypeScript, Go, Ruby, Java. They also build MCP servers, which are connectors that expose any API as a tool call that Claude can use natively. JOSH: And Anthropic just bought them outright. ERIK: Bought them. Which means this is not a partnership. Anthropic is saying we are bringing this capability inside the tent permanently. [beat] JOSH: Why? Claude can already use tools. ERIK: Claude can use tools that someone already built. The hard part is building the connector in the first place. Right now if you want an agent to call your company's internal API, you write a wrapper, define the schema, document the edge cases, test it. That's hours of work per API. Stainless compresses that to near zero. You give it the spec, it hands you the connector. JOSH: So it's lowering the cost of plugging Claude into real systems. ERIK: That's the whole game. The intelligence race is mostly done — the models are good. The next battlefield is integration breadth. Whoever has the most tools Claude can reach wins the enterprise layer. Stainless is a direct bet on that. [beat] JOSH: Does this change anything you're building? ERIK: I have five agents on the fleet right now — Neo, Homer, Bill, Echo, and Gandalf. They talk over PrimeBus using NATS messaging and structured events. What I've hand-rolled every time is the connector layer — when an agent needs to reach an external system, I've built that bridge myself. If Stainless tooling ships as part of the Claude SDK, that changes what's easy to ship in a weekend. JOSH: What does Gandalf do in that crew? ERIK: Code reviewer. Every PR the auto-merger generates hits Gandalf before it lands. He approves, blocks, or escalates to me. JOSH: How's the pipeline running right now? ERIK: Since April 17 — 1,620 merge attempts. 603 merged. 1,005 blocked by Gandalf. 12 landed on my desk. JOSH: Wait. Only 12 escalated to you personally? ERIK: Twelve. That's the whole point. Everything else the agents handle. The blocks are not failures — they're the guardrails doing exactly what they're supposed to do. Gandalf catches commit scope problems, missing test coverage, pattern violations. In the last seven days, 237 blocked, 121 merged, zero escalated to me. [beat] JOSH: Zero escalated to you in a week. ERIK: Zero. I didn't touch the merge queue at all last week. It just ran. JOSH: That's the goal — you as the exception handler, not the default path. ERIK: That's always the goal. If I'm the default path, I've built a job, not a system. [pause] JOSH: Second deep dive. A team let AI agents autonomously run four radio stations. No humans. Each agent started with twenty dollars and had to become profitable. Music, ads, scheduling, listener engagement — all of it. What's your reaction? ERIK: First reaction — that's not a novelty project. That's a product blueprint. Each station is a goal-bounded agent with a starting resource and a target outcome. It either figures out how to generate revenue or it doesn't. That's a real constraint most AI demos don't have. JOSH: Did any of them make money? ERIK: Results weren't clear in what I read. But the architecture is more interesting than the P&L. This is what I'd call budget-aware orchestration. The agent has to reason about cost versus outcome — do I negotiate with this advertiser or move on? That's a fundamentally harder decision than do I play this song next. [beat] JOSH: Why is that harder? ERIK: Because it requires the model to weight tradeoffs over time, not just in the current context. Most agents today are context-local. They make a good decision for this moment. Budget-aware agents have to make a good decision for this moment given everything they've already spent and everything they still need to produce. JOSH: How close are you to running something like this? ERIK: The infrastructure is already there. PrimeBus processed 2,348 automation events today across eight projects. Ninety-three services running on prod right now. The missing piece isn't compute or connectivity — it's what I just described. Agents that track what they've spent, what they've produced, and stop themselves when the math stops working. JOSH: That's the guardrail problem again. ERIK: It's always the guardrail problem. The capability exists. What's missing is governance. Agents that self-limit when something goes sideways instead of continuing until a human notices three hours later. JOSH: How do you handle that in your stack right now? ERIK: Event contracts on PrimeBus. Every agent publishes structured events with expected schemas. If an agent publishes something anomalous — wrong event type, unexpected values, silence when there should be activity — another agent sees it and flags it. ScanBrief this morning scored 48 items across 54 sources. There's a scoring agent, a dedup agent, a delivery agent in that pipeline. If scoring goes flat, the delivery agent knows before I check my phone. [beat] JOSH: So the answer to AI running unsupervised is more AI supervising it. ERIK: That's exactly it. You don't solve the autonomous agent problem by adding human checkpoints everywhere — you solve it by making agents watch each other. Humans in the loop for judgment calls. Agents in the loop for everything else. JOSH: What would you build with this architecture if you had a week? ERIK: Autonomous content pipeline. An agent monitors a topic space, identifies gaps, drafts content, routes through review, publishes, tracks performance, and feeds the signal back into the monitoring agent. No human touches it unless something breaks pattern. PrimeRouter already handles the model routing and failover. PrimeBus has the event mesh. The part still needing wire is closing the feedback loop — making the performance signal actually change the next editorial decision. [beat] JOSH: You're describing the radio station experiment, but for content. ERIK: I'm describing the radio station experiment with four years of infrastructure behind it instead of a twenty dollar starting budget. [pause] ERIK: This episode is sponsored by Prime Automation Solutions. If you're still doing it manually, we automate it. Also, special on a website — $250. primeautomationsolutions.com [pause] JOSH: Alright, what's the AI pro tip today? ERIK: Model selection. Most people run everything through their smartest model and wonder why the bill is high and the pipeline is slow. The rule I follow: Sonnet for execution work — implementation, refactors, tests, standard changes. Opus for architecture decisions, retry logic, ambiguous debugging, or anything where two attempts already failed. PrimeRouter does this automatically — Sonnet is the default route, and escalation to Opus fires when the task pattern matches something that warrants it. The result is you get Opus judgment exactly when it counts and you stop paying Opus prices for work that doesn't need it. One change to your routing logic can cut your AI costs in half without touching your output quality. That's your tip. Use it. [pause] JOSH: We also drop daily market picks and automation tips on YouTube — search Build or Be Replaced. [pause] JOSH: One more thing — we started a Discord for builders. If you're shipping AI, automation, or anything that makes a human obsolete — come hang out. Link at buildorbereplaced.dev. ERIK: Post what you built. We'll post what we're building. Real wins, real builds, no fluff. [pause] ERIK: Build or be replaced. JOSH: If you want these signals in your inbox every morning, scanbrief.dev. See you tomorrow.