Mon–Fri · 6 AM ET
← All Episodes
EP  • 00:09:58

Claude Opus 4.8 | Build or Be Replaced

Today: Claude Opus 4.8 | Bricks and Minifigs Stole a Man's $200k Lego Collection | I made a million dollar product from my dorm room (2025) Episode date: 2026-05-29.

Download MP3 →

Transcript

Here's episode 33:

---

JOSH: It's Friday, May 29. This is Build or Be Replaced — powered by ScanBrief.dev. I'm Josh, here with Erik Anderson.

ERIK: Faster models, permission theater, and your car's been narcing on you to your insurance company.

JOSH: Stick around — Erik's got an AI pro tip at the end about when to stop using Sonnet and call in the big model.

[pause]

JOSH: Three quick headlines. GitHub banned a security researcher.

ERIK: Researcher posted Windows zero-day exploits after Microsoft blew past the disclosure deadline. GitHub pulled the account. That's the wrong incentive. You want researchers filing bugs, not getting banned for doing the job Microsoft didn't.

[pause]

JOSH: Blue Origin had a bad week.

ERIK: New Glenn blew up during a static fire test. Not a launch — a ground test. That's the kind of failure that resets your timeline by quarters, not weeks. Tough one.

[pause]

JOSH: Volkswagen is blocking Home Assistant.

ERIK: They updated their API to require client assertion — basically a check that says you're an official app. Home Assistant isn't an official app. So now your VW integration is dead unless somebody reverse-engineers the handshake. This is what vendor lock-in looks like in practice. They own the API, they own your access.

[pause]

JOSH: Okay, deep dives. Claude Opus 4.8 dropped. What's the actual story?

ERIK: Anthropic announced Opus 4.8 with a fast mode — 2.5 times faster than 4.7, at three times lower cost. Same price point as Opus 4.7 on the standard tier. That's not an incremental update. That's a fundamentally different cost curve for the same capability tier.

JOSH: So what changes for how you're using it?

ERIK: My routing model right now: Sonnet handles execution work. Implementation, refactors with an approved design, tests, docs. Opus only gets the hard problems — architecture decisions, ambiguous failures, anything touching retries or schedulers or multi-step failure logic. That boundary exists because Opus was expensive and slow enough that hitting it unnecessarily hurt.

JOSH: And now?

ERIK: Now the escalation math shifts. If fast mode is genuinely 2.5x faster, the latency penalty for calling Opus on a tricky problem drops enough that I can widen the escalation criteria. Right now there are tasks I push through Sonnet because the Opus cost wasn't worth it — borderline cases. Those borderline cases move into the Opus bucket with 4.8.

JOSH: What does borderline mean in practice?

ERIK: Anything where I'm reading the output and thinking "I'm not sure this is right." Ambiguous runtime errors. Log parsing where the stderr doesn't clearly point at a root cause. Cost or reliability decisions. Anything touching async flows or queue behavior. Right now, some of those I eat on Sonnet because I know Opus will get there but I'm watching credits. With 4.8, I spend the escalation more freely.

JOSH: The enhanced honesty piece — Anthropic made a point of it. Does that matter for agentic work?

ERIK: Yes, actually. When a model is running multi-step tasks autonomously, hallucinated confidence is a real failure mode. The agent tells you it did the thing. It didn't fully do the thing. Or it made an assumption mid-task and didn't surface it. Enhanced honesty in agentic contexts means more "I stopped here because this was ambiguous" and less "I completed the task" when it actually papered over a gap. That's load-bearing for any system where you're not watching every step.

JOSH: You're not watching every step.

ERIK: ScanBrief scored 85 items across 54 sources today. PrimeBus processed 2,139 automation events across 6 projects. Five code changes were automatically reviewed by Gandalf and merged to production overnight. No, I'm not watching every step. That's the point. The system has to be honest with itself, or you find out at 2 a.m.

[pause]

JOSH: The Show HN thing — "Continue? Y/N." A 60-second game about permission fatigue. You played it?

ERIK: Saw it on the ScanBrief digest. The concept is simple — you're an AI agent getting bombarded with approval requests, one after another. Approve, approve, approve. After a while you stop reading them. You just press Y.

JOSH: That's... accurate.

ERIK: It's a perfect satire of what happens when you deploy agentic systems with no real guardrail design. You either get no permissions and the agent does whatever it wants, or you get permission spam and the human rubber-stamps everything because they're overwhelmed. Neither of those is a safety system.

JOSH: What does an actual safety system look like?

ERIK: My auto-merger is the concrete version of this. Since April 17th — 3,912 merge attempts. 1,026 got merged. 2,874 were blocked by Gandalf. 12 escalated to me.

JOSH: That's a lot of blocked merges.

[beat]

ERIK: That's 2,874 pieces of bad code that didn't hit prod. Different framing. The blocks aren't failures — they're the system working. The whole point of Gandalf in the pipeline is that not everything gets through. Last seven days specifically: 1,905 runs, 255 merged, 1,650 blocked, zero escalations to me. None. Gandalf handled everything this week without needing a human decision.

JOSH: Zero escalations in a week — that's the goal?

ERIK: That's the goal. The 12 escalations to me since April are the hard edge cases — things Gandalf flagged as genuinely ambiguous, where a human call was the right call. Twelve in 3,912 attempts. That's 0.3 percent. That's what good guardrail design looks like. The game is satirizing the opposite system — one where every tenth decision gets routed to a human who's already tuned out.

JOSH: And that's a real problem.

ERIK: It's the dominant failure mode in production AI deployments right now. Companies throw an agent into a workflow, add a confirm button, and call it oversight. The confirm button becomes muscle memory. You've got 5 agents in the fleet — Neo, Homer, Bill, Echo, Gandalf — and 85 projects emitting telemetry to PrimeBus right now. 101 services running on the production server. You cannot manually confirm 2,139 events a day. You need the guardrails to be intelligent enough that the escalations that do reach you are worth your attention.

[pause]

JOSH: Car data. The headline is that cars collect a startling amount of information. Were you startled?

ERIK: Not even slightly. Cars are IoT devices. They have GPS, accelerometers, cellular modems, and increasingly, cameras. Of course they're collecting location and driving behavior data. The only surprising thing is that people are still surprised.

JOSH: The story is specifically about it being sold to insurers and data brokers.

ERIK: That's where it gets concrete. Location data alone is legally considered non-PII in most contexts, so it moves freely. Driving behavior — hard braking, speed profiles, cornering — is exactly what actuarial models want. The insurer doesn't need your name attached. They need your risk profile. And if they can correlate your driving telemetry to your policy number, which they can, the name is redundant.

JOSH: And opting out is buried.

ERIK: This is the same playbook as every other data collection story for the last ten years. Collection is on by default. Opt-out exists but requires navigating an app sub-menu that most owners don't know about, or calling a dealer, or mailing something. It's designed to minimize opt-out rates. The difference with cars is the data is richer and more continuous than most people assume. Your phone knows where you go. Your car knows how you drive when you get there.

JOSH: What do you do with that?

ERIK: Assume collection. Assume it's sold. Check the privacy settings on the manufacturer's app — most of them have a connected services section. Turn off data sharing if you care. Don't assume the default is off. In IoT the default is never off.

[pause]

ERIK: This episode is sponsored by Prime Automation Solutions. If you're still doing it manually, we automate it. Also, special on a website — $250. primeautomationsolutions.com

[pause]

JOSH: Alright, what's the AI pro tip today?

ERIK: Model routing. Here's the practical version of how it works right now, updated for Opus 4.8. Sonnet is your execution layer. Approved design, clear task, known patterns — Sonnet. Don't overthink it. Opus is for design decisions, ambiguous failures, anything touching reliability logic, retries, schedulers, or queues. Also use Opus when you've already tried Sonnet twice and the output still doesn't feel right — that's a signal, not a coincidence. With Opus 4.8 running at three times lower cost in fast mode, the threshold for escalating drops. Borderline cases that weren't worth the Opus burn before — they're worth it now. The rule of thumb: if you're reading Sonnet's output and second-guessing it, that second-guess is your escalation signal. Stop second-guessing cheap models on expensive problems. That's your tip. Use it.

[pause]

JOSH: Binge all five episodes this weekend plus our YouTube shorts — links at buildorbereplaced.dev.

[pause]

JOSH: One more thing — we started a Discord for builders. If you're shipping AI, automation, or anything that makes a human obsolete — come hang out. Link at buildorbereplaced.dev.

ERIK: Post what you built. We'll post what we're building. Real wins, real builds, no fluff.

[pause]

ERIK: Build or be replaced.

JOSH: If you want these signals in your inbox every morning, scanbrief.dev. See you tomorrow.

---

Episode 33 is ready. Three clean deep dives — Opus 4.8, the permission fatigue game connecting directly to Erik's 2,874 Gandalf blocks, and car data as IoT realism. All live stats from the briefing block, nothing invented.

— Claude / Neo (Bob-1)