Transcript
ERIK: It's Monday, September 28th. This is Build or Be Replaced — powered by ScanBrief.dev. I'm Josh, here with Erik Anderson.
JOSH: Episode 115. ScanBrief already pulled 57 items across 56 sources before either of us finished coffee.
ERIK: 115 episodes. That's not a milestone number, that's just a Monday.
JOSH: Stick around — Erik's got an AI pro tip at the end about getting more work out of fewer tokens.
[pause]
JOSH: Alright, headlines. First one — somebody's owed a billion dollars in Nvidia stock?
ERIK: Early Nvidia technical advisor, stock options from 1993, texture-mapping code that ended up in the NV1. Vested at about a billion bucks. That's not luck, that's someone who understood the platform before there was a platform.
JOSH: Next — "When did Google get so weird?" Apparently search is auto-generating overviews for the most random queries now.
ERIK: Yeah, and misreading intent while it does it. Google's not indexing the web anymore, it's guessing at what you meant and showing you that instead. Different product wearing the same logo.
JOSH: Third — "Don't couple your Go code to GitHub."
ERIK: Simple rule, people ignore it constantly. If your business logic imports a GitHub SDK type, you don't have Go code anymore, you have GitHub code that happens to compile. I've made that mistake. I don't make it twice.
JOSH: And last one — a survey's going around saying most enterprise "AI agent" pilots never make it to production.
ERIK: That tracks. Everybody wants the demo, nobody wants the guardrails. A pilot that works once in a sandbox isn't a system, it's a screenshot. You find that out the first time it touches real data and nobody built the part that catches it when it's wrong.
[pause]
JOSH: Let's go deeper on a few of these. Start with the one that caught my eye — Ember-1.
ERIK: Fireworks Research took Kimi K3 and built a specialized coding model on top of it. Comparable quality, about 40% fewer tokens to get there. They did it by teaching the model to reason more efficiently, not by making it dumber.
JOSH: What does 40% fewer tokens actually buy you?
ERIK: Money and speed, in that order. Every token you don't generate is a token you don't pay for and don't wait on. I run Hermes locally on the Mac M3 — qwen3-coder, about 70 tokens a second — and the whole reason that setup exists is because I got tired of paying cloud prices for reasoning I could get for free on hardware I already own.
JOSH: So is Ember-1 something you'd actually route work to?
ERIK: Not today. It's Fireworks' serverless stack, it's new, and I don't put unproven models anywhere near a pipeline that touches production. But I track it. The pattern matters more than the specific model — efficient reasoning is the trend, not a one-off. Anthropic's been chasing the same thing, OpenAI too. The winners aren't whoever has the biggest model, they're whoever wastes the fewest tokens getting to the right answer.
JOSH: How do you even test that without burning money finding out?
ERIK: You don't test it live. You test it in a lab with a fixed budget and a scorecard, and you don't promote anything until you've got real numbers, not vibes. Same discipline as anything else I run. Any new backend sits dark behind a flag until it's proven — no exceptions, doesn't matter how good the benchmark looks.
[pause]
JOSH: Speaking of discipline — one of today's stories is basically about your job. "There is more to code review than automatable detection."
ERIK: That one's correct, and it's the whole argument for why PrimeBus exists the way it does. The article's point is that a linter catching a bug isn't the same as a human catching a bad decision. I agree completely. So I split the job.
JOSH: Split it how?
ERIK: PrimeBus handles the automatable part — tests fail, it generates fix variants, it tries to merge the good one. Gandalf sits in front of that as the judgment layer. It's not there to catch syntax errors, it's there to catch "this technically works but it's a bad idea."
JOSH: Does that actually hold up at volume?
ERIK: Since June 5th, PrimeBus has run 2,295 merge attempts. 1,507 got merged. 788 got blocked by Gandalf. Zero escalated to me.
JOSH: Wait, 788 blocked — that sounds like a lot of failures.
ERIK: That's the part people get backwards. That's not a failure rate, that's the guardrail doing its job 788 times without waking me up. If Gandalf wasn't blocking anything, I'd be worried, not relieved. The story people want to tell is "AI review replaced human review." The real story is AI review replaced the boring 65% of human review and got stricter about the rest.
JOSH: So the automatable detection article — you'd say PrimeBus agrees with it, not fights it.
ERIK: Completely agrees. The mistake is building a pipeline that thinks passing tests equals a good merge. Mine doesn't think that. That's why zero of 2,295 attempts needed me directly — not because nothing went wrong, but because the system knows the difference between "broken" and "wrong," and only escalates the second one. And today's digest is still coming in hot — 93 repo changes, 93 test runs, 29 alerts already this morning. Gandalf's already working through it before we finished recording.
[pause]
JOSH: Last deep dive — "Malleable software," restoring user agency in a world of locked-down apps.
ERIK: This is the one I'd tattoo on the wall. The argument is that modern software stopped letting people adapt it — everything's a sealed app, you get the buttons they give you and nothing else. That's exactly the problem PrimeDash solves for me at home.
JOSH: PrimeDash is your dashboard for everything running in the lab, right?
ERIK: Right now that's 153 services on the production server alone, and 140 distinct projects feeding telemetry into PrimeBus. Nobody ships me a dashboard that shows all of that together — I had to build one, because the commercial tools assume you have three services, not a hundred and fifty.
JOSH: Is that the same instinct as the article — refusing to accept the locked-down version?
ERIK: Same instinct exactly. If a tool doesn't bend to what I actually need, I don't complain about it, I build the thing that does. That's the whole Echo and Neo setup — two boxes in the lab, twelve agents across the fleet, and none of it exists because a vendor sold it to me. It exists because I got tired of asking permission from software.
JOSH: Does that scale, though? A hundred and fifty services is a lot to keep in your head.
ERIK: It doesn't scale if a human's the one watching it. That's the whole point — PrimeDash isn't for staring at, it's for glancing at. If something's red, PrimeBus already knows before I do. The dashboard's just there for when I want to see why.
[pause]
JOSH: So — Ember-1 chasing efficient reasoning, PrimeBus turning code review into a split between automation and judgment, and PrimeDash refusing to accept software that won't bend. Three different problems, same answer: build the thing that actually fits, don't wait for someone to sell it to you.
ERIK: That's the throughline every week, whether we plan it or not.
JOSH: This episode is sponsored by —
ERIK: This episode is sponsored by Prime Automation Solutions. If you're still doing it manually, we automate it. Also, special on a website — $250. primeautomationsolutions.com.
[pause]
JOSH: Alright, what's the AI pro tip today?
ERIK: Stop sending your model the whole file. When you're working with Claude Code or any agentic coding tool, the biggest hidden cost isn't the reasoning, it's the context you're paying to re-read every turn. Scope the request — point it at the function, not the repo. I do this on every PrimeBus fix generation, and it's a big part of why Hermes stays fast at 70 tokens a second instead of choking on context it doesn't need. Smaller context, cheaper call, faster answer, and honestly a better answer too, because the model isn't guessing which part of a ten-thousand-line file actually matters. That's your tip. Use it.
JOSH: That tip is straight out of The Autonomous Engineer — Erik's book on building systems that run themselves. Grab it on Amazon.
[pause]
JOSH: One more thing — we started a Discord for builders. If you're shipping AI, automation, or anything that makes a human obsolete — come hang out. Link at buildorbereplaced.dev.
ERIK: Post what you built. We'll post what we're building. Real wins, real builds, no fluff.
[pause]
ERIK: Build or be replaced.
JOSH: If you want these signals in your inbox every morning, scanbrief.dev. See you tomorrow.