Mon–Fri · 6 AM ET
← All Episodes
EP  • 00:07:34

Yes, and | Build or Be Replaced

Today: Yes, and | I hired an illustrator to draw my house. Now it's my Home Assistant dashboard | Man discovers his parents' coffee machine used 1TB of data in 10 days Episode date: 2026-10-09.

Download MP3 →

Transcript

ERIK: You're going to want to hear the Claude Code story at the bottom of this one.
JOSH: It's Friday, October 9th. This is Build or Be Replaced — powered by ScanBrief.dev. I'm Josh, here with Erik Anderson.
ERIK: Friday means guardrails did their job all week and nobody had to escalate anything to me. That's the whole theme today.
JOSH: Stick around — Erik's got an AI pro tip at the end about picking the right model for the job, not just the newest one.

[pause]

JOSH: Okay, headlines. First one — OpenAI pulled back three mathematical results they'd published. What happened there?
ERIK: They published results, somebody checked the math, it didn't hold up, and they retracted it. That's it. The story isn't the retraction — it's that nobody caught it before publish.
JOSH: That's wild for a company that size.
ERIK: Verification is still the hard part. Generating a claim is cheap. Checking it is expensive. Everybody's learning that the hard way right now.

[beat]

JOSH: Next — somebody's parents' coffee machine used a full terabyte of data in ten days.
ERIK: That coffee maker had a bug spamming its own internal traffic until it saturated the home Wi-Fi. Nobody audited the device, it just ran until it broke something.
JOSH: A coffee machine took down the internet.
ERIK: Every smart device on your network is a tiny unmonitored service. If you wouldn't ship an API without logging, don't plug in a coffee maker without knowing what it's phoning home.

[beat]

JOSH: Last headline — Ask HN: what do you run on a five dollar VPS that's worth keeping online 24/7?
ERIK: Good question to ask yourself honestly. Most people's answer is "nothing, I just forgot to cancel it." Mine would be a cron job and a health check. If it's not watching something or reacting to something, it doesn't need to be always-on.

[pause]

JOSH: Alright, let's go deeper on one — why isn't the industry freaking out about DeepSeek 4.1 Flash? That headline's been everywhere.
ERIK: Because it's basically frontier-model quality at a fraction of the cost, and that's supposed to be terrifying if you're OpenAI or Anthropic. But from where I sit, it's not scary, it's just another lane.
JOSH: What do you mean, a lane?
ERIK: PrimeRouter is the gateway that sits in front of every model call I make. It doesn't care if the answer comes from Claude, Gemini, or a cheap Ollama Cloud model — it picks based on priority and falls over to the next backend if one's unavailable. Something like DeepSeek 4.1 Flash is exactly the kind of model that slots into a cheap lane for high-volume, low-stakes work.
JOSH: So you'd actually route to it?
ERIK: For the right job, sure. Not for anything protected — I'm not putting a discount model in my execution chain or my code review path without staring at it for a while first. But for a single-shot summarization task where I just need throughput and I'm watching the cost? That's exactly what a cheap-but-capable model is for.
JOSH: How does that compare to just always using the best model?
ERIK: Always using the best model is how you burn money on tasks that don't need it. The point of a router isn't "use the smartest thing available," it's "use the right thing for this specific call, and have a fallback when it's not there." That's the whole argument for multi-provider failover instead of betting the business on one vendor.

[pause]

JOSH: Okay, second deep dive, and this one's my favorite headline of the week — someone says they found a planet nobody knew existed, using Claude Code.
ERIK: Yeah, somebody ran an agentic coding session against astronomy data looking for patterns, and Claude Code ended up surfacing a signal that looked like an undiscovered planet. Whether it holds up to peer review is a separate question, but the workflow is the real story.
JOSH: Why's the workflow the story and not the planet?
ERIK: Because six months ago that's a research team with a grant and a grad student running analysis for months. Now it's one person pointing an agentic coding tool at a public dataset and having it iterate on the hypothesis itself.
JOSH: Is that the same kind of thing you're doing with your agents?
ERIK: Same category, smaller scale. ScanBrief scored 95 items across 56 sources again overnight, average relevance around 52.9, and that ranking isn't me reading every headline — it's an agent doing the scoring so I only look at what's worth looking at. The planet guy did the same trick, just pointed at a telescope instead of a news feed.
JOSH: What does that mean for people who aren't engineers? Does this change what's possible for a regular person with an idea?
ERIK: It means the bottleneck used to be "can you write the code to test this idea," and now the bottleneck is "do you have a good idea worth testing." That's a much more interesting problem to have. I've got twelve agents in the fleet right now — Claude instances and Gandalf doing GPT-side review — and none of them are geniuses individually. They're useful because they run the boring iteration loop without getting tired.

[pause]

JOSH: Let's do one more — Whistle, the speech-to-text model that's under 17 megabytes.
ERIK: That one's underrated. Everybody's chasing bigger models, and this is the opposite bet — how small can you make something and still have it work in real time, fully on-device, no API call.
JOSH: Why does the size matter so much?
ERIK: Because a model that small runs on a Raspberry Pi, a wearable, a robot — anything without a GPU and without a network connection. It's a convolutional front end with a handful of attention blocks, nothing exotic, they just clearly optimized hard for footprint.
JOSH: Do you run anything like that yourself?
ERIK: I run local Ollama backends on Bill, my Mac M3, specifically for stuff I don't want leaving the house or don't want to pay cloud tokens for. A model like Whistle is the same philosophy — not every job needs a frontier model phoning home. Sometimes the win is just "runs anywhere, costs nothing, good enough."
JOSH: How does that compare to the DeepSeek story — isn't that also "cheap and good enough"?
ERIK: Same instinct, different axis. DeepSeek is cheap because of economics — somebody else is eating the compute cost. Whistle is cheap because of engineering — it's just genuinely tiny. Both are telling you the same thing though: 2026 is the year "good enough and local" stops being the compromise option.

[pause]

ERIK: This episode is sponsored by Prime Automation Solutions. If you're still doing it manually, we automate it. Also, special on a website — $250. primeautomationsolutions.com.

[pause]

JOSH: Alright, what's the AI pro tip today?
ERIK: Stop defaulting to your most expensive model for every call. Build yourself even a dumb priority list — cheap model first, escalate on failure or low confidence, save the frontier model for the stuff that's actually hard or actually protected. I run this for real with PrimeBus: 2,295 merge attempts since June, 1,507 went through clean, 788 got blocked by Gandalf before they ever touched prod, and zero had to come to me directly. That's not a 65% success rate story, that's a guardrails-caught-the-rest story. Your model routing should work the same way — cheap by default, escalate on evidence, never skip the check.
ERIK: That's your tip. Use it.

[beat]

JOSH: Binge all five episodes this weekend plus our YouTube shorts — links at buildorbereplaced.dev.

[pause]

JOSH: One more thing — we started a Discord for builders. If you're shipping AI, automation, or anything that makes a human obsolete — come hang out. Link at buildorbereplaced.dev.
ERIK: Post what you built. We'll post what we're building. Real wins, real builds, no fluff.

[pause]

ERIK: Build or be replaced.
JOSH: If you want these signals in your inbox every morning, scanbrief.dev. See you tomorrow.