Transcript
JOSH: It's Thursday, May 7th. This is Build or Be Replaced — powered by ScanBrief.dev. I'm Josh, here with Erik Anderson.
ERIK: Today we're talking about the gap between looking like you know what you're doing and actually knowing.
JOSH: Stick around — Erik's got an AI pro tip at the end about picking the right model for the right job.
[pause]
JOSH: First up — the US Library of Congress just added SQLite to its recommended storage formats list. Is this actually significant?
ERIK: Very. The Library of Congress lists formats they trust to still be readable in 50 years. They're not recommending it because it's trendy. Single file, no server required, fully documented, open spec. If you've been going back and forth between JSON, CSV, and SQLite for long-term data storage, that debate is over. Federal endorsement.
[beat]
JOSH: Valve open-sourced the hardware design files for the Steam Controller. CAD files, Creative Commons. What's the play?
ERIK: They dropped STP and STL files — the native CAD format and the 3D print format. Non-commercial, attribution required. That scope is enough for the community to build replacement parts, better grips, modded accessories. Valve gets a longer product life without spending a dollar on it. Smart move wrapped in generosity.
[beat]
JOSH: And Google just relaunched reCAPTCHA as something called Fraud Defense. What changed?
ERIK: reCAPTCHA was built to separate humans from bots. That problem is mostly solved. Fraud Defense is built for a web where AI agents are transacting alongside humans — buying tickets, filling forms, making API calls. It classifies which of three actors is operating: human, bot, or AI agent. The fact that Google shipped this is an admission that the agentic web isn't coming. It's already here.
[pause]
JOSH: Alright, there was a piece on Hacker News today I want to dig into. Title: "Appearing productive in the workplace." The argument is that generative AI lets non-experts produce output that's indistinguishable from expert work. That feels like something you have strong opinions about.
ERIK: The piece splits this into two cases. A novice mimicking a senior in the same field — that's the obvious one. You find out eventually when something breaks and they can't debug it. The more dangerous case is someone producing work in a domain they have no background in at all. A project manager generating network architecture diagrams. An analyst writing security policies. Nobody in the room knows what good looks like.
JOSH: And the output looks completely authoritative.
ERIK: Polished, confident, grammatically correct. The model doesn't signal uncertainty the way a junior engineer does. It just delivers. And if the reviewer doesn't have the domain knowledge to catch what's missing, the decision gets made on something that looks like expert output but has no grounding.
[beat]
JOSH: You use AI to move faster than anyone else. How is what you're doing different from this?
ERIK: I know when the output is wrong. That part doesn't come with the tool. It comes from years of breaking things in production, debugging at 2am, watching systems fail in ways you didn't predict. The model can generate a PrimeBus event routing config that looks perfect. I can tell in thirty seconds if the retry logic is off. Someone who's never built a message bus can't make that call.
JOSH: So the AI accelerates execution but doesn't replace the judgment.
ERIK: The judgment is what you're paying for. Always was.
[beat]
JOSH: How do you protect your own systems from this? You're running AI that reviews and merges code automatically.
ERIK: Guardrails in the pipeline. I run an automated code reviewer — Gandalf — between every change and production. Since mid-April the auto-merger has run 997 times. 384 changes made it to production. 602 got blocked.
JOSH: Wait — 602 blocked. That's not a problem?
ERIK: That's the system working. A 38.5% merge rate means the guardrails are doing their job. If I was seeing 90% merges, I'd be auditing the reviewer. Eleven cases escalated to me directly — those were the genuinely ambiguous ones where Gandalf couldn't make the call. Every block either caught a real issue or enforced an explicit rule.
JOSH: What does it actually catch?
ERIK: Asyncio task handling, SQL race conditions, retry logic that could burn API credits in a loop. Every time Gandalf blocks something I haven't seen before, I write a rule and drop it in the fix-guides directory for that project. Next service I build starts with those lessons already baked in. You can't skip that part and call it automation.
[beat]
JOSH: That's the thing the article is missing — most companies don't have a Gandalf.
ERIK: Most companies have humans doing code review after they've already looked at a polished output. ScanBrief scored 37 items this morning across 54 sources. If the relevance scoring was misconfigured, every brief downstream would be wrong — and it would arrive in your inbox looking perfect. The format doesn't tell you if the signal is bad.
JOSH: So the lesson is: build the review layer before you give AI the surface area.
ERIK: Build the cage first. Then let it run.
[pause]
JOSH: There's a related piece today — "Vibe coding and agentic engineering are getting closer than I'd like." What's the actual convergence?
ERIK: Vibe coding is describe-and-generate. You tell the model what you want, it writes code, you run it. Agentic engineering is the model taking autonomous action — executing commands, modifying files, making decisions in a loop without a human reviewing each step. They've been separate workflows. The new LLM toolkit in that piece starts merging them. Same interface, now with autonomous execution baked in.
JOSH: Why is that a problem if the individual pieces work fine?
ERIK: Because the person who's been vibe coding doesn't know what constraints an autonomous agent needs. They've been treating the model as a code generator. Now it's a code executor. That's a different thing entirely. The surface area for unintended behavior goes up fast.
[beat]
JOSH: Give me a concrete example of what goes wrong.
ERIK: Early on I had an automation touching files in a directory I hadn't explicitly excluded from its scope. Nothing blew up, but it could have. That was the lesson. Now every agent has an explicit whitelist — what it can read, what it can write, what it can delete. Anything outside that list, it stops and escalates. I've got 93 services running in production right now. A runaway agent with undefined scope is a bad day.
JOSH: So you learned this before the tools were this capable.
ERIK: PrimeBus processed 2,285 automation events today across nine projects. Overnight, one code change went through Gandalf and merged while I was asleep. That's the system working. But it only works because I defined the scope before I deployed anything. Echo is my dev environment — everything new runs there first. When it breaks on Echo I write a rule. When it passes Echo it goes to prod with the guardrails already tested.
JOSH: How does someone starting with these new tools avoid the trap?
ERIK: Answer four questions before you deploy anything autonomously. What can it read? What can it write? What can it delete? What does it do when it's not sure? If you can't answer all four, you're not ready. The tools moving fast doesn't mean your governance has to lag. Write the rules first.
[beat]
JOSH: The author said "closer than I'd like." You sound less worried than they are.
ERIK: The capability isn't what worries me. I want more of it. What worries me is people deploying capability without the architecture to hold it. That's not an AI problem — that's just engineering. You test in the lab before you go to production. That rule didn't change because the thing doing the work is a language model.
JOSH: Same boring engineering answer.
ERIK: It's always the same boring answer. Nobody wants to hear it until something breaks.
[pause]
ERIK: This episode is sponsored by Prime Automation Solutions. If you're still doing it manually, we automate it. Also, special on a website — $250. primeautomationsolutions.com
[pause]
JOSH: Alright, what's the AI pro tip today?
ERIK: Model selection. Most people default to the biggest, most capable model for everything. That's expensive and slow on the wrong tasks. You need tiers. Fast, cheap model for high-frequency, low-stakes work — routing, classification, quick summaries. Heavy model for architecture questions, ambiguous errors, anything where being wrong costs real time or money. In PrimeRouter I route by priority tier. Routine calls hit Sonnet. Anything involving retry logic, failure analysis, or a decision I'd want to audit — that escalates to Opus. I define the tiers before I write a line of code. That's what keeps token costs flat even as the system grows. Define your tiers first. That's your tip. Use it.
[pause]
JOSH: Track your freedom score and net worth with the Freedom Blueprint app — free download, link in the show notes.
[pause]
JOSH: One more thing — we started a Discord for builders. If you're shipping AI, automation, or anything that makes a human obsolete — come hang out. Link at buildorbereplaced.dev.
ERIK: Post what you built. We'll post what we're building. Real wins, real builds, no fluff.
[pause]
ERIK: Build or be replaced.
JOSH: If you want these signals in your inbox every morning, scanbrief.dev. See you tomorrow.