Transcript
JOSH: It's Monday, August 24. This is Build or Be Replaced — powered by ScanBrief.dev. I'm Josh, here with Erik Anderson.
ERIK: The signal today is simple. Cheap AI is winning attention, but systems with guardrails are winning production.
JOSH: Stick around — Erik's got an AI pro tip at the end about writing an agent.md that actually keeps your coding agents out of the ditch.
[pause]
JOSH: First headline. Anthropic's best model is apparently struggling to attract users while cheaper tools keep growing. Is this just price pressure?
ERIK: Price matters, but workflow matters more. If your best model is amazing but expensive, people save it for hard problems. The daily driver wins the habit.
JOSH: Second. A post about learning to build LLMs from scratch blew up on Hacker News. Should builders go that deep?
ERIK: Yes, if they're serious. You don't need to train frontier models in your basement, but understanding transformers, tokens, attention, evals, and inference cost changes how you build with Claude or GPT.
JOSH: Third. Someone used Claude Opus 5 to reverse engineer everyday devices and get shells, exploit firmware, even mess with recording indicators. That's not comforting.
ERIK: Nope. That's the next security problem. AI doesn't just write web apps. It reads firmware, explains weird protocols, and turns dusty hardware bugs into weekend projects.
[pause]
JOSH: ScanBrief scored 78 items across 56 sources this morning. The cheap-AI story jumped out first. Why?
ERIK: Because it's the market telling us something builders already know. Capability alone doesn't win. Availability wins. Cost wins. Latency wins. The tool people can run 200 times a day becomes the tool they trust.
JOSH: But if the premium model is better, shouldn't people want the best one?
ERIK: For some work, absolutely. I use the strongest model when the blast radius is high. Architecture changes. Security review. Weird failure analysis. But for daily edits, summaries, migrations, test repair, and boring glue code, the cheaper model usually gets the call first.
JOSH: So the best model becomes the specialist.
ERIK: Exactly. Think about networking. You don't send every packet through a senior engineer. You use routing tables, runbooks, telemetry, and escalation paths. Same thing with AI. Small model handles the normal case. Strong model handles the weird case. Human handles the business call.
[beat]
JOSH: That sounds like your PrimeBus setup.
ERIK: It is. PrimeBus doesn't treat every failure like a board meeting. Events come in. Agents subscribe. The system picks a path. The live stat this morning is 2,295 auto-merger attempts since June 5. 1,507 merged. 788 blocked by Gandalf. Zero escalated to me.
JOSH: Wait, really? Zero escalated?
ERIK: Zero. And the blocked count is the good part. That's not failure. That's guardrails doing work. Gandalf says no when a patch smells wrong, when tests don't prove enough, or when the diff touches something it shouldn't.
JOSH: So cheaper models can do more work if the system around them is strict.
ERIK: That's the whole point. People keep arguing model versus model. Wrong argument. The system around the model is where production happens. Queues. Tests. permissions. review gates. telemetry. rollback. That's where trust is built.
JOSH: What does that mean for a team choosing AI tools right now?
ERIK: Don't buy one premium seat and call it an AI strategy. Build a routing strategy. Cheap model for repeatable work. Strong model for hard judgment. Agent rules in the repo. CI that actually runs. A human approval point for anything with money, customer data, or prod access.
JOSH: And if they don't have that?
ERIK: Then they're basically letting a very confident intern SSH into prod. Bold choice. Not my favorite.
[pause]
JOSH: The second story is the one about learning LLMs from scratch. Hacker News loved it. You're an automation guy, not an AI researcher. Why does that matter to your world?
ERIK: Because the API hides the machine, but the bill and the bugs still show up. If you understand how the model works, you stop asking it impossible things. You learn why context matters. Why examples matter. Why tool calls fail. Why a short prompt can be worse than no prompt.
JOSH: Do you think every builder needs to learn the math?
ERIK: Not all of it. But learn enough to stop being a passenger. Know what tokens are. Know why attention gets expensive. Know the difference between training, fine-tuning, retrieval, and just stuffing a giant prompt with junk. Those are different tools.
JOSH: What's the practical version of that?
ERIK: Build a tiny transformer tutorial. Run a local model once. Break it. Watch latency change when context gets bigger. Then go back to Claude or GPT and your prompts get better because now you understand the shape of the thing you're talking to.
[beat]
JOSH: How does that show up in your own systems?
ERIK: BookForge is a good example. Writing a chapter is not one prompt. That's amateur hour. The system has outline state, style rules, continuity checks, citations when needed, editorial passes, and rejection criteria. The model is one worker inside the shop.
JOSH: So the model isn't the product.
ERIK: Right. The process is the product. Same with the PAS website-pipeline. It doesn't just say, build a website, good luck. It gathers business context, generates copy, builds pages, runs checks, then hands me the parts that need taste or pricing judgment.
JOSH: Where do people get this wrong?
ERIK: They ask one giant prompt to be designer, developer, QA, security reviewer, copywriter, and project manager. Then they act shocked when it drops something. Duuude, you gave one brain seven jobs and no checklist.
JOSH: That's fair.
ERIK: Build smaller loops. One agent writes. One reviews. One runs tests. One checks policy. One decides if the result can merge. That's how you turn AI from a chat window into infrastructure.
JOSH: Does learning LLM internals make someone better at that?
ERIK: Yes. Because you stop thinking in magic. You start thinking in inputs, state, memory, tools, and verification. That's engineering.
[pause]
JOSH: Third deep dive. The reverse-engineering story. Everyday hardware getting popped with help from Claude Opus 5. Is that hype, or is that a real shift?
ERIK: Real shift. Hardware hacking used to have a nasty entry fee. Datasheets, firmware dumps, obscure protocols, weird tooling, hours of staring at binary output. AI lowers that pain. It doesn't make a beginner elite overnight, but it turns a motivated person into a much faster problem finder.
JOSH: The story mentioned things like microphones, webcams, lights, firmware, memory over Wi-Fi. That's a lot.
ERIK: That's the lesson. The attack surface isn't just your laptop anymore. It's every cheap device with firmware, Wi-Fi, USB, Bluetooth, and an update process someone forgot about.
JOSH: And there was also that Android automotive head unit malware story.
ERIK: Same family of problem. Cars are rolling computers with entertainment stacks, app stores, Bluetooth, USB, network paths, and sometimes very questionable vendor firmware. When that firmware gets infected upstream, your security team doesn't even see the beginning of the story.
JOSH: What does that mean for normal companies?
ERIK: Asset inventory has to include weird stuff. Conference room cameras. badge printers. kiosks. TV panels. head units in fleet vehicles. Anything that runs code and touches the network deserves a name, an owner, and a patch story.
[beat]
JOSH: That's not how most IT inventories look.
ERIK: Nope. Most inventories are laptops, servers, cloud accounts, maybe phones. Then there's a cabinet full of mystery devices using default creds and firmware from 2021. Very normal. Very bad.
JOSH: How would you handle it?
ERIK: Put discovery on a schedule. Nmap, DHCP logs, switch MAC tables, wireless controller data, vulnerability scans, whatever you have. Feed it to a bus. When a new device appears, classify it. If it can't be classified, quarantine or alert. Don't wait for someone to remember the lobby display exists.
JOSH: That's where PrimeRestorer comes in too, right?
ERIK: Different side of the same discipline. PrimeRestorer exists because backups are not vibes. Databases, system state, S3 offsite, restore paths. The question is not do we have backups. The question is did we prove restore works. Security without restore is theater.
JOSH: You sound annoyed by this one.
ERIK: Because people buy tools and skip the boring part. The boring part saves you. Asset names. owners. restore tests. allowlists. blocked merges. boring logs. My production server has 144 services running right now. If I don't have telemetry, naming, and guardrails, that's not engineering. That's a pile.
JOSH: And the AI angle makes the pile riskier.
ERIK: Yes. AI helps defenders, but it helps curious attackers too. The difference is process. If your process is stronger than theirs, you're fine. If your process is a shared spreadsheet and hope, you're going to have an educational week.
[pause]
JOSH: There's also a thread in today's list about complex systems failing. That feels connected.
ERIK: It is. Complex systems don't fail because one part is evil. They fail because normal things line up. A firmware bug. A missing owner. A weak update path. A noisy alert nobody reads. A backup nobody tested. Then Tuesday happens.
JOSH: So the answer isn't one bigger alert?
ERIK: Please no. The answer is small controls everywhere. Make failure boring. PrimeBus blocks bad merges before I see them. ScanBrief ranks signals before I read them. PrimeRestorer checks backup paths before I need them. That's the shape.
JOSH: Smaller loops, stronger gates.
ERIK: Exactly. And write it down. If your AI agent can't read the rule, your AI agent can't follow the rule. Same for humans, honestly.
[pause]
ERIK: This episode is sponsored by Prime Automation Solutions. If you're still doing it manually, we automate it. Also, special on a website — $250. primeautomationsolutions.com
[pause]
JOSH: Alright, what's the AI pro tip today?
ERIK: Put an agent.md file in every repo. Not a cute manifesto. A working instruction file.
JOSH: What goes in it?
ERIK: Commands first. How to install. How to test. How to lint. How to run one focused test. Then architecture notes. Where the app starts. Where shared helpers live. What files are generated. What folders are off limits.
JOSH: That's it?
ERIK: Add merge rules. Example: never touch migrations without a test. Never edit secrets. Never change public API names without updating docs. Run tests before claiming done. If tests fail, report the exact command and failure.
JOSH: So it's a runbook for the agent.
ERIK: Yes. And keep it short. Agents read long documents like humans read Terms of Service. Badly. Give them the commands and the boundaries. Then make CI enforce the rest.
JOSH: What changes after you add it?
ERIK: Fewer dumb diffs. Faster repairs. Less babysitting. Your agent stops guessing how your repo works. That's your tip. Use it.
[pause]
ERIK: That tip is straight out of The Autonomous Engineer — my book on building systems that run themselves. Grab it on Amazon.
[pause]
JOSH: One more thing — we started a Discord for builders. If you're shipping AI, automation, or anything that makes a human obsolete — come hang out. Link at buildorbereplaced.dev.
ERIK: Post what you built. We'll post what we're building. Real wins, real builds, no fluff.
[pause]
ERIK: Build or be replaced.
JOSH: If you want these signals in your inbox every morning, scanbrief.dev. See you tomorrow.