Transcript
JOSH: It's Friday, June 12. This is Build or Be Replaced — powered by ScanBrief.dev. I'm Josh, here with Erik Anderson.
ERIK: AI agents are cheap until you hand them a cloud account and no leash.
JOSH: Stick around — Erik's got an AI pro tip at the end about putting hard spend limits around agent work before the agent gets creative.
[pause]
JOSH: First headline. An AI agent tried to scan DN42 and left its operator with a six thousand five hundred thirty one dollar AWS bill. Erik, that's not a bug, that's a mortgage payment.
ERIK: That's what happens when autonomy meets an unlimited credit card. The agent did what it was asked to do. The system around it failed.
JOSH: Next one. Hacker News is arguing that if you're asking for human attention, you need to show human effort. Is that just manners, or is it now an engineering requirement?
ERIK: It's engineering hygiene. If Claude wrote the patch, say Claude wrote the patch. Add what you checked, what you believe, and what you need reviewed. Don't dump a mystery box on another engineer and call it collaboration.
JOSH: Third headline. An old paper is back: nobody gets credit for fixing problems that never happened. That's painfully familiar.
ERIK: Preventative work is invisible until it stops happening. PrimeBus processed 968 automation events across 16 projects today, and the whole point is that most of them should be boring.
[pause]
JOSH: Let's start with the AWS bill story, because that's the one everyone is going to remember.
ERIK: Yeah. An agent tried to join and index DN42, which is this decentralized network lab people use for BGP, tunnels, experiments, that kind of thing. It spun up cloud resources, kept going, and the operator got hit with a $6,531.30 AWS bill.
JOSH: Wait, really?
ERIK: Really. And the part people miss is that the agent wasn't evil. It wasn't conscious. It wasn't sitting there twirling a mustache. It was given a goal, access to infrastructure, and not enough guardrails.
JOSH: So the failure is the box around the agent.
ERIK: Exactly. Agentic systems need budgets the same way production systems need CPU limits, rate limits, and rollback plans. If your Kubernetes cluster can stop a container from eating the node, your AI runner should stop a task from eating your AWS account.
[beat]
JOSH: What does that mean in practical terms?
ERIK: Hard limits. Not vibes. Set AWS budgets. Set service quotas. Use separate accounts. Use IAM roles that only allow the exact actions needed. Put agents in a sandbox account first. Then make the agent ask for approval before it crosses a spend threshold or creates resource types it hasn't used before.
JOSH: That sounds obvious when you say it.
ERIK: It is obvious. That's why people skip it.
JOSH: That's cold.
ERIK: It's true. Everyone wants the cool demo. Nobody wants to wire the breaker panel. Then the demo becomes a financial incident.
[pause]
JOSH: How do you handle that in your own systems?
ERIK: PrimeBus is event-driven, but it doesn't mean every event gets a blank check. Gandalf reviewed and merged 17 code changes to production overnight, but that pipeline has gates. Since June 5, the auto-merger has run 637 attempts: 438 merged, 199 blocked by Gandalf, 0 escalated to me. That's the story. The blocked count means the guardrails worked.
JOSH: Zero escalated is the number that jumps out.
ERIK: Right. That's what I want. Not because the system is magic. Because the system has a narrow job, clear tests, separate agents, logs, review rules, and stopping conditions. The model can suggest. It can patch. It can test. It can merge if the evidence passes. It doesn't get to improvise production policy at 3 AM.
JOSH: And this AWS agent basically improvised production policy.
ERIK: That's how I read it. The operator said go do network discovery, and the surrounding system didn't say, cool, but your max spend is fifty bucks, your max runtime is two hours, and if you create anything expensive, stop and ask.
JOSH: Is this an AI problem or just an automation problem wearing a new hat?
ERIK: Both. Old automation broke things fast because scripts are dumb and literal. AI automation breaks things creatively because the model can find a path you didn't imagine. That's useful. That's also why it needs tighter walls.
[beat]
JOSH: What's the wall you would build first?
ERIK: Money. Always. Cloud budget alarm is not enough. Alarm means it already happened. You need budget enforcement, quota enforcement, and kill switches. If the agent starts ten instances when it usually starts one, stop it. If it touches a region it has never touched before, stop it. If the estimated cost crosses the job budget, stop it.
JOSH: And then a human reviews?
ERIK: Or another system reviews. HumanRail is the pattern, even if you don't use my system. When confidence drops or risk goes up, route it to a human. Don't make every task human-in-the-loop. That's slow. Make the dangerous transitions human-in-the-loop.
[pause]
JOSH: Second deep dive. Human attention. The post says if you're asking for attention, demonstrate human effort. Why did that hit a nerve?
ERIK: Because code review is already expensive. Now people are pasting AI-generated patches, AI-generated summaries, AI-generated explanations, and asking a human to untangle what nobody owned.
JOSH: So the issue isn't using AI.
ERIK: No. Use AI all day. I build with Claude every day. The issue is pretending AI output is the finished thought. If you send me a pull request, I want to know what changed, why it changed, what you ran, what you're unsure about, and what you want me to focus on.
JOSH: Label the machine part.
ERIK: Yes. Say, Claude generated the first pass. I reviewed the database migration. Tests passed locally. I'm not sure about the retry behavior. Please look there. That's a good review request.
JOSH: What's the bad version?
ERIK: Here's a thousand lines. Thoughts?
[beat]
JOSH: I felt that one.
ERIK: Everyone has seen it. It's not new, but AI made the volume worse. Before, a lazy request took time to write. Now a lazy request can be beautifully formatted in ten seconds.
JOSH: That's the trap, right? It looks polished.
ERIK: Polished is not the same as inspected. I don't care if the summary has bullets and nice wording. I care if the person sending it knows where the sharp edges are.
[pause]
JOSH: How do you avoid that in an automated review pipeline?
ERIK: Make the agents produce evidence, not prose. PrimeBus doesn't just say, looks good. It emits events. Tests ran. Files changed. Risk category. Gandalf decision. Merge decision. Block reason. If something fails, the event trail is there.
JOSH: So the review isn't a vibe check.
ERIK: Correct. And with email-reactor, same idea. A client emails a website change request. The system can parse it, draft the change, and prepare the update, but it has to keep the original request tied to the diff. Otherwise you're trusting a summary of a summary. That's how little mistakes turn into embarrassing invoices.
JOSH: What does that mean for teams using AI in code reviews right now?
ERIK: Set a house rule. AI-generated work is allowed, but it must include human notes. What did you verify? What did you not verify? What changed from the model's first version? Where should the reviewer spend attention?
JOSH: That sounds like it would cut review noise fast.
ERIK: It does. It also makes junior engineers better faster, because they have to form an opinion. The model can help them write code, but it can't be the only one who understands the code.
[beat]
JOSH: There's a cultural piece there too.
ERIK: Big time. If you ask for my attention, you're spending my time. Show me you spent some of yours. That's not gatekeeping. That's respect for the queue.
[pause]
JOSH: Third story. The old paper about not getting credit for problems that never happened. Why is that showing up again now?
ERIK: Because prevention is becoming the main engineering job. The loud work used to be outages, war rooms, dashboards going red. The valuable work is making sure that never starts.
JOSH: But companies reward the firefighter.
ERIK: They do. Person runs into the burning building, everyone claps. Person installs the sprinkler system, nobody notices. Engineering has the same sickness.
JOSH: How do you prove the value of prevention then?
ERIK: You instrument it. You count what was blocked, what was caught, what was automatically repaired, what never reached a human. ScanBrief scored 90 items across 54 sources today. Most don't become show topics. That's not wasted work. That's filtering doing its job.
JOSH: Same logic as Gandalf blocking merges.
ERIK: Exactly. A blocked merge isn't a failure. It's a prevented incident. Same with PrimeRestorer. Nobody gets excited about backups until the database is gone. Then suddenly the backup orchestrator is the most important system in the building.
[beat]
JOSH: How do you make that visible without turning it into dashboard theater?
ERIK: Tie prevention to decisions. Don't make a wall of charts nobody reads. Send the right signal to the right place. A weekly report that says 193 risky auto-merger runs were blocked in the last 7 days is useful. A dashboard with fourteen colors and no action is office wallpaper.
JOSH: That's a very specific complaint.
ERIK: I've seen some dashboards, man.
JOSH: Fair.
ERIK: The point is, preventative work needs receipts. If your runbook saved an outage, log it. If your agent stopped a bad deploy, record why. If your backup test passed, keep the restore time. Engineers don't get credit for ghosts. Give the ghost a ticket number.
[pause]
JOSH: How does this apply outside infrastructure? Say someone is building a small app or solo business.
ERIK: Same rules. Add guardrails before you add features. If InkEngine is generating book drafts, I care about versioning, prompt records, source files, and output checks. If ScanBrief is ranking stories, I care about source diversity and scoring. If email-reactor is changing a client site, I care about approval states.
JOSH: You're saying boring systems first.
ERIK: Boring systems make the fun stuff possible. Nobody wants to hear that, but it's true. You want agents shipping while you sleep? Cool. Then give them logs, budgets, tests, rollback, and a way to stop themselves.
JOSH: That sounds like the theme of the day.
ERIK: Yeah. Autonomy is not permission to be sloppy. Autonomy is a reason to be more strict.
[pause]
ERIK: This episode is sponsored by Prime Automation Solutions. If you're still doing it manually, we automate it. Also, special on a website — $250. primeautomationsolutions.com
[pause]
JOSH: Alright, what's the AI pro tip today?
ERIK: Before you run an agent against anything real, write a one-page operating contract. Not a giant policy doc. One page.
JOSH: What's in it?
ERIK: Inputs, allowed tools, forbidden actions, max runtime, max spend, stop conditions, and who gets paged. Then paste that into the agent prompt and enforce the same rules outside the prompt with actual limits.
[beat]
ERIK: Prompt rules are not security. They're instructions. Use them, but back them with code. If the agent can create cloud resources, put a budget in the cloud account. If it can edit files, run tests before merge. If it can send email, require approval before external send.
JOSH: So the tip is prompt plus guardrail.
ERIK: Exactly. Tell the agent the rules, then build the system so the rules still matter when the model gets weird. That's your tip. Use it.
[pause]
JOSH: Binge all five episodes this weekend plus our YouTube shorts — links at buildorbereplaced.dev.
[pause]
JOSH: One more thing — we started a Discord for builders. If you're shipping AI, automation, or anything that makes a human obsolete — come hang out. Link at buildorbereplaced.dev.
ERIK: Post what you built. We'll post what we're building. Real wins, real builds, no fluff.
[pause]
ERIK: Build or be replaced.
JOSH: If you want these signals in your inbox every morning, scanbrief.dev. See you tomorrow.