Transcript
JOSH: It's Friday, June 5. This is Build or Be Replaced — powered by ScanBrief.dev. I'm Josh, here with Erik Anderson.
ERIK: AI made code cheap. Trust is still expensive.
JOSH: Stick around — Erik's got an AI pro tip at the end about using agents as reviewers, not interns.
[pause]
JOSH: First headline. Anthropic is moving toward an IPO, but enterprise buyers are pushing back on AI spend. Is that the first real crack in the AI gold rush?
ERIK: Cost pressure is real. The people paying the invoices finally found the usage tab, and now every model call has to prove it belongs in the workflow.
JOSH: Second headline. Microsoft Build was packed with agentic AI, new Copilot tooling, and Nvidia's RTX Spark dev box. What matters there?
ERIK: The desktop is becoming an agent workbench. Local compute, cloud models, GitHub, Windows, all stitched together. That's not a demo anymore. That's the new dev machine.
JOSH: Third headline. Nvidia is pitching enterprise software companies on building AI agents with its stack. Hardware company, software layer, agent runtime. Too much?
ERIK: Not too much. Totally expected. Whoever controls the agent runtime controls where the inference bill lands. Nvidia knows that better than anyone.
[pause]
JOSH: Let's start with the Anthropic story. IPO talk and cost backlash at the same time feels messy.
ERIK: Messy, but honest. That's what happens when a technology goes from magic trick to line item. Early AI adoption was, "Duuude, this thing writes code." Now it's, "Why did finance get a bill that looks like a small car?"
JOSH: Wait, really? That's where we are already?
ERIK: Yeah. And it should happen. If an AI agent saves ten engineer-hours, the bill can be big and still be fine. If it's summarizing summaries of meetings nobody needed, kill it. No drama. No committee. Kill it.
JOSH: That sounds simple. But enterprises are not famous for simple.
ERIK: Correct. They buy the tool, skip the instrumentation, then act surprised when nobody can explain the cost. Same thing happened with cloud. Same thing happened with Kubernetes. Everybody wanted elastic everything until the invoice showed up wearing steel-toed boots.
JOSH: What should they measure?
ERIK: Outcome per model call. Not tokens by themselves. Tokens are fuel. You don't judge a truck by gallons burned unless you know what it hauled.
JOSH: Give me the builder version.
ERIK: If Claude edits a Terraform module, runs tests, opens a patch, and a human approves it, that cost is tied to a deliverable. If an agent reads Slack all day and writes "noted" with better grammar, that's expensive noise.
JOSH: How do you handle that in your own systems?
ERIK: PrimeBus treats AI work like event-driven infrastructure. A failure hits the bus. A reviewer agent looks at the event. A fix agent gets scoped context. A verifier runs tests. PrimeSentinel watches for stuck jobs. Nobody gets unlimited context just because the model asked nicely.
JOSH: That's the guardrail.
ERIK: That's the budget guardrail too. The fastest way to waste money with AI is to let every agent read everything. That's not intelligence. That's a buffet with a company credit card.
[beat]
JOSH: There's a trust angle too, right? Anthropic is known for safety. But buyers are asking about spend.
ERIK: Trust includes spend. That's the part people miss. If I can't predict what a system costs, I don't trust it. If I can't trace what it changed, I don't trust it. If I can't roll it back, I don't trust it. Safety isn't only "did the model say something spooky." Safety is "can this thing operate inside my business without wrecking the place."
JOSH: That's less glamorous.
ERIK: Better. Glamour gets you a keynote. Boring controls keep production alive.
JOSH: How does this compare to network automation?
ERIK: Same pattern. In Cisco NSO, you don't let a workflow blast configs across devices with no dry run, no diff, no rollback, and no owner. AI agents should be held to at least that standard. Probably higher, because they can sound confident while doing something very dumb.
JOSH: That one hurts.
ERIK: It should. Confidence is not correctness. Claude is good. I build with Claude every day. But my systems don't trust the sentence. They trust the evidence. Tests passed. Diff is small. Scope is right. Service health is clean. Then we move.
[pause]
JOSH: Microsoft Build is the next one. Agentic AI everywhere. Copilot app, developer tooling, runtime talk, local AI hardware. What's the actual signal?
ERIK: Microsoft wants the agent to sit where the work already happens. That's the signal. Not a chatbot off to the side. Agent in the editor. Agent in GitHub. Agent near Windows. Agent near your files. Agent near your terminal.
JOSH: That sounds useful and also kind of dangerous.
ERIK: Both. Useful tools are usually dangerous. A shovel is useful. Still don't hand it to someone in the data center and say, "Go explore."
JOSH: Nice visual.
ERIK: Accurate visual. The mistake is treating an AI coding agent like an intern. Interns need mentoring, but they usually don't have shell access, repo access, cloud credentials, and the confidence of a senior architect after three prompts.
JOSH: Then what is it?
ERIK: Treat it like a junior automation engine with a change window. Give it a ticket. Give it the repo slice. Give it the test command. Give it the allowed files. Make it produce a diff. Make something else review it.
JOSH: Something else meaning another agent?
ERIK: Sometimes. A reviewer agent is good at being picky if you prompt it right. I like separate roles. Builder agent makes the patch. Reviewer agent attacks it. Verifier agent runs the boring checks. PrimeBus coordinates the event flow. Nobody gets to grade their own homework.
JOSH: That's the pro tip tease, isn't it?
ERIK: Pretty much. The end version has fewer words and more yelling.
[beat]
JOSH: Where does local compute fit into this? Nvidia RTX Spark, dev boxes, all that.
ERIK: Local compute matters for latency, privacy, and cheap iteration. Not every task needs the biggest cloud model. Some tasks need fast local classification. Is this log event worth waking up Claude? Is this diff risky? Is this alert a duplicate? Smaller local models can answer that and save the expensive model for the hard part.
JOSH: So local isn't about replacing the frontier model.
ERIK: Not for me. It's about triage. My ScanBrief pipeline is a good mental model. Gather sources, dedupe, score, rank, then spend attention on what matters. Same with agents. Don't call the monster model for every tiny decision. Put a bouncer at the door.
JOSH: A bouncer at the door is a very technical architecture term.
ERIK: It should be in more diagrams.
JOSH: How would a normal team start?
ERIK: Pick one workflow that already has a clean pass-fail. Pull request review. Terraform plan review. Ansible runbook generation. NSO service package test fixes. Don't start with "AI, run my company." Start with "AI, inspect this diff and tell me what can break."
JOSH: And then?
ERIK: Log everything. Prompt, context, model, cost, output, test result, human decision. If you don't log it, you can't improve it. If you can't improve it, you're not building automation. You're doing vibes with an invoice.
JOSH: That's going on a mug.
ERIK: Send me one.
[pause]
JOSH: Nvidia pitching enterprise agents feels like the other side of this. They're not just selling GPUs. They're selling the factory around the GPUs.
ERIK: Exactly. GPU vendors learned from cloud vendors. The money isn't only in the metal. It's in the stack that keeps the metal busy.
JOSH: What does that mean for builders?
ERIK: Model choice is going to get more tied to infrastructure choice. That's the part to watch. If your agent runtime, vector store, eval pipeline, and deployment system all live inside one vendor's lane, moving later gets painful.
JOSH: Is that bad?
ERIK: Not automatically. Vendor lock-in is fine when you're getting paid more than you're trapped. That's the honest answer. If Nvidia's stack gets your product shipped and customers pay you, use it. But know where the exits are.
JOSH: What are the exits?
ERIK: Standard interfaces. Store events in a format you control. Keep prompts in your repo. Keep evals portable. Don't bury business logic inside a hosted agent builder where the export button is decorative.
JOSH: Decorative export button. Painfully believable.
ERIK: Seen it too many times. In networking, we learned this with controllers. The GUI is friendly until you need the thing it doesn't support. Then you're staring at an API doc at midnight wondering who hurt you.
JOSH: How does HumanRail fit here?
ERIK: HumanRail is built around the idea that agents should be able to ask for help without pretending. If confidence is low, route to a person. If policy says human approval is required, stop and ask. That's not weakness. That's how you keep the system moving without giving the model a fake badge.
JOSH: So humans are still in the loop.
ERIK: Humans are in the right loop. Big difference. I don't want a human clicking approve on every tiny thing. That's checkbox theater. I want humans where judgment matters. Risky diff. Weird customer request. Legal language. Production blast radius. The rest should move.
JOSH: Where does PrimeSentinel come in?
ERIK: PrimeSentinel watches the machinery. Agentic systems fail in boring ways. Job stuck. Queue growing. Worker alive but not doing work. Selenium session hung. NATS consumer lagging. The demo never shows that part. Production absolutely will.
JOSH: That's the part nobody wants to keynote.
ERIK: Because it doesn't sparkle. But stuck-job detection is how you sleep. If the system can't tell you it's jammed, it's not autonomous. It's abandoned.
[beat]
JOSH: There's also a policy headline this week. The U.S. moved toward a shorter review period for advanced AI models. Does that matter to builders, or is that just Washington doing Washington?
ERIK: It matters because regulation will hit the release process. Thirty days, ninety days, whatever the number is, the bigger point is model releases are becoming governed events. System cards, evals, safety reviews, enterprise contracts. The cowboy phase is ending for the big labs.
JOSH: Does that slow everyone down?
ERIK: Big labs, yes. Builders, not necessarily. If anything, it helps small teams because the model layer gets more stable and documented. You can build on top of it instead of waking up to surprise behavior changes every Tuesday.
JOSH: But don't builders need the newest model?
ERIK: Sometimes. Usually they need the right model. Big difference. Use the best reasoning model for hard planning. Use a cheaper model for classification. Use local for routing. Use deterministic code for anything that should never be creative.
JOSH: Deterministic code. Old fashioned.
ERIK: Still undefeated. If a regex solves it, don't summon a frontier model wearing a cape.
JOSH: How do you decide when an agent should touch production?
ERIK: Same way I decide for automation. Lab first. Small scope. Clear rollback. Evidence before action. In my world, Cisco NSO taught that lesson hard. Service changes need a model, a diff, a commit path, and a rollback path. AI doesn't get a pass because it used nice words.
JOSH: That feels like the whole episode in one sentence.
ERIK: Good. Put it on the invoice.
[pause]
JOSH: Before the tip, give me the Friday builder takeaway. What should someone change after hearing all this?
ERIK: Put accounting, review, and routing into your agent flow this weekend. Not next quarter. This weekend. Track model cost per task. Split builder and reviewer roles. Add one stop condition where a human must approve. Then run it on a boring workflow first.
JOSH: Boring workflow first.
ERIK: Always. Boring workflows are where money hides. Password reset runbooks. Config drift checks. Pull request summaries. RFP scanning. GovContractScanner pulling items overnight is not sexy, but finding the right opportunity before everyone else sees it is useful.
JOSH: And useful beats flashy.
ERIK: Every day. Flashy gets applause. Useful pays the server bill.
[pause]
ERIK: This episode is sponsored by Prime Automation Solutions. If you're still doing it manually, we automate it. Also, special on a website — $250. primeautomationsolutions.com
[pause]
JOSH: Alright, what's the AI pro tip today?
ERIK: Use agents as reviewers, not interns. That's the tip.
JOSH: Give me the version someone can use today.
ERIK: Take a pull request, a Terraform plan, or a runbook change. Don't ask the model, "make this better." That's vague. Ask it to review the change against five rules.
JOSH: What kind of rules?
ERIK: Rule one, identify blast radius. Rule two, list assumptions. Rule three, find missing tests. Rule four, call out security or credential risk. Rule five, decide whether a human must approve before merge.
JOSH: That's very specific.
ERIK: Specific is the whole point. Then make the output structured. JSON if you're wiring it into automation. A short checklist if a human reads it. Keep the builder model separate from the reviewer model if you can. If you can't, at least use a separate prompt and make it adversarial.
JOSH: Adversarial how?
ERIK: Tell it, "Your job is to block unsafe changes, not be helpful." Helpful models wave things through. Reviewer models need to be a pain. That's their job.
JOSH: And this works without a huge platform?
ERIK: Yes. Start with one script. Read the diff. Send it to Claude. Save the review. Fail the CI job if it returns a blocking issue. That's it. Then improve it when it catches something real. That's your tip. Use it.
[pause]
JOSH: Binge all five episodes this weekend plus our YouTube shorts — links at buildorbereplaced.dev.
[pause]
JOSH: One more thing — we started a Discord for builders. If you're shipping AI, automation, or anything that makes a human obsolete — come hang out. Link at buildorbereplaced.dev.
ERIK: Post what you built. We'll post what we're building. Real wins, real builds, no fluff.
[pause]
ERIK: Build or be replaced.
JOSH: If you want these signals in your inbox every morning, scanbrief.dev. See you tomorrow.