Transcript
JOSH: It's Wednesday, October 7. This is Build or Be Replaced — powered by ScanBrief.dev. I'm Josh, here with Erik Anderson.
ERIK: Today's theme is simple: stop asking one giant model to do every job. That's lazy architecture.
JOSH: Stick around — Erik's got an AI pro tip at the end about using cheap decision gates before you wake up the expensive agent.
JOSH: [pause]
JOSH: First headline: Google released EmbeddingGemma 2, a multimodal embedding model built for on-device use. Why should builders care?
ERIK: Because memory just moved closer to the machine. Text, images, audio, and video can land in the same search space without calling a frontier model for every little comparison.
JOSH: Second headline: OpenAI's Decisions API is in public beta, powered by GPT-6 Luna. Typed decisions, routing, scoring, classification. Is that boring in a good way?
ERIK: Very good way. Boring is where automation becomes production. I don't need poetry from a gatekeeper. I need a yes, a no, a score, and a reason code.
JOSH: Third headline: AI safety and regulation are back in the news, with federal attention on OpenAI, Anthropic, and other labs.
ERIK: That's the bill coming due. If companies want agents touching money, infrastructure, health, and code, people are going to ask who gets blamed when the agent does something dumb.
JOSH: [pause]
JOSH: Let's start with EmbeddingGemma 2. Most people hear embeddings and think search. Is that still the point?
ERIK: Search is the easy explanation. The better explanation is this: embeddings let software compare meaning without making a model reason from scratch every time.
ERIK: [beat]
ERIK: That matters because most automation doesn't need deep reasoning. It needs recognition. "Have I seen this ticket before?" "Does this log look like last Tuesday's BGP flap?" "Is this screenshot basically the same broken layout the client complained about last month?"
JOSH: So the model isn't the worker. It's more like the index.
ERIK: Yeah. It's the memory layer. And when it's multimodal, it gets way more useful.
ERIK: Think about email-reactor. A client sends an email saying the button looks wrong. Maybe they attach a screenshot. Maybe the wording is vague because normal humans don't talk like Jira tickets. A multimodal embedding system can compare the screenshot, the email text, the site history, and previous fixes before Claude gets pulled in.
JOSH: That sounds less flashy than an agent, but more useful.
ERIK: Exactly. Agents get the attention because they look alive. Embeddings are the filing system that keeps the agent from acting like it was born five seconds ago.
JOSH: That's a good way to put it.
ERIK: Without memory, every agent is a new intern with admin rights. Bad combo.
JOSH: On-device is the other piece. Why does that matter?
ERIK: Cost, speed, and privacy. Pick the order based on your pain.
ERIK: In network automation, you don't casually ship configs, device names, IP space, topology notes, and incident logs to whatever API looked cool on Hacker News. You need local processing where possible. You need redaction. You need policy checks. Then maybe you call a bigger model.
JOSH: That's the enterprise angle.
ERIK: It's also the home lab angle. PrimeDash watches my lab services. PrimeBus moves events around. ScanBrief ranks tech signals. Those systems should understand their own history locally. If every similarity check turns into an API call, you built a tax meter.
JOSH: So the stack is local retrieval first, then model call later.
ERIK: Right. Local embeddings first. Cheap decision model second. Expensive agent last.
ERIK: That's the architecture. Not because it sounds elegant. Because it keeps the system fast enough to use and cheap enough to leave running.
JOSH: Where do people mess this up?
ERIK: They treat the biggest model as the default. Every log line goes to the smartest thing they can afford. Every email becomes a full agent loop. Every support ticket gets a long prompt with twelve paragraphs of instructions and a prayer at the end.
ERIK: [beat]
ERIK: That's not engineering. That's burning money with YAML.
JOSH: Fair.
ERIK: The smarter pattern is routing. If the issue matches a known pattern, run the known workflow. If the confidence is low, ask a smaller model to classify it. If the blast radius is real, then call the big model or a human.
JOSH: That's infrastructure, not a demo.
ERIK: Yep. Demos love chat boxes. Infrastructure loves queues, typed outputs, retries, NATS, logs, and dashboards.
JOSH: [pause]
JOSH: That takes us right into the Decisions API. Why is a dedicated decision endpoint different from just prompting a model to choose A, B, or C?
ERIK: Because production systems hate mushy outputs. A chat model will answer your question, explain itself, add a caveat, apologize for the caveat, then maybe wrap the answer in Markdown because it felt cute.
JOSH: I have seen that.
ERIK: Everybody has. A decision endpoint is built around the thing the system actually needs: classification, scoring, routing, approval, rejection.
ERIK: I want something like: action equals retry, confidence equals high, reason equals test failure matched known pattern. That's useful. That can drive a workflow.
JOSH: Give me the PrimeBus version.
ERIK: PrimeBus gets events from my systems. Tests fail, services warn, agents finish jobs, reviewers post results. The bus doesn't need a motivational speech. It needs to know what happens next.
ERIK: Does this failure match a known runbook? Did the fix touch a risky file? Did the reviewer catch a real issue or hallucinate one? Should the patch be blocked, retried, or sent to me?
JOSH: And you can make each one a small decision.
ERIK: Exactly. Small gates. Strong logs. Clear contracts.
ERIK: If the decision comes back typed, I can throw it into PrimeDash and see the pattern. I can replay it. I can compare versions. I can ask why a gate blocked something yesterday and passed a similar thing today.
JOSH: That's a big difference from reading chat transcripts.
ERIK: Huge difference. Chat transcripts are evidence after the fact. Typed decisions are control points.
JOSH: Wait, really?
ERIK: Really. If you're building agentic systems, the words are not the product. The state transitions are the product.
JOSH: Say that again.
ERIK: The state transitions are the product. An agent is only useful if it moves work from one safe state to another safe state.
ERIK: Open ticket to classified. Classified to patched. Patched to tested. Tested to reviewed. Reviewed to shipped or blocked. That's the system.
JOSH: Where does Claude fit in that?
ERIK: Claude is still great when I need reasoning, code edits, explanation, or tradeoff analysis. I use Claude because it can actually work through messy context.
ERIK: But I don't need Claude to decide if a file extension is allowed. I don't need Claude to decide if tests returned exit code zero. I don't need Claude to decide if a ticket is about billing, DNS, or CSS.
JOSH: Use the small thing for the small job.
ERIK: Yep. And more importantly, don't let the expensive agent be the only source of truth. The agent proposes. The pipeline decides. The tests decide. The guardrails decide.
JOSH: That's a little less magical.
ERIK: Good. Magic is expensive to debug.
JOSH: [pause]
JOSH: The third story is safety and regulation. This keeps coming back. Builders hear that and tune out because it sounds political. Should they?
ERIK: No. They should pay attention, but not panic.
ERIK: Safety sounds abstract until your agent has write access. Then it's very concrete.
ERIK: Can it send an email? Can it change a route policy? Can it merge code? Can it spend money? Can it delete customer data? Those are not philosophy questions. Those are permissions questions.
JOSH: So the builder version of AI safety is access control.
ERIK: Access control, audit trails, blast radius, rollback, human approval, and boring logs.
ERIK: The big labs can argue about frontier risk. Fine. In the trenches, the question is: what can this system touch, and how do I know what it did?
JOSH: That maps to network automation pretty cleanly.
ERIK: Completely. Network engineers learned this the hard way years ago. You don't give every script full access to every device and hope the regex had a good childhood.
JOSH: That's specific.
ERIK: It's true. You stage changes. You diff configs. You validate. You run pre-checks. You run post-checks. You keep rollback plans. Cisco NSO, Ansible, Terraform, Kubernetes, same idea. Desired state is nice. Guarded state is better.
JOSH: How does that compare to AI agents?
ERIK: Agents need the same discipline. Maybe more.
ERIK: A traditional script is dumb in a predictable way. An agent is smart in an occasionally weird way. That means the guardrails have to be explicit.
ERIK: PrimeBus is built around that idea. Events move through the bus. Workers subscribe. Reviewers check. Gates block risky output. The architecture assumes the agent can be wrong.
JOSH: That's not anti-AI.
ERIK: Not even close. That's how you use AI more. If I trust the guardrails, I can let the system do more work while I'm asleep.
ERIK: The people who ignore safety end up babysitting their automation. That's the worst of both worlds. You paid for the robot and still got the night shift.
JOSH: Dry but accurate.
ERIK: Very accurate.
JOSH: What should a smaller team copy from that?
ERIK: Start with permissions. Don't start with prompts.
ERIK: Give the agent read access first. Then let it propose changes. Then let it open pull requests. Then maybe let it merge low-risk changes after tests pass. You earn write access one boundary at a time.
JOSH: That's slower than the hype version.
ERIK: It's faster in the long run because you don't spend Friday night explaining why your AI assistant renamed half the database columns.
JOSH: That happened to somebody.
ERIK: Definitely. Maybe not today, but definitely.
JOSH: [pause]
JOSH: Bring these three stories together for me. Embeddings, Decisions API, regulation. What's the pattern?
ERIK: The agent stack is splitting into layers.
ERIK: Embeddings are memory. Decision models are traffic cops. Frontier models are heavy reasoning. Policy and audit are the brakes. Queues and event buses are the nervous system.
ERIK: If you wire those together, you get automation that can run. If you don't, you get a chat box with confidence issues.
JOSH: And the winner isn't just the person with the biggest model.
ERIK: Correct. The winner is the person with the best system around the model.
ERIK: That's why I keep building around PrimeBus, ScanBrief, email-reactor, and PrimeDash. Different jobs. Different signals. Same lesson. The model is a component. The system is the advantage.
JOSH: That sounds like the show title.
ERIK: Yeah. Build or be replaced. Subtle branding. Very tasteful.
JOSH: [pause]
ERIK: This episode is sponsored by Prime Automation Solutions. If you're still doing it manually, we automate it. Also, special on a website — $250. primeautomationsolutions.com
JOSH: [pause]
JOSH: Alright, what's the AI pro tip today?
ERIK: Put a cheap decision gate in front of your expensive agent.
ERIK: Here's the exact pattern. Before you call Claude, GPT, Gemini, whatever, ask a smaller decision step three questions.
ERIK: First: what type of work is this? Bug, content, support, code review, research, billing, unknown.
ERIK: Second: what is the risk level? Read-only, low-risk write, customer-facing, money-moving, infrastructure-changing.
ERIK: Third: what should happen next? Ignore, route, retrieve context, ask a human, or call the big model.
JOSH: That's easy to add.
ERIK: Very easy. Use structured output. JSON schema. Enum fields. No cute prose.
ERIK: Then log every decision. After a week, you'll see where your automation is wasting money. You'll see which jobs never needed the big model. You'll see which jobs need a human every time.
ERIK: That's how you build an agent that gets cheaper and safer as it runs. That's your tip. Use it.
JOSH: [pause]
ERIK: If you're building toward financial independence through automation, my first book walks through the whole path. Free chapter at erikandersonbook.com.
JOSH: One more thing — we started a Discord for builders. If you're shipping AI, automation, or anything that makes a human obsolete — come hang out. Link at buildorbereplaced.dev.
ERIK: Post what you built. We'll post what we're building. Real wins, real builds, no fluff.
JOSH: [pause]
ERIK: Build or be replaced.
JOSH: If you want these signals in your inbox every morning, scanbrief.dev. See you tomorrow.