Transcript
JOSH: It's Tuesday, June 9. This is Build or Be Replaced, powered by ScanBrief.dev. I'm Josh, here with Erik Anderson.
ERIK: Today's theme is simple: AI isn't slowing down. The bottleneck moved to compute, trust, and who owns the workflow.
JOSH: Stick around, Erik's got an AI pro tip at the end about making agents prove their work before they touch production.
[pause]
JOSH: First up, Apple looks like it's rebuilding Apple Intelligence around Google Gemini models. Big deal, or just branding?
ERIK: Big deal. Apple doesn't need the flashiest demo. They need AI that works inside the phone, respects privacy, and can actually touch apps without making a mess.
JOSH: And Siri is apparently getting the real treatment this fall.
ERIK: Siri has been the punchline for years. If Apple gives it context, app actions, visual search, and a sane privacy model, it stops being a voice assistant and starts being a local operator.
JOSH: Next one is weird. xAI is starting to look like a datacenter business.
ERIK: That's the story. Frontier AI is not only about models anymore. It's power, GPUs, buildings, fiber, cooling, and contracts. The lab with the most compute wins more often than the lab with the best tweet.
JOSH: Last headline. Microsoft's open source tools were hacked to steal passwords from AI developers.
ERIK: That one should get everyone's attention. AI builders are wiring keys, tokens, agents, and cloud access together fast. If your dev toolchain gets poisoned, the attacker doesn't need to beat your model. They just ride your credentials.
[pause]
JOSH: Let's start with Apple. Every year people say Siri is fixed. Why is this one different?
ERIK: Because Apple is changing the architecture, not just the voice. That's the part that matters.
[beat]
ERIK: Old Siri was a command parser with a nice microphone. Set a timer. Call my wife. Play a song. It had narrow actions, and it failed in ways that made you stop trusting it.
JOSH: Very familiar.
ERIK: Exactly. Once a tool embarrasses you twice, you stop using it. That's the real Siri problem. Not capability. Trust.
JOSH: So what changes if Gemini-style models are underneath it?
ERIK: You get better reasoning over messy input. You get multimodal support. You get a model that can look at context instead of waiting for perfect phrasing. And if Apple wires it into apps properly, it can take actions across the system.
JOSH: Like what?
ERIK: Pull the tracking number from one email, find the delivery photo, text it to the contractor, and add a reminder if nobody replies by noon. That's not magic. That's workflow glue.
[beat]
ERIK: The trick is Apple has to make it feel boring. Boring is good here. If Siri becomes an agent on a billion phones, the win isn't some wild demo on stage. It's ten tiny actions a day that don't fail.
JOSH: Where does privacy fit into that?
ERIK: That's Apple's best angle. Run what you can on device. Send only what you must to cloud models. Make the boundary visible enough that normal people don't have to read a system diagram to feel safe.
JOSH: You're building a lot of agent stuff. Does this match what you see in your own systems?
ERIK: Completely. PrimeBus processed 866 automation events across 14 projects today. That sounds like a big number, but the reason it works is boring plumbing. Events come in. Agents subscribe. Guardrails check outputs. Actions happen only when the rules say they can happen.
JOSH: That's the phone version?
ERIK: Pretty much. The phone is the bus. Apps are services. Siri becomes the agent layer. The calendar, mail, messages, photos, reminders, files, all of that becomes callable infrastructure.
JOSH: Wait, really?
ERIK: Yeah. People think of phones as apps. Builders should think of phones as local infrastructure with sensors, identity, payment, storage, and user context. Apple owns the best runtime for personal agents if they don't fumble the permissions model.
JOSH: What's the risk?
ERIK: Half measures. If Siri can summarize but can't act, people shrug. If it can act but asks for confirmation on every tiny step, people turn it off. If it acts without proof, people stop trusting it.
[beat]
ERIK: Agent design is not "let the model do stuff." That's how you wake up to weird calendar invites and an apologetic status page. Good agent design is intent, plan, evidence, action, audit trail.
JOSH: That sounds like your production pipeline.
ERIK: It is. The PrimeBus auto-merger has run 367 attempts since 2026-06-05: 282 merged, 85 blocked by Gandalf, 0 escalated to Erik. That's the part I care about. Not just the 76.8% merge rate. The 85 blocked tells me the guardrails are doing work.
JOSH: So Apple needs a Gandalf.
ERIK: Everybody needs a Gandalf. Especially your phone.
[pause]
JOSH: The xAI story is stranger. The headline says it looks more like a datacenter REIT than a frontier lab. What does that mean in normal words?
ERIK: It means compute is becoming the product. Not the model. Not the chatbot. The GPUs.
JOSH: That's wild.
ERIK: It is, but it makes sense. Frontier labs are hungry all the time. Training needs massive clusters. Inference needs massive clusters. Peak usage is brutal. If a company controls enough power and hardware, it can sell capacity to other labs and suddenly it's not only competing with them. It's feeding them.
JOSH: So xAI partnering with Anthropic and Google for compute isn't as weird as it sounds?
ERIK: Not weird at all. It's infrastructure economics. If you have idle capacity, sell it. If someone else has a capacity crunch, buy time. Everyone talks about intelligence like it's floating in the air. It isn't. It's sitting in racks, pulling power, waiting on network fabric.
JOSH: How does that change what builders should care about?
ERIK: Stop treating model choice like religion. You need routing. You need fallbacks. You need telemetry. You need cost controls. The model is one dependency in the path.
[beat]
ERIK: That's why PrimeRouter exists in my stack. I don't want one provider outage turning my systems into a very expensive paperweight. Route by priority. Track latency. Watch token spend. Fail over when the provider gets weird.
JOSH: Is that overkill for a small builder?
ERIK: No. A simple version is not overkill. Even a basic gateway that says "use Claude for code review, use a cheaper model for summaries, fall back if latency crosses a threshold" will save you pain.
JOSH: Where does local AI fit?
ERIK: Local matters more every month. Hermes runs on a Mac M3 with qwen3-coder at 70 tokens per second. That's not replacing every cloud model, but it's enough for private code tasks, small reviews, extraction, routing decisions, and offline work.
JOSH: So cloud for power, local for control?
ERIK: That's the clean split. Cloud models are your heavy machinery. Local models are your shop tools. You don't rent a crane to tighten a bolt.
JOSH: Nice.
ERIK: And you don't send every secret to a cloud model just because the API is easy. Builders need to get serious about data boundaries. What can run local? What needs the frontier model? What needs a human?
JOSH: Is this why companies are acting like compute is oil?
ERIK: Yeah, except oil doesn't ask for more context window every quarter.
[beat]
ERIK: The labs know demand is not linear. Better models create more use cases. More use cases create more inference. More inference creates more infrastructure demand. That's the loop.
JOSH: And if you own the loop?
ERIK: You print options. You can train. You can sell capacity. You can negotiate with other labs. You can survive spikes. Compute is leverage in the plain English sense. It's power.
JOSH: You used the forbidden word.
ERIK: Fair. Compute is a big hammer. Better?
JOSH: Much better.
[pause]
JOSH: The Microsoft toolchain hack feels like the one that should scare regular developers. What happened there?
ERIK: Attackers went after open source tools used by AI developers and stole passwords. That's the simple version. The important part is the target selection.
JOSH: Why AI developers specifically?
ERIK: Because AI developers are connecting everything. GitHub tokens. Cloud credentials. Slack bots. Vector databases. Browser automation. Deployment keys. Agents with permissions. If you compromise that workstation or package path, you might get access to the whole build system.
JOSH: So the agent becomes the blast radius.
ERIK: Exactly. A normal script with bad permissions is bad. An agent with bad permissions is worse because it can discover things, call tools, and chain actions together.
JOSH: How do you defend against that?
ERIK: First, don't give agents broad credentials. Give them scoped tokens. Read-only where possible. Short-lived where possible. Separate dev from prod. Separate review from merge. Separate observation from action.
[beat]
ERIK: Second, log every tool call. If your agent can run Terraform, Selenium, kubectl, NSO actions, or GitHub writes, you need an audit trail. Not vibes. Logs.
JOSH: What should those logs include?
ERIK: Who requested it, what model produced it, what tool ran, what input went in, what output came back, what changed, and what gate approved it. That's minimum.
JOSH: Sounds heavy.
ERIK: It sounds heavy until an agent deletes the wrong thing and your only incident note is "Claude seemed confident."
JOSH: Fair.
ERIK: Third, put review gates where money or production is involved. In my world, Gandalf reviews changes before merge. It blocked 85 auto-merger attempts in the last 7 days, and 0 escalated to me. That's not bureaucracy. That's the system catching risk before I have to.
JOSH: What about developers using random packages?
ERIK: Pin versions. Use lockfiles. Watch dependency changes. Don't install mystery CLIs with your main shell profile loaded. Run suspicious tools in a container. Keep secrets out of env files that every tool can read.
JOSH: People won't love that.
ERIK: People also don't love rotating every credential after a compromised npm package. Pick your pain.
[beat]
JOSH: Does AI make supply chain security harder?
ERIK: Much harder. It increases speed. Speed exposes weak process. If a human used to copy one command from a README, now an agent may run ten commands, inspect files, install packages, and open PRs before coffee.
JOSH: That sounds like the future and a headache.
ERIK: It is both. The answer is not "don't use agents." That's losing. The answer is make agents operate like junior engineers with badges that only open certain doors.
JOSH: What do you mean by badges?
ERIK: Permission tiers. One agent can read logs. Another can propose code. Another can open a PR. A different gate can merge after tests and review. Don't give one model the keys to the building because it wrote a nice summary.
JOSH: How does this connect to network automation?
ERIK: Same pattern. In Cisco NSO, you don't let every workflow push service changes everywhere with no dry run. You validate intent. You check diffs. You commit through controlled paths. AI doesn't remove that discipline. It makes it mandatory.
JOSH: So the boring controls win again.
ERIK: Always. Boring controls are undefeated.
[pause]
JOSH: There was also OpenCV 5 in the feed today. Quick builder take?
ERIK: Computer vision is quietly becoming normal infrastructure. Cameras, OCR, screenshots, UI tests, factory lines, document processing. OpenCV 5 matters because vision isn't only research now. It's another sensor input for automation.
JOSH: Would you wire that into agents?
ERIK: Absolutely. Selenium can drive the browser, but vision can verify what actually rendered. That's useful. DOM says the button exists. Vision says the modal covered it. Big difference.
JOSH: That's painfully specific.
ERIK: Because it happens.
[pause]
ERIK: This episode is sponsored by Prime Automation Solutions. If you're still doing it manually, we automate it. Also, special on a website — $250. primeautomationsolutions.com
[pause]
JOSH: Alright, what's the AI pro tip today?
ERIK: Make your agent produce evidence before it takes action.
[beat]
ERIK: Not a paragraph saying "I checked it." Evidence. Command output. Test results. Diff summary. File paths. API response. Screenshot. Whatever proves the action is safe.
JOSH: Give me the actual prompt.
ERIK: Use this: "Before making any write action, list the exact evidence you have, the command or tool that produced it, the risk if you're wrong, and the smallest reversible action you can take next."
[beat]
ERIK: Then enforce it. If the evidence field is empty, the tool call doesn't run. If the risk is high, route to a human. If the action isn't reversible, require a second review.
JOSH: That's practical.
ERIK: That's how you move fast without pretending the model is magic. Agents don't need trust. They need boundaries and receipts. That's your tip. Use it.
[pause]
JOSH: We also drop daily market picks and automation tips on YouTube — search Build or Be Replaced.
[pause]
JOSH: One more thing — we started a Discord for builders. If you're shipping AI, automation, or anything that makes a human obsolete — come hang out. Link at buildorbereplaced.dev.
ERIK: Post what you built. We'll post what we're building. Real wins, real builds, no fluff.
[pause]
ERIK: Build or be replaced.
JOSH: If you want these signals in your inbox every morning, scanbrief.dev. See you tomorrow.