Transcript
JOSH: It's Wednesday, August 26. This is Build or Be Replaced — powered by ScanBrief.dev. I'm Josh, here with Erik Anderson.
ERIK: The theme today is simple. The AI stack is moving out of the cloud brochure and into real hardware, real guardrails, and real accountability.
JOSH: Stick around — Erik's got an AI pro tip at the end about making agents prove their work before they touch production.
[pause]
JOSH: ScanBrief scored 95 items across 56 sources today. First headline: OpenAI's Jalapeño ASIC is being pitched as better than Nvidia Blackwell for LLM inference. That's a big claim, right?
ERIK: Massive claim. If it's real, the story isn't just faster chips. It's hardware and software being designed together for inference instead of treating GPUs like magic rectangles that solve every problem.
JOSH: Second headline: Apple introduced M6, M5 Ultra, and new Macs around them. Is this just Apple doing Apple, or does it matter for builders?
ERIK: It matters. Local AI is getting less weird. My Hermes backend on a Mac M3 already runs qwen3-coder around 70 tokens a second, so when Apple adds more Neural Engine and memory bandwidth, that changes what you can run at the edge.
JOSH: Third headline: C2PA cameras do not survive contact with reality. That's a pretty brutal headline.
ERIK: Yeah, and it's probably correct. Provenance is useful, but if your trust model depends on the camera, the app, the upload path, and the viewer all behaving perfectly, I have bad news. Humans are involved.
[pause]
JOSH: Let's start with the OpenAI chip. Jalapeño sounds like a fake product name, but the claim is very real: better inference than Blackwell across open-source models. What's the actual signal there?
ERIK: The signal is specialization. Nvidia won because CUDA became the default language of AI infrastructure. Everybody built around it. Training, inference, libraries, clusters, monitoring, procurement. The whole thing grew around GPUs. But inference is not training. Serving models all day is a different job than building the model in the first place.
[beat]
ERIK: If OpenAI taped out an ASIC in about 16 months and it's beating general GPU platforms on LLM inference, that tells you where the money is going. Not into another shiny chatbot. Into the pipes. Tokens are the product now. Cheap tokens, predictable latency, and enough capacity that you don't wake up to a provider outage and start rewriting your app at 2 AM.
JOSH: So this is OpenAI trying to own more of the stack?
ERIK: Exactly. They don't want to rent the entire future from Nvidia forever. Nobody with that much traffic wants that. When you run enough inference, the chip bill becomes strategy. You stop asking, "What API call should I make?" and start asking, "How many watts per useful answer?"
JOSH: That's wild. Most people still think the model is the product.
ERIK: The model is part of the product. The serving layer is where a lot of the pain lives. Routing, retries, context windows, caching, rate limits, fallbacks, billing, telemetry. That's why I built PrimeRouter. I don't want an app to know or care whether a request lands on Claude, GPT, or a local model. It asks for a capability and a priority tier. PrimeRouter handles the rest.
[beat]
ERIK: Same reason PrimeBus exists. This morning the auto-merger stat is 2,295 attempts since June 5: 1,507 merged, 788 blocked by Gandalf, 0 escalated to me. That's not a cute demo. That's infrastructure saying, "I can move fast because the guardrails are doing actual work."
JOSH: Wait, really? Zero escalated to you?
ERIK: Yeah. And the important part is not the merge rate. The important part is the blocked count. People hear "blocked" and think failure. Wrong. That's Gandalf doing the job. A bad automated merge that reaches production is the expensive one. A blocked merge is the system saying, "Nope, not today."
JOSH: How does that connect back to chips?
ERIK: Same principle. The winners aren't going to be the ones with one impressive benchmark. The winners are going to be the ones with the whole system working. Silicon, compiler, model runtime, scheduler, observability, pricing, failover. If Jalapeño is real and general-purpose enough for multiple open-source models, that's OpenAI telling the market, "We're not just buying shovels anymore. We're making the mine."
JOSH: Does Nvidia lose here?
ERIK: Not tomorrow. Nvidia has the datacenter relationships, the software moat, and years of operational trust. But pressure shows up at the edges first. Big customers build custom silicon. Smaller providers rent whatever gives them margin. Builders like us route across providers because no single vendor deserves blind trust.
[beat]
ERIK: The practical takeaway is this: don't weld your product to one model vendor or one inference path. Use an internal gateway. Log the request, the model, the latency, the cost, the output quality. Then switch when the math says switch. Feelings don't pay the GPU bill.
[pause]
JOSH: That brings us right to Apple. New M6, M5 Ultra, new Mac mini, new Mac Studio. You mentioned Hermes. Is local AI finally becoming normal?
ERIK: It's becoming boring, which is better. Boring means useful. When I can run a local coding model on a Mac and get usable speed, that changes architecture. Not every task needs a frontier model. Some tasks need cheap, private, close-to-the-files inference.
JOSH: What's an example?
ERIK: Log triage. Small code edits. Test failure summaries. Drafting a runbook. Classifying incoming events. You don't need the smartest model on Earth to say, "This Selenium test failed because the selector changed." You need something fast, local, and connected to your event bus.
[beat]
ERIK: In my setup, Hermes can sit behind PrimeRouter as another provider. If a task is low-risk, local model. If it's high-risk or needs stronger reasoning, send it to Claude or GPT. If a provider times out, fail over. The app doesn't need a religious opinion about models. It needs an answer with a trace.
JOSH: So the new Macs matter because they make that cheaper?
ERIK: Cheaper, quieter, and easier to own. Apple's unified memory model is a big deal for local inference. The specs in ScanBrief say M6 has a 12-core CPU, 12-core GPU, dual 16-core Neural Engines, and 170 gigabytes per second of unified memory bandwidth. That's not a toy. That's enough machine for serious edge work, especially if you're smart about model size.
JOSH: But most companies still want cloud AI, right?
ERIK: They want cloud AI because it's easy to start. They want local AI the first time legal asks where the data went. Or when the cloud bill lands. Or when a vendor has an outage and the NOC is staring at a blank dashboard like it's going to apologize.
[beat]
ERIK: Local doesn't mean everything runs on a laptop. It means you have choices. A Mac Studio under a desk. A mini server in a branch. A Kubernetes node with a local model. A worker that handles boring tickets before a human sees them. The point is control.
JOSH: How would you decide what runs local and what goes to a hosted model?
ERIK: Three questions. Is the data sensitive? Is the task repetitive? Is the cost of being slightly wrong low? If yes, local is a candidate. Example: summarizing yesterday's alerts. Local. Drafting a customer-facing incident report. Maybe frontier model, maybe human review. Auto-merging code. Absolutely guarded. PrimeBus doesn't get to be brave. It gets to be measured.
JOSH: That's a good line. "Doesn't get to be brave."
ERIK: Bravery in automation is usually somebody forgot the rollback plan.
[pause]
JOSH: The third story is C2PA cameras not surviving contact with reality. This is about proving images are authentic. Why should an automation engineer care?
ERIK: Because provenance is just another trust chain, and trust chains break at the weakest handoff. C2PA tries to attach metadata so you can know where media came from and whether it was edited. Good idea. But real workflows are messy. Screenshots happen. Apps strip metadata. Platforms recompress files. People record screens with another camera. Suddenly your beautiful cryptographic story is a pile of maybe.
JOSH: So you think it's useless?
ERIK: No. I think it's useful evidence, not truth. Big difference. Same with logs. Same with CI results. Same with AI evals. A signed artifact tells you something. It doesn't tell you everything.
[beat]
ERIK: This is where builders get themselves in trouble. They add one trust signal and call the problem solved. "The image has provenance." Cool. What about the capture device? What about the account that uploaded it? What about the chain between camera and viewer? What about edits that are allowed? What about missing metadata? Missing evidence is also evidence.
JOSH: That sounds a lot like your Gandalf setup.
ERIK: It is. Gandalf doesn't trust a single signal. PrimeBus gets telemetry from the repo, tests, review output, runtime events, and policy checks. A merge doesn't happen because one model said, "Looks good." That's how you wake up unemployed. The system wants corroboration.
JOSH: What does corroboration look like for media?
ERIK: Multiple layers. Signed capture when available. Platform attestation when available. Hashes. Account reputation. Time and location consistency. Human review for high-impact cases. And clear UI that says what you actually know. Not "verified true." More like, "Captured by this device, edited by this app, uploaded by this account, no breaks detected after this point."
[beat]
ERIK: Boring wording saves you from lawsuits. Marketing wording gets screenshots in discovery.
JOSH: Dry but fair.
ERIK: Same applies to AI agents. Don't say "the agent fixed it" unless you can show the ticket, the diff, the tests, the reviewer, the merge event, and the deploy result. That's why telemetry matters. Today 140 distinct projects have emitted telemetry to PrimeBus, and 144 services are running on the production server right now. I need to know what happened without reading tea leaves.
JOSH: So provenance is really an operations problem?
ERIK: Yep. Identity, state, policy, audit trail. That's operations. C2PA is trying to bring that thinking to media. Good. But the real world always attacks the clean diagram first.
[pause]
JOSH: Before the tip, sponsor read.
ERIK: This episode is sponsored by Prime Automation Solutions. If you're still doing it manually, we automate it. Also, special on a website — $250. primeautomationsolutions.com
[pause]
JOSH: Alright, what's the AI pro tip today?
ERIK: Make your agent produce a proof bundle before it changes anything important. Not a paragraph. A bundle. Input, plan, files touched, commands run, test output, risk notes, and rollback step. Then have a second model or policy gate read the bundle and approve or block.
[beat]
ERIK: You can do this today with Claude Code, GitHub Actions, and a simple JSON file. Agent writes the patch. Agent writes proof.json. CI runs tests. Reviewer model reads proof.json plus the diff. If anything is missing, it fails closed. That's how you stop vibes from becoming production changes.
[beat]
ERIK: Start with one repo and one rule. "No merge unless tests ran and the rollback is named." You'll catch embarrassing stuff immediately. That's your tip. Use it.
[pause]
ERIK: If you're building toward financial independence through automation, my first book walks through the whole path. Free chapter at erikandersonbook.com.
[pause]
JOSH: One more thing — we started a Discord for builders. If you're shipping AI, automation, or anything that makes a human obsolete — come hang out. Link at buildorbereplaced.dev.
ERIK: Post what you built. We'll post what we're building. Real wins, real builds, no fluff.
[pause]
ERIK: Build or be replaced.
JOSH: If you want these signals in your inbox every morning, scanbrief.dev. See you tomorrow.