Mon–Fri · 6 AM ET
← All Episodes
EP  • 00:12:57

Bun is being ported from Zig to Rust | Build or Be Replaced

Today: Bun is being ported from Zig to Rust | Google Chrome silently installs a 4 GB AI model on your device without consent | How OpenAI delivers low-latency voice AI at scale Episode date: 2026-05-05.

Download MP3 →

Transcript

JOSH: It's Tuesday, May 5. This is Build or Be Replaced powered by ScanBrief.dev. I'm Josh, here with Erik Anderson.
ERIK: Thirty-six items hit ScanBrief this morning, PrimeBus already chewed through 2,155 events, and the pattern is the same every day now: the teams building systems are pulling away from the teams still clicking buttons.
JOSH: Stick around. Erik's got an AI pro tip at the end about getting better output by splitting one prompt into three jobs.
[pause]
JOSH: First headline. Bun is moving from Zig to Rust. Big engineering move or nerd drama?
[beat]
ERIK: Big engineering move. When a runtime that serious rewrites core pieces, that's not style points. That's a bet on safety, hiring, tooling, and surviving the next five years.
[beat]
JOSH: Next one. Chrome silently dropping a four gig AI model onto user devices. Is that as bad as it sounds?
[beat]
ERIK: Worse, because the problem isn't only size. It's consent. If your browser installs a model, re-downloads it after deletion, and doesn't make the contract obvious, you've got a trust problem before you've got an AI feature.
[beat]
JOSH: Third headline. OpenAI talking about low-latency voice at scale. Why should builders care?
[beat]
ERIK: Because voice stops being a toy when latency gets low enough that people stop noticing the machine. Once that happens, support, dispatch, field ops, triage, all of it changes fast.
[pause]
JOSH: Let's stay on that Chrome story, because people heard four gigs and immediately thought, wait, what is my browser doing.
[beat]
ERIK: Yeah. That's the right reaction. A browser is supposed to render pages, manage sessions, run extensions, and stay out of your way. When it starts acting like a silent model deployment platform, the bar changes. You don't get to say, well, it's for helpful features. That's not the point. The point is the install happened without a clean, obvious choice.
[beat]
JOSH: So what's the actual problem? Storage? Privacy? Battery?
[beat]
ERIK: All three. Storage is the easiest one to see. Four gigs is not a rounding error on normal laptops. Battery matters if that thing gets invoked locally. But privacy and trust are the real issue. If a vendor quietly places model weights on a device, users are going to ask what else is happening quietly. And they should.
[beat]
JOSH: Some people will say, hold on, local AI is better than sending everything to the cloud.
[beat]
ERIK: Local can be better. I run local. Hermes is on a Mac M3 in my stack pushing qwen3-coder around 70 tokens a second. I like local inference. I like keeping jobs in my own walls. But local by choice and local by stealth are not the same thing. If I deploy Hermes, I know exactly why it's there, what port it's on, what jobs hit it, and what box pays the power bill. Chrome doing a surprise drop is the opposite of that.
[beat]
JOSH: So this isn't anti-local. It's anti-sneaky.
[beat]
ERIK: Exactly. Put a clear setting in front of me. Tell me the size. Tell me the feature set. Tell me when it downloads, when it runs, and how to remove it for real. Then we're adults. Right now it sounds like product people got excited and skipped the part where the user is supposed to matter.
[beat]
JOSH: What does this mean for people building apps on top of these browsers?
[beat]
ERIK: It means you need to assume users are getting more suspicious, not less. If you're building browser-based AI features, you have to over-communicate. Show your work. Say what runs local. Say what goes remote. Say what gets cached. Say what gets deleted. Because if the platform burns trust, your app inherits the smoke.
[beat]
JOSH: That's a rough inheritance.
[beat]
ERIK: It is. And it's avoidable. In ops, hidden automation is how you get paged at 2 AM. Same rule here. If a system makes changes out of band, people stop trusting the system. PrimeRestorer taught me that. Backups aren't useful because they exist. They're useful because I know when they ran, what they covered, where the copies landed, and whether restore tests passed. Silent behavior is fake reliability.
[beat]
JOSH: That sounds like the bigger lesson.
[beat]
ERIK: It is the lesson. Hidden convenience becomes visible pain later. Every time.
[pause]
JOSH: On the OpenAI voice story, I think normal people hear low latency and think, cool, less awkward pauses. You're hearing something else.
[beat]
ERIK: I'm hearing the interface change. Text is patient. Voice is not. With text, users tolerate seconds. With voice, they start judging at a few hundred milliseconds. That means the whole stack has to behave differently. Audio chunking, turn detection, interruption handling, model handoff, backpressure, retry logic, all of it matters.
[beat]
JOSH: In plain English?
[beat]
ERIK: If the machine talks too slow, people talk over it or hang up. Then your great model doesn't matter.
[beat]
JOSH: Better.
[beat]
ERIK: Good. Low-latency voice at scale is a systems story, not just a model story. Everybody wants to talk about the magic sentence the AI says. I care about the plumbing. Can it stream partials? Can it recover from jitter? Can it keep context without sounding drunk? Can it do it for thousands or millions of sessions at once?
[beat]
JOSH: And can it do it cheaply.
[beat]
ERIK: Also that. Cheap matters. Always. If your demo works at ten calls and dies financially at ten thousand, you built a hobby. Not a product.
[beat]
JOSH: Where do you see this landing first?
[beat]
ERIK: Internal ops. Not consumer assistants. Builders should start where the value is obvious and the blast radius is controlled. NOC triage. Help desk routing. Field dispatch. After-hours call intake. Runbook lookup. Those are high-friction, repetitive flows where shaving thirty seconds actually compounds.
[beat]
JOSH: Because you've already got structure there.
[beat]
ERIK: Right. A lot of that work already has scripts, queues, escalation paths, humans on standby. Voice becomes the front door, not the brain. That's important. People keep trying to make voice the whole product. Bad move. Voice should collect intent, confirm context, route to tools, and get out of the way.
[beat]
JOSH: That sounds a lot like how you build your systems.
[beat]
ERIK: Yeah. PrimeBus works because every service doesn't try to do everything. One job per worker. Publish events. Subscribe where needed. Keep contracts tight. If I were building voice ops, I'd have the transcript service publish into NATS, an intent classifier subscribe, then a job runner trigger Terraform, NSO, Selenium, whatever the task needs. HumanRail or a similar gate sits there when confidence drops. That's boring architecture. Boring architecture wins.
[beat]
JOSH: You're saying the flashy part is the least important part.
[beat]
ERIK: The flashy part is what gets clipped into a demo video. The boring part is what keeps it alive in production. I've got 89 services running on the production server right now. None of them care about vibes. They care about contracts, retries, logs, and whether the next message showed up.
[beat]
JOSH: What does low-latency voice change for a smaller builder? Somebody with one product, not sixty-four projects and a home lab that sounds mildly haunted.
[beat]
ERIK: It lowers the threshold for useful automation. A small builder can now put a voice layer on top of an existing workflow without hiring a call center or building a giant speech stack from scratch. But they still need guardrails. Keep the scope narrow. Don't start with open-ended conversation. Start with five intents. Maybe ten. Status check. Password reset. Order lookup. Appointment confirm. Escalate. That's enough to print money if the flow was painful before.
[beat]
JOSH: Wait, really?
[beat]
ERIK: Yeah. Builders waste time chasing sophistication when friction is sitting right in front of them. Remove one expensive manual step and you get paid. That's the whole trick.
[beat]
JOSH: So if somebody hears this and says, I want a voice agent tomorrow, your answer is not go buy the biggest model.
[beat]
ERIK: Correct. My answer is map the workflow, measure the call path, define the failure path, then pick the model. Model selection comes after job definition. Not before. If the job is fuzzy, the result will be fuzzy and expensive.
[pause]
JOSH: The third one I want to hit is Agent Skills. Because a lot of people are feeling this already. The AI writes code fast, and then the team spends the next day cleaning it up.
[beat]
ERIK: That's exactly the problem. Speed without process looks amazing for six minutes. Then you own the mess. Agents are very willing to skip the parts senior engineers rely on: specs, tests, code review, design checks, rollback plans. The machine doesn't feel shame, so you have to install it manually.
[beat]
JOSH: That's a sentence.
[beat]
ERIK: It's true. I use Claude every day. Multiple instances. PrimeBus fans work out across the mesh. One agent can draft a fix, another can review, another can run comparison logic. But the reason that system works is I don't let one output become production truth by default.
[beat]
JOSH: Give me the real version. What happens in your stack?
[beat]
ERIK: PrimeBus has logged 214 auto-fixes with a 78 percent success rate. People hear that and focus on the 214. I focus on the remaining 22 percent. Those are the cases that prove whether your guardrails are real. Bad test? Rejected. Weak confidence? Route it. Two variants both pass but differ materially? Compare and hold. That machinery matters more than the first draft.
[beat]
JOSH: So Agent Skills is basically trying to force that discipline back into the loop.
[beat]
ERIK: Yeah. And it needs to happen. Right now too many teams are treating coding agents like interns who somehow also have root. Terrible idea. If the agent can write code, it also needs a path through spec validation, test execution, policy checks, and ideally a second opinion.
[beat]
JOSH: That's where people push back. They say, if I add all that process, I lose the speed.
[beat]
ERIK: Then they don't understand the math. Raw generation is not the bottleneck anymore. Correction is the bottleneck. Rework is the tax. If the agent saves you twenty minutes writing code and costs you three hours in regressions, congrats, you bought a faster shovel and dug the hole deeper.
[beat]
JOSH: Dry, but fair.
[beat]
ERIK: Stripe's overnight formatting story connects to this too. At huge scale, the move isn't heroics. It's building controlled machinery that can touch a giant surface area safely. Same pattern. Small trusted steps. Strong rollback. Measurable outcomes. That's how you operate with agents too.
[beat]
JOSH: Where do smaller teams usually get this wrong first?
[beat]
ERIK: They ask the model for implementation before they force a crisp definition of done. That poisons everything downstream. Then they skip test scaffolding because the code looks plausible. Then they let the same agent review its own work, which is comedy. Then they act surprised when prod gets weird.
[beat]
JOSH: How do you fix that without building a giant platform team?
[beat]
ERIK: Start stupid simple. One prompt writes the spec. Second prompt writes the code against the spec. Third prompt acts like a reviewer and only looks for defects, missing tests, and contract violations. Then run the tests. Then decide. That's already better than most setups I see.
[beat]
JOSH: And that's still fast.
[beat]
ERIK: Very fast. Also cleaner. ScanBrief is the same shape, by the way. Fifty-six sources come in noisy. The system doesn't trust the first thing it sees. It scores, dedupes, ranks, summarizes. That's why the brief is useful at 6:30 in the morning instead of looking like a junk drawer with RSS tags.
[beat]
JOSH: So the pattern across all three stories is basically this: the winners are the ones with systems, not stunts.
[beat]
ERIK: That's it. Chrome shows what happens when power outruns trust. Voice shows what happens when the plumbing gets good enough to change behavior. Agent Skills shows the next correction in coding: less magic, more discipline. Builders should pay attention to all three.
[pause]
ERIK: This episode is sponsored by Prime Automation Solutions. If you're still doing it manually, we automate it. Also, special on a website — $250. primeautomationsolutions.com
[pause]
JOSH: Alright, what's the AI pro tip today?
[beat]
ERIK: Stop asking one prompt to do three jobs. Split it. First prompt: write the spec and edge cases. Second prompt: implement only against that spec. Third prompt: review like a grumpy senior engineer and look for breakage, missing tests, and bad assumptions. Different jobs, different mindset. If you've got agents, wire them that way. If you're solo in Claude, do it in sequence. You'll get less pretty nonsense and more code you can actually ship. That's your tip. Use it.
[pause]
JOSH: We also drop daily market picks and automation tips on YouTube — search Build or Be Replaced.
[pause]
JOSH: One more thing — we started a Discord for builders. If you're shipping AI, automation, or anything that makes a human obsolete — come hang out. Link at buildorbereplaced.dev.
ERIK: Post what you built. We'll post what we're building. Real wins, real builds, no fluff.
[pause]
ERIK: Build or be replaced.
JOSH: If you want these signals in your inbox every morning, scanbrief.dev. See you tomorrow.