macOS Container Machines | Build or Be Replaced
Today: macOS Container Machines | Claude Fable 5 | Upcoming breaking changes for npm v12 Episode date: 2026-06-10.
Download MP3 →
Build or Replaced
Today: macOS Container Machines | Claude Fable 5 | Upcoming breaking changes for npm v12 Episode date: 2026-06-10.
Download MP3 →JOSH: It's Wednesday, June 10. This is Build or Be Replaced — powered by ScanBrief.dev. I'm Josh, here with Erik Anderson. ERIK: AI is getting better, package managers are getting stricter, and if your pipeline can't handle both, it's already behind. JOSH: Stick around — Erik's got an AI pro tip at the end about making agents prove their work before they touch prod. [pause] JOSH: First headline. Claude Fable 5 showed up in the feed today, and ScanBrief ranked it right near the top. Is this another model launch, or something bigger? ERIK: Bigger. The signal is the class name. Mythos-class means Anthropic is changing how they describe capability, not just bumping a model number. When naming shifts, product boundaries usually shift right behind it. JOSH: Second one. npm v12 has breaking changes coming for installs. Why should builders care? ERIK: Because install scripts are a supply-chain mess. npm making Git and remote dependency resolution opt-in by default is annoying for lazy builds and great for real systems. July 2026 sounds far away until your CI breaks on a Friday. JOSH: Third headline. A German ruling says Google can be liable for false answers in AI Overviews. That feels like a big line in the sand. ERIK: It is. The era of “the model said it, not us” is ending. If you put AI output in front of users, you own it. Same as automation. Same as code. Same as a bad runbook at 2 AM. [pause] JOSH: Let's start with Claude Fable 5. The feed called it a new Mythos-class AI with strong software engineering and research performance. What's the real signal? ERIK: The real signal is not “new model good.” That's boring. The signal is Anthropic is carving out a new lane for agent work. [beat] ERIK: Current Claude naming is musical. Opus, Sonnet, Haiku. Fable and Mythos sound different. That matters because names are product architecture. You don't change the shelf label unless the shelf changed. JOSH: What would that mean in practice? ERIK: It probably means longer-running work, higher agency, and stronger tool use. Not just chat. Not just “write me a function.” More like, “take this repo, inspect the failures, propose two fixes, run the tests, explain the risk, then stop at the approval gate.” JOSH: That sounds like what you're already doing with PrimeBus. ERIK: Exactly. PrimeBus processed 1073 automation events across 18 projects today. That's not a toy demo. That's my actual systems throwing events on a NATS bus, agents reading those events, and workflows firing based on what happened. [beat] ERIK: The important part is not the model. The important part is the contract around the model. Claude can be brilliant and still need boundaries. My agents don't get to freestyle against prod because they feel confident. Gandalf reviews the change. Tests run. Telemetry gets emitted. The bus records the whole thing. JOSH: And you had code changes merge overnight, right? ERIK: Nineteen code changes were automatically reviewed by Gandalf and merged to production overnight. That's the number that matters. Not a benchmark. Not a tweet. Real code, reviewed by an agent, shipped while I was asleep. JOSH: Wait, really? ERIK: Yeah. And the auto-merger has run 445 attempts since June 5. 331 merged, 114 blocked by Gandalf, 0 escalated to me. That's the story. The blocked count is not failure. That's guardrails working. JOSH: So when people say the new model is safer or more capable, you don't trust that by itself. ERIK: Never. Model safety cards are useful, but production safety comes from the system around the model. Permissions. Logs. Test gates. Rollback paths. Human escalation. That boring stuff is what lets you move fast without turning your infrastructure into a science project. [beat] ERIK: Fable 5 can be amazing. Great. Put it behind PrimeRouter. Route low-risk tasks to cheaper models. Send high-risk code review to the strongest model. Fail over when a provider gets weird. Track every call with OTEL. Now you have engineering instead of vibes. JOSH: How do you compare that to old automation? Like NSO, Terraform, or scripts? ERIK: Old automation is deterministic. NSO says, “given this service model, push this config.” Terraform says, “given this state, converge this infrastructure.” That's good. Keep that. [beat] ERIK: Agentic automation is different. It handles ambiguity. A test failed. The error is ugly. The dependency changed. The README lied. That's where Claude shines. But you still wrap it in deterministic rails. The agent can propose. The system decides what gets deployed. JOSH: That's the difference between building with AI and just asking it for answers. ERIK: Exactly. Asking AI questions is fine. Building with AI means the answer becomes an event, the event becomes a workflow, and the workflow leaves evidence. If it can't leave evidence, it doesn't touch anything important. [pause] JOSH: Let's move to npm v12. The headline says Git and remote dependency resolution become opt-in by default. Why is that a big deal? ERIK: Because `npm install` has been too trusting for too long. Developers treat install like a harmless download. It's not. It's code execution wearing a hoodie. JOSH: That's a pretty dark hoodie. ERIK: Accurate hoodie. Package installs can run scripts. Remote dependencies can change. Git URLs can point at things you didn't audit. One compromised dependency and your build machine becomes the softest target in the room. JOSH: So npm is tightening the default. ERIK: Right. npm v12 is saying, “you want this remote thing? Prove it.” That will break some builds. Good. Broken builds are feedback. Silent compromise is worse. [beat] ERIK: Builders should upgrade to npm 11.16.0 or later now and start testing the stricter behavior before July 2026. Don't wait until the release lands. Run it in CI. Find the packages that depend on sketchy install behavior. Replace them or explicitly approve them. JOSH: Is this just a JavaScript problem? ERIK: No. JavaScript is just louder because npm is everywhere. Python has the same class of issue. Containers have it. GitHub Actions have it. Terraform providers have it. Any system that pulls code from the internet and runs it during build has a supply-chain problem. JOSH: What does that mean for a small team? ERIK: Small teams need fewer dependencies and stricter gates. That sounds boring, but it's how you stay alive. Pin versions. Use lockfiles. Block unknown post-install scripts. Cache known-good artifacts. Scan dependency diffs like you scan code diffs. [beat] ERIK: In my world, PrimeRouter doesn't get a dependency update because “npm said yes.” It goes through review. Same with ScanBrief. Same with InkEngine. If a package changes behavior, I want a system catching that before it becomes a morning surprise. JOSH: Where does AI fit into that? ERIK: This is a perfect agent job. Not “please secure my app.” That's junk. Give the agent a lockfile diff. Tell it to classify risk. New maintainer? New install script? New transitive dependency with network access? New Git dependency? Have it produce a short review with evidence. JOSH: And then Gandalf gets involved? ERIK: Yeah. Gandalf can block the merge if the risk is wrong or the proof is thin. That's how agents should work. One agent proposes. Another reviews. The pipeline enforces. Nobody gets a magic wand. JOSH: That's a very specific architecture. ERIK: It has to be. “AI in CI” without policy is just a faster way to create a mess. Put the agent in one lane. Give it inputs. Require outputs. Validate those outputs. Store the decision. Then you can trust the pattern over time. [beat] ERIK: Also, builders need to stop pretending convenience is free. Every package you add is a vendor relationship. Maybe tiny. Maybe unpaid. Still a relationship. If that maintainer gets compromised, bored, angry, or bought, your build inherits the problem. JOSH: That's grim, but fair. ERIK: It's not grim. It's normal engineering. Networks taught us this years ago. You don't let random devices join a production routing domain because they asked nicely. Same idea. Dependencies need admission control. [pause] JOSH: The third deep story is Google and AI Overviews. A German ruling says Google can be liable for false answers. Why does that matter beyond Google? ERIK: Because it moves AI output from novelty to responsibility. If your product summarizes something and a user relies on it, you may own the damage. That's the whole shift. JOSH: Is that going to make companies pull back? ERIK: Some will. The sloppy ones should. If your AI feature is just a text box glued to a model, legal pressure is going to hurt. If your system has citations, audit logs, confidence thresholds, and a route to human review, you're in better shape. JOSH: That sounds like ScanBrief. ERIK: ScanBrief scored 99 items across 54 sources today. It doesn't just ask a model, “what's news?” It pulls from feeds, scores relevance, ranks stories, and keeps source context. That matters because summaries are cheap. Traceable summaries are the product. [beat] ERIK: If an AI says “Company X got acquired” and it's wrong, I want to know which source caused it, what the score was, what model touched it, and whether another source confirmed it. Without that, you're just publishing fog. JOSH: So the ruling is really about accountability. ERIK: Yes. And builders should take it seriously before regulators force them to. Every AI output should have a provenance story. Where did the data come from? What transformed it? What confidence did the system have? Was there a human approval path for risky cases? JOSH: What counts as risky? ERIK: Anything with money, health, legal claims, employment, security, or reputation. If the answer can cost somebody cash, access, a job, or public trust, don't treat it like a chatbot response. Treat it like a change request. JOSH: That's a good way to frame it. ERIK: Same pattern as infrastructure. A config change can take down a network. A false AI answer can take down trust. Different blast radius, same discipline. [beat] ERIK: This is where HumanRail-style thinking matters. When confidence is low, route to a person. When sources disagree, don't hide it. When the model is guessing, say it's guessing or don't publish it. JOSH: But doesn't that slow everything down? ERIK: A little. Good. Speed without control is not speed. It's gambling with better fonts. JOSH: That's going on a mug. ERIK: Please don't. [beat] ERIK: Look, I want AI everywhere. I'm the guy building agents to replace manual work. But replacing manual work doesn't mean removing judgment. It means putting judgment at the right choke points so humans aren't clicking the same button 400 times. JOSH: How does this apply to a builder shipping a small app? ERIK: Add source links. Store prompts and model responses for important actions. Add a confidence threshold. Put a review queue in front of anything risky. Let users report bad output. And when the AI makes a claim, make the UI show why. JOSH: Not flashy. ERIK: Correct. Useful beats flashy. The builders who win here won't be the ones with the cutest demo. It'll be the ones whose systems can survive contact with users, lawyers, outages, and Tuesday. [pause] ERIK: This episode is sponsored by Prime Automation Solutions. If you're still doing it manually, we automate it. Also, special on a website — $250. primeautomationsolutions.com [pause] JOSH: Alright, what's the AI pro tip today? ERIK: Make your agent produce a receipt before it changes anything. [beat] ERIK: Here's the pattern. Give the agent a task, but require five fields before execution: observed problem, files touched, tests to run, rollback plan, and risk level. If any field is missing, the agent stops. If the risk level is high, it routes to review. JOSH: So the prompt is basically a gate. ERIK: The prompt is part of it. The system enforcement is the real part. Don't trust the model to remember your policy. Put the policy in code. Have CI reject an agent patch if the receipt is missing. Have another model review the receipt against the diff. Make the boring paperwork automatic. [beat] ERIK: You can do this today with Claude, GitHub Actions, and a simple JSON schema. The model writes the receipt. CI validates the schema. Tests run. A reviewer agent checks whether the receipt matches the patch. Then a human or your merge policy decides. JOSH: That's practical. ERIK: It also changes behavior. Agents do better work when they know they have to explain the work in a structured way. Same as engineers. That's your tip. Use it. [pause] ERIK: If you're building toward financial independence through automation, my first book walks through the whole path. Free chapter at erikandersonbook.com. [pause] JOSH: One more thing — we started a Discord for builders. If you're shipping AI, automation, or anything that makes a human obsolete — come hang out. Link at buildorbereplaced.dev. ERIK: Post what you built. We'll post what we're building. Real wins, real builds, no fluff. [pause] ERIK: Build or be replaced. JOSH: If you want these signals in your inbox every morning, scanbrief.dev. See you tomorrow.