The top secret URSALA, RAQUEL, and FARRAH satellites (2025) | Build or Be Replaced
Today: The top secret URSALA, RAQUEL, and FARRAH satellites (2025) | LinkedIn Larpmaxxing | Gemini 4 Argon Episode date: 2026-10-01.
Download MP3 →
Build or Replaced
Today: The top secret URSALA, RAQUEL, and FARRAH satellites (2025) | LinkedIn Larpmaxxing | Gemini 4 Argon Episode date: 2026-10-01.
Download MP3 →ERIK: Build or be replaced. JOSH: It's Thursday, October 1st. This is Build or Be Replaced — powered by ScanBrief.dev. I'm Josh, here with Erik Anderson. ERIK: Today's theme is machines doing the babysitting so humans don't have to — inference engines that tune themselves, functions that boot before you finish the request. JOSH: Stick around — Erik's got an AI pro tip at the end about picking the right model for the job instead of defaulting to the biggest one. [pause] JOSH: Okay, headlines. First one — there's a report out that a popular CI/CD action got compromised. A maintainer's token leaked through a malicious dependency and it rode into thousands of pipelines before anyone caught it. ERIK: This is the one that should scare people more than it does. Nobody reads their CI config line by line. They just trust the marketplace action and move on. I run my own pipeline through PrimeBus with Gandalf review sitting in front of every merge specifically because I don't trust third-party anything by default — including my own code half the time. JOSH: Next — a major cloud provider's status page went down during their own outage. The thing meant to tell you it's down was also down. ERIK: That's the funniest kind of infrastructure failure. You put your status page on the same blast radius as the thing it monitors and it goes dark right when everyone needs it most. PrimeDash exists on my side so I'm never trusting a vendor's own self-report — I want my own eyes on every service in the lab, independent of whether the vendor's dashboard is even up. [beat] JOSH: Last one — 56k.rip. Somebody rebuilt the 1996 dial-up experience. ERIK: The modem screech, the page load, all of it. Fun nostalgia hit. Also a good reminder — we complain about a two-second LLM response time now. We used to wait two minutes for a GIF. [pause] JOSH: Alright, let's go deep. Gemini 4 Argon dropped — Google's calling it a frontier model for long-horizon engineering work. What's actually different here? ERIK: It's not optimized to sound smart in a single answer. It's built for long-horizon workflows — stuff that takes fifty steps, not five. Legal, finance, cybersecurity. And they're rolling it out through something called the Fairwind program first — trusted cyber-defense partners only, not a public API day one. JOSH: That's a weird way to launch a model. Why gate it like that? ERIK: Because a model that's actually good at multi-step reasoning over sensitive domains is also a model that's actually good at finding exploits or drafting attacks. You don't hand that out broadly on day one. I run PrimeRouter as the thing that decides which backend gets which job in my fleet, and a restricted-access frontier model is exactly the kind of lane I'd put behind extra gates before anything customer-facing touches it. Dark by default, prove it in a sandbox first. [beat] JOSH: Does that match how you'd roll out something like that yourself? ERIK: Almost exactly. I don't let a new model touch production traffic the first week it exists. I run it in isolation, watch what it does, then open the lane a little at a time. With PrimeRouter that means the model gets registered as a backend, it sits behind a feature flag that defaults off, and it only gets a chain slot after I've watched it handle real work in a sandbox without a human catching it doing something dumb. Gemini 4 Argon doing that at Google scale just tells me the pattern holds no matter how big the lab is. Big labs and one-guy home labs end up at the same answer: you don't trust a new model with real traffic until it's earned it. JOSH: Is there a version of this you'd actually use once it opens up? ERIK: If the long-horizon claims hold, yeah — anything in my pipeline that's fifty steps instead of five is a candidate. Fixer chains, multi-file refactors, that kind of thing. But I'd want to see it inside PrimeSentinel's watch first, same as any other new backend, before it gets anywhere near a lane that matters. [pause] JOSH: Next one — Magnitude, a Y Combinator company, built what they're calling a self-optimizing inference engine for agents. That sounds like it's aimed right at what you do all day. ERIK: It is. The pitch is the engine watches how your agent actually performs — latency, cost, success rate — and adjusts routing and parameters on its own instead of you hand-tuning it. That's the exact problem I solve manually right now. I've got agents on the mesh — Neo, Homer, Bill — and when one's slow or expensive I go in and retune it myself. JOSH: So this would do that automatically? ERIK: In theory. I'm skeptical of "self-optimizing" claims until I see the failure mode, though. What happens when it optimizes for speed and quietly tanks quality? I'd want to see it under PrimeSentinel-style watching before I trusted it unsupervised — something keeping an eye on it for stuck jobs or silent degradation, not just taking the vendor's word that it's smart. [beat] JOSH: You've built your own version of that watchdog already, haven't you? ERIK: That's basically what PrimeSentinel is for me — it catches jobs that stall instead of fail loud. The scary ones aren't crashes, they're the agent that just quietly stops making progress and nobody notices for six hours. Any "self-optimizing" system needs that same kind of outside eye on it, or you're trusting a black box to grade its own homework. JOSH: And meanwhile your own pipeline is already running at scale without a human in the loop most of the time. ERIK: PrimeBus auto-merger has run 2,295 attempts since June 5th. 1,507 got merged, 788 got blocked by Gandalf review, and zero — zero — got escalated to me. That's a 65.7% merge rate, and I want to be clear, the 788 blocked ones aren't failures. That's the guardrail working exactly like it's supposed to. A system that merges everything isn't safer, it's just unreviewed. JOSH: So if Magnitude showed up on your doorstep tomorrow, what would it actually need to prove before you trusted it? ERIK: Same thing I ask of anything new in the fleet — show me the blocked cases, not just the merged ones. Anybody can demo the happy path. I want to see what it refused to do and why. That's where the real signal is. [pause] JOSH: Last deep dive — 5x faster Edge Functions, moving from V8 isolates to Firecracker microVMs. Translate that for me. ERIK: V8 isolates are the lightweight sandboxing Cloudflare Workers made popular — fast to start, but limited in what you can run inside them. Firecracker is the microVM tech Amazon built for Lambda — heavier isolation, but historically slower cold starts. Getting a 5x speedup while moving to the heavier isolation model is the interesting part. Usually you trade one for the other. JOSH: Why does cold-start speed matter that much? ERIK: Because every serverless function, every edge route, every webhook handler lives or dies on cold start time. ScanBrief pulled in 105 items across 56 sources again this morning, average relevance score 51.8 — all of that runs through scoring functions that spin up, do their job, and die. Shave milliseconds off every one of those and it adds up fast across hundreds of runs a day. [beat] JOSH: Does this change anything for how you'd build going forward? ERIK: It's a signal, not a switch I'm flipping today. Right now I've got 144 services running across the lab and 140 distinct projects feeding telemetry into PrimeBus. If microVM cold-starts get genuinely competitive with isolates, that changes the math on where I put short-lived work — things like the PAS website-pipeline that spin up a build, check a client site, and tear down. Right now I keep those on long-running containers because cold start was too expensive to pay every time. If that cost drops, more of that work goes event-driven instead of always-on. JOSH: What would it take for you to actually move something over? ERIK: Real numbers from a real workload, not a vendor benchmark. I'd run one low-stakes function — something out of the PAS pipeline, not anything customer-facing — on the new runtime for a week and watch it next to the old one in PrimeDash before I trusted a blog post's 5x claim. Infrastructure vendors round in their own favor. Always have. [pause] ERIK: This episode is sponsored by Prime Automation Solutions. If you're still doing it manually, we automate it. Also, special on a website — $250. primeautomationsolutions.com. [pause] JOSH: Alright, what's the AI pro tip today? ERIK: Stop defaulting to the biggest model for everything. I run PrimeRouter specifically so every job gets matched to the right backend instead of me hand-picking a model every time — classification and quick lookups go to something small and cheap, long multi-step reasoning goes to a frontier model, and routine code fixes go wherever's fastest and available. If you're paying frontier-model prices to summarize a paragraph, you're burning money for no reason. Build yourself even a dumb if-else router before you build anything fancier — priority in, backend out. You'll cut your spend without losing quality on the stuff that actually needs the big model. ERIK: That's your tip. Use it. [beat] JOSH: Track your freedom score and net worth with the Freedom Blueprint app — free download, link in the show notes. [pause] JOSH: One more thing — we started a Discord for builders. If you're shipping AI, automation, or anything that makes a human obsolete — come hang out. Link at buildorbereplaced.dev. ERIK: Post what you built. We'll post what we're building. Real wins, real builds, no fluff. [pause] ERIK: Build or be replaced. JOSH: If you want these signals in your inbox every morning, scanbrief.dev. See you tomorrow.