Gemini 3.5 Flash | Build or Be Replaced
Today: Gemini 3.5 Flash | I’ve built a virtual museum with nearly every operating system you can think of | Google changes its search box Episode date: 2026-05-20.
Download MP3 →
Build or Replaced
Today: Gemini 3.5 Flash | I’ve built a virtual museum with nearly every operating system you can think of | Google changes its search box Episode date: 2026-05-20.
Download MP3 →JOSH: It's Wednesday, May 20. This is Build or Be Replaced — powered by ScanBrief.dev. I'm Josh, here with Erik Anderson. ERIK: Guardrails are having a moment. Google just put them on search. An 8B model went from 53% to 99% because of them. My pipeline shipped 18 code changes to production overnight because of them. That's the theme today. JOSH: Stick around — Erik's got an AI pro tip at the end about routing workloads to the right model and actually cutting your costs. [pause] JOSH: Okay, first headline — Google is calling this the biggest search overhaul in 25 years. Gemini 3.5 Flash is now the default AI model for search. What's the short version? ERIK: The search box is an agent now. Not smarter autocomplete. An agent. You can give it multi-step tasks and it executes them. That's a real architectural shift, not a UI refresh. [beat] JOSH: Next — GitHub is investigating unauthorized access to their internal repositories. ERIK: Internal repos at GitHub means potentially source code, pipeline configs, infrastructure logic. Details are thin but scope could be significant. Worth watching. [beat] JOSH: And OpenAI is now using Google's SynthID watermarking system for AI-generated images. ERIK: Two competitors sharing infrastructure to solve the same problem. When that happens, the problem is real and neither of them wants to be the one left without a solution. [pause] JOSH: Let's go deep on Google first. Twenty-five years is a long time to not change something. ERIK: It is. And the reason search didn't change is it didn't have to. Type keywords, get a list of links, click through. That loop worked. The problem is that loop was built for a world where information was hard to find. Now information is everywhere and the bottleneck is synthesis — turning ten sources into one answer. That's what Gemini 3.5 Flash does inside search now. It doesn't just retrieve. It reasons. JOSH: What does that actually look like for someone using it? ERIK: You can ask it something like "find me an open-source Python library for PDF extraction that handles scanned pages and has active maintenance" and it comes back with a curated answer, not ten blue links you have to go read yourself. The complexity of the query used to determine how long it took you to get an answer. Now the model absorbs that complexity. Your time to answer drops. JOSH: Is Gemini 3.5 Flash actually good enough for that? ERIK: That's the smart question. They're saying 4x faster than competing models at frontier-level intelligence. Speed matters at Google's scale — they're serving billions of queries. You need fast and good. Haiku-tier speed with Sonnet-tier reasoning is roughly the target. If they hit that, it works. [beat] JOSH: ScanBrief flagged this one high today? ERIK: Pulled 51 items across 54 sources this morning. This one scored near the top. And it's episode 26 for us — the model story has been showing up almost every week since we started. The pace isn't slowing down. JOSH: How does this affect what you're building? ERIK: It doesn't threaten ScanBrief. Google search and ScanBrief are solving different problems. Google brings you what's indexed and popular. ScanBrief scores by relevance to what I actually care about — automation, infrastructure, AI tooling. It's curation, not retrieval. What does affect me is that Gemini 3.5 Flash is now in the mix as a routing option. PrimeRouter already handles multi-provider failover. I'll evaluate it on speed and token cost against what I'm running now. [beat] JOSH: You already have that infrastructure built to slot it in. ERIK: PrimeBus processed 3,891 automation events across 12 projects today. Any model with a better cost-to-speed ratio gets tested against that load. That's just how the fleet runs. [pause] JOSH: Okay, the guardrails story. Forge. An 8-billion parameter model at 53% on agentic tasks — that's not great. Then guardrails push it to 99%. How? ERIK: The number is the whole story. 53% on agentic tasks means the model is failing roughly half the time on anything multi-step. That's not a capability problem — that's a reliability problem. Guardrails don't make the model smarter. They intercept the outputs before they execute and check them against known failure patterns. Wrong tool call, hallucinated parameter, a loop that won't terminate — these are predictable failure modes. You pattern-match against them and block the bad output before it does damage. JOSH: So the model doesn't change at all? ERIK: Same weights, same training. You're adding a validation layer on top. The model still generates the bad output — you just never let it leave the system. That's why the jump is so dramatic. You go from "model fails half the time" to "model's failures are caught before they matter." JOSH: Is that how Gandalf works in your pipeline? ERIK: Exactly how. Since April 17, the auto-merger has run 1,641 attempts. 621 made it to production. 1,008 got blocked by Gandalf. [beat] JOSH: So more than half get blocked. That sounds like a lot. ERIK: It's supposed to. Every blocked PR is a guardrail working. Gandalf is checking for a specific set of failure patterns — the ones that burned us before. Asyncio task garbage collection, SQL-level race conditions, mixed import signatures, fire-and-forget calls that eat exceptions. Those aren't random failures. They're predictable. So they're blockable. JOSH: What about the 12 that got escalated to you? ERIK: Those are the ones where the fix was technically correct but the approach was wrong. You can't guardrail judgment calls. If a service is trying to solve a concurrency problem and the auto-fix works locally but creates a worse race condition at scale — Gandalf can't see that. That comes to me. Everything else runs without me. Eighteen code changes shipped to prod overnight. JOSH: You wake up in the morning and there are already eighteen changes in production. ERIK: And I didn't touch a terminal. That's 81 distinct projects emitting telemetry to PrimeBus right now. The fleet is running. My job is the judgment calls, not the execution. [beat] JOSH: The Forge paper basically proves what you built empirically. ERIK: The academic version, yeah. They ran the benchmarks. I ran it on 95 live services. Same conclusion: guardrails are the delta between a model that's almost useful and one you can trust with production code. The 8B model story matters because it means you don't need frontier-scale compute to build something reliable. You need the right architecture around whatever model you're running. JOSH: That's a big deal for cost. ERIK: Massive. If a guardrailed 8B model performs at 99% on your specific task, you have no reason to pay for a 70B. The conversation in AI right now is about model capability. The real variable is guardrail quality. Nobody's writing papers about that fast enough. [pause] ERIK: This episode is sponsored by Prime Automation Solutions. If you're still doing it manually, we automate it. Also, special on a website — $250. primeautomationsolutions.com [pause] JOSH: Alright, what's the AI pro tip today? ERIK: Stop defaulting everything to the most capable model you have access to. Every task has a cost-to-quality curve and most builders are ignoring it. Simple classification tasks — does this text match this category, yes or no — don't need Opus. They need the fastest model that gets it right 95% of the time. Save the expensive calls for the hard problems: architecture decisions, ambiguous error logs, anything where being wrong costs more than the API call itself. I route explicitly in PrimeRouter. Routine summarization goes to the fastest tier. Complex multi-step analysis escalates up. When I stopped defaulting everything to the top model, API costs dropped by more than half. Latency dropped too. Know what your workload actually requires, then match your model to that requirement. Not to your anxiety about whether the cheaper model is good enough. Test it. Most of the time it is. That's your tip. Use it. [pause] ERIK: If you're building toward financial independence through automation, my first book walks through the whole path. Free chapter at erikandersonbook.com. [pause] JOSH: One more thing — we started a Discord for builders. If you're shipping AI, automation, or anything that makes a human obsolete — come hang out. Link at buildorbereplaced.dev. ERIK: Post what you built. We'll post what we're building. Real wins, real builds, no fluff. [pause] ERIK: Build or be replaced. JOSH: If you want these signals in your inbox every morning, scanbrief.dev. See you tomorrow.