Mon–Fri · 6 AM ET
← All Episodes
EP  • 00:13:30

Olo (Color) | Build or Be Replaced

Today: Olo (Color) | GPT 5.6 Sol is the best "vision" model OpenAI ever released | Sun Clock Episode date: 2026-08-18.

Download MP3 →

Transcript

JOSH: It's Tuesday, August 18. This is Build or Be Replaced, powered by ScanBrief.dev. I'm Josh, here with Erik Anderson.
ERIK: Cheap AI is here, agentic fixes are getting real, and the internet is becoming training data warfare.
JOSH: Stick around, Erik's got an AI pro tip at the end about routing cheap models without letting them touch production.
JOSH: [pause]
JOSH: First headline. OpenAI cut GPT-5.6 Luna pricing by 80 percent and Terra by 20 percent, but Sol stayed the same. Why does that matter?
ERIK: Because the cheap tier just got useful enough for boring work. That means summaries, triage, classifiers, screenshot reads, and pre-checks can run all day without lighting the budget on fire.
JOSH: So not a frontier model story.
ERIK: No. It's a plumbing story. Sol staying full price tells you OpenAI still wants premium money for the hardest work, but Luna is now the default worker bee.
JOSH: [beat]
JOSH: Second headline. GitHub's agentic autofix for code scanning is in public preview, and Copilot for Jira is now generally available. Helpful or terrifying?
ERIK: Both. Assigning a security alert to an agent and getting a pull request back is useful. Letting that agent wander around without policy is how you create a new incident while fixing the old one.
JOSH: That's comforting.
ERIK: Accurate, though.
JOSH: [beat]
JOSH: Third headline. Reports say Israel funded websites designed to influence AI chatbot answers about Gaza. Is that just propaganda with better SEO?
ERIK: It's worse. It's source poisoning for machines. If your AI system reads the web and doesn't score source trust, you're letting the loudest contractor write your context.
JOSH: [pause]
JOSH: Start with pricing. Luna down 80 percent, Terra down 20 percent, Sol unchanged. What's the real story?
ERIK: The real story is Jevons paradox with tokens. When the price drops, people don't spend less. They use more. They move AI calls into places they avoided before because the math was ugly.
JOSH: Like what?
ERIK: Ticket cleanup. Log summaries. Test failure triage. Email classification. PR description checks. Compliance evidence. Screenshot inspection. All the little jobs engineers ignore because nobody wants to pay frontier-model prices for a sentence that says, yeah, the button moved.
JOSH: And now Luna handles that?
ERIK: For a lot of it, yes. You don't need Sol to classify a failed cron job. You don't need the most expensive model to summarize a Terraform plan. You need the cheapest model that can do the job and a system that knows when to escalate.
JOSH: That's where PrimeRouter comes in?
ERIK: Exactly. PrimeRouter is my LLM gateway. Apps don't get to freeload their own model choice. They send the task, priority, context size, and tolerance for mistakes. The router picks the provider and model. If Luna can do it, Luna gets it. If it needs vision, it routes to vision. If it needs hard reasoning, it can move up.
JOSH: That sounds like overkill until the bill shows up.
ERIK: Most infrastructure sounds like overkill until it saves you. Then everyone pretends it was obvious.
JOSH: [beat]
JOSH: What changes when the cheap model also has a huge context window?
ERIK: People start stuffing entire projects into prompts. That's the dangerous part. Bigger context makes lazy design easier. You can paste the whole repo, but should you? Usually no.
JOSH: Why not?
ERIK: Because context isn't free just because the model is cheaper. More context means more places to get confused, more secrets to accidentally include, more stale files in the prompt, and more cost across thousands of calls. Cheap does not mean careless.
JOSH: So what's the better pattern?
ERIK: Retrieval with boundaries. Give the model the files it needs, not the entire attic. PrimeBus sends an event, PrimeRouter gets the task, and the tool layer pulls relevant code, logs, or docs. Then the answer gets validated against something real.
JOSH: What counts as real?
ERIK: Tests passed. Schema valid. Diff stayed inside the target files. No secrets touched. No deployment policy changed. No auth code edited without a human gate. That's real. A model saying "I am confident" is not real.
JOSH: Wait, really? You don't use confidence?
ERIK: I use confidence as a hint, not a decision. Models are great at sounding sure while standing in a hole. Outcome checks matter more.
JOSH: [beat]
JOSH: How does this affect a builder listening today?
ERIK: Put a router in front of your model calls. Even a simple one. One function. Task type in, model out. Start with three tiers: cheap, normal, expensive. Log every call. Log the reason. Log the cost. Then review a week later and you'll see the waste immediately.
JOSH: That simple?
ERIK: Duuude, yes. Builders keep waiting for the perfect AI platform. Start with a table and a policy. "Summaries use Luna. Code review uses Terra. Production deployment decisions require human approval." That alone beats random API calls buried across ten scripts.
JOSH: And Sol?
ERIK: Use Sol when being wrong costs more than the model call. Architecture decisions. Security analysis. Hard debugging. Anything touching money, identity, or production. Don't use it to rewrite a Slack message unless you enjoy donating to cloud margins.
JOSH: [pause]
JOSH: That brings us to GitHub. Agentic autofix can take code scanning alerts and open a PR. What's different now?
ERIK: Classic autofix gave you a suggested patch. Agentic autofix behaves more like a junior developer. It explores files, proposes a fix, runs analysis again, and opens a pull request. GitHub also has Copilot for Jira generally available now, so the agent can be triggered from work items and stream progress back into Jira.
JOSH: That sounds useful.
ERIK: It is useful. That's why you need to be serious about it. Useful tools get permissions. Permissions create blast radius.
JOSH: Give me the practical fear.
ERIK: The practical fear is not "the AI becomes evil." That's movie stuff. The real fear is boring. It fixes a security alert by widening an exception. It adds a dependency you don't trust. It changes a test instead of the code. It edits a config file that changes runtime behavior. It closes the alert and quietly creates a worse problem.
JOSH: The green checkmark lies.
ERIK: The green checkmark tells one story. You need the rest of the story.
JOSH: [beat]
JOSH: You run auto-review and auto-merge through PrimeBus. How do you cage it?
ERIK: Events first. PrimeBus doesn't say "go do whatever." It emits specific events. Test failed. Build passed. Lint failed. Service stuck. Dependency changed. Then agents subscribe to the event type and only get the tools they need.
JOSH: Tool permissions per job.
ERIK: Exactly. A lint fixer doesn't need deploy access. A docs updater doesn't need secrets. A test triage agent doesn't need write access to infrastructure. If the agent has every key on the ring, that's not automation. That's a breach with a friendly username.
JOSH: Dry, but fair.
ERIK: Also true in network automation. With Cisco NSO, you don't let every script push device config because it feels convenient. You model services, validate intent, check diffs, and commit through guardrails. AI doesn't change that. It makes it more important.
JOSH: How does Gandalf fit in?
ERIK: Gandalf is my gatekeeper. It reviews proposed changes before merge. It looks for risky file paths, suspicious diffs, missing tests, dependency changes, secret patterns, and policy violations. If it smells wrong, it blocks. No drama. Just no.
JOSH: You like the blocked PRs.
ERIK: Correct blocks are wins. People measure automation by how often it says yes. Wrong metric. Measure whether it says no when it should.
JOSH: [beat]
JOSH: What should teams do before turning on agentic autofix?
ERIK: First, branch protection. Required reviews for sensitive code. Required tests. Required code scanning. Required secret scanning. Second, CODEOWNERS for auth, billing, deployment, and infrastructure. Third, custom instructions that say what the agent cannot touch. Fourth, a policy bot that checks the diff.
JOSH: That's a lot before the fun part.
ERIK: The fun part is not waking up to a self-inflicted incident. Ask me how I know from networking. Every "quick script" becomes production if it works twice.
JOSH: Is Jira integration a risk too?
ERIK: It can be. Jira has messy human context. Tickets include partial requirements, pasted logs, internal links, sometimes credentials because people are people. If that context flows into a public pull request or an agent session, you need to know exactly what gets copied.
JOSH: So the ticket becomes part of the prompt.
ERIK: Yep. And prompts are not magic private thoughts. They're data moving through systems. Treat them like logs. Classify them. Redact them. Keep the agent from dragging sensitive text into places it doesn't belong.
JOSH: What's the builder move here?
ERIK: Start with read-only. Let the agent explain the fix. Then let it open PRs. Then let it auto-merge only low-risk categories after a long audit trail. Don't jump from "cool demo" to "write access to prod" because the vendor video had nice lighting.
JOSH: [pause]
JOSH: Third story. Reports say an Israeli government-funded campaign created websites aimed at shaping AI chatbot answers about Gaza. What matters technically?
ERIK: The technical story is that AI search and training pipelines trust the web too much. If someone can publish enough polished content, get it indexed, get it crawled, and get chatbots to cite it, they can influence the answer without ever convincing a human audience.
JOSH: So the websites don't need readers.
ERIK: Right. Humans are optional. The crawler is the audience. That's the weird new part.
JOSH: That's wild.
ERIK: It's search engine manipulation with a different target. Old SEO wanted your click. This wants to become the answer.
JOSH: [beat]
JOSH: How does a builder defend against that?
ERIK: Source scoring. Provenance. Recency. Cross-checking. Domain reputation. Disclosure detection. Multiple independent sources. And stop treating a chatbot citation as truth. A citation means the model found text. It doesn't mean the text is clean.
JOSH: How does ScanBrief handle that?
ERIK: ScanBrief doesn't just grab a headline and call it done. It compares sources, dedupes stories, ranks relevance, and I still care where the signal came from. If a source is unknown, partisan, recently created, or only echoed by copies of itself, that should lower trust.
JOSH: Can small teams actually do that?
ERIK: Yes. You don't need a research lab. Add fields to your ingestion pipeline: domain age if you can get it, source type, author presence, outbound citations, whether other reputable sources confirm it, and whether the article discloses sponsorship. Score it. Even a crude score beats blind ingestion.
JOSH: What about RAG systems inside companies?
ERIK: Same problem, smaller room. If your internal AI reads Confluence, Jira, Google Drive, and Slack, bad context can poison answers. Maybe it's not propaganda. Maybe it's an old runbook. Maybe it's a draft policy. Maybe it's a random Slack thread from 2022 that everyone forgot.
JOSH: The model doesn't know it's stale.
ERIK: Unless you tell it. Put dates on docs. Mark owners. Expire runbooks. Add approval status. PrimeSentinel watches jobs for stuck behavior, but the same idea applies to knowledge. Stale knowledge needs alerts too.
JOSH: That's a nice way to think about it.
ERIK: It's infrastructure. Knowledge has health checks. If your runbook hasn't been reviewed since before your last architecture change, it's not documentation. It's a trap with headings.
JOSH: [beat]
JOSH: This makes AI answers feel less reliable.
ERIK: Good. They should feel conditional. AI is incredibly useful, but you need to know what it touched. Which sources? Which files? Which tools? Which assumptions? Builders who track that will move fast. Builders who don't will get a beautiful answer built on garbage.
JOSH: What's the connection between all three stories today?
ERIK: Cheap models increase usage. Agentic tools increase action. Source poisoning attacks the input. That means the full loop matters now. Input trust, model routing, tool permissions, output validation. Miss one and you're guessing.
JOSH: That's the episode right there.
ERIK: Pretty much. Build the loop, not the demo.
JOSH: [pause]
ERIK: This episode is sponsored by Prime Automation Solutions. If you're still doing it manually, we automate it. Also, special on a website — $250. primeautomationsolutions.com
JOSH: [pause]
JOSH: Alright, what's the AI pro tip today?
ERIK: Put a cheap-model gate in front of every expensive model call. Here's the pattern. First pass uses Luna or your cheapest reliable model. It must return three things: answer, confidence reason, and escalation flag. Then your code checks the output against rules. If the task touches protected files, missing citations, secrets, auth, billing, deploy config, or low confidence, route up to Terra or Sol. Don't let the model decide alone. Make the software decide with policy.
JOSH: Give me one example.
ERIK: Code review. Cheap model summarizes the diff and labels risk. Your script checks the file paths. If it sees Terraform, Kubernetes manifests, NSO service packages, auth handlers, or migrations, it escalates. If it's docs or a typo fix, it stays cheap. You save money and keep sharp tools away from boring work.
JOSH: That's usable today.
ERIK: Yep. One wrapper function. One policy file. One log table. That's your tip. Use it.
JOSH: [pause]
JOSH: We also drop daily market picks and automation tips on YouTube — search Build or Be Replaced.
JOSH: One more thing — we started a Discord for builders. If you're shipping AI, automation, or anything that makes a human obsolete — come hang out. Link at buildorbereplaced.dev.
ERIK: Post what you built. We'll post what we're building. Real wins, real builds, no fluff.
JOSH: [pause]
ERIK: Build or be replaced.
JOSH: If you want these signals in your inbox every morning, scanbrief.dev. See you tomorrow.