GPT‑Live | Build or Be Replaced
Today: GPT‑Live | Show HN: Microsoft releases Flint, a visualization language for AI agents | TypeScript 7 Episode date: 2026-07-09.
Download MP3 →
Build or Replaced
Today: GPT‑Live | Show HN: Microsoft releases Flint, a visualization language for AI agents | TypeScript 7 Episode date: 2026-07-09.
Download MP3 →JOSH: It's Thursday, July 9. This is Build or Be Replaced — powered by ScanBrief.dev. I'm Josh, here with Erik Anderson. ERIK: Benchmarks are getting exposed, agents are getting diagrams, and tractors just became a software freedom story. JOSH: Stick around — Erik's got an AI pro tip at the end about making agent fixes prove themselves before they touch production. [pause] JOSH: First headline. OpenAI says some coding benchmarks have too much noise. Erik, why should builders care? ERIK: Because fake signal makes people pick the wrong model. If a benchmark says a model is great, but the tasks are broken, flaky, or weirdly scored, you're not measuring engineering ability. You're measuring trivia with a compiler attached. JOSH: Next up, Microsoft released Flint, a visualization language for AI agents. Useful, or just another diagram tool? ERIK: Useful if it shows decisions, tool calls, state, and failure paths. Agent systems are not magic. They're distributed systems with prompts. If you can't see the graph, you're guessing. JOSH: Third headline. John Deere owners are getting repair rights under an FTC settlement. Software meets tractors. ERIK: That's the real story. The machine is physical, but the lock is software. Right to repair is becoming right to inspect, right to patch, and right to run your own tooling. [pause] JOSH: Let's start with OpenAI calling out coding evaluations. That feels inside baseball, but it keeps showing up everywhere. ERIK: It matters because everybody is buying AI off scoreboard screenshots right now. SWE-Bench, internal evals, vendor charts, leaderboards. Cool. But if the eval has noisy tasks, bad expected answers, or flaky setup, the model ranking starts lying to you. JOSH: Wait, really? One bad benchmark can move buying decisions? ERIK: Absolutely. Managers see one chart and suddenly the whole company is migrating. Engineers do it too. We all pretend we're above it, then we see a model jump three points and start rewriting our toolchain. [beat] ERIK: The better way is boring. Build your own evals from your own failures. Pull real bugs. Pull real tickets. Pull the exact kind of repo context your agents will see. Then score the thing that matters: did the patch pass tests, did it stay inside the blast radius, did it explain the change, and did the review gate agree? JOSH: That's basically how you treat PrimeBus, right? ERIK: Yeah. PrimeBus doesn't care if a model has a shiny benchmark score. It cares if the code works. The auto-merger has run 2295 attempts since June 5: 1507 merged, 788 blocked by Gandalf, 0 escalated to me. That's the metric I trust. Not because the merge rate is pretty. Because the blocked count proves the guardrails are doing their job. JOSH: That blocked number sounds like the point, not the failure. ERIK: Exactly. If Gandalf blocks 788 bad or risky merges, that's not waste. That's the system saying no before I wake up to a mess. People want agents to be brave. Wrong. I want agents to be productive and cowardly near production. [beat] JOSH: How would a normal team copy that without your whole Bobaverse setup? ERIK: Start with one repo and one workflow. Don't build the galaxy. Take your last ten bugs. Turn them into test fixtures. Give the model the issue, the failing test, and the repo. Make it produce a patch. Then run the exact same checks every time. ERIK: The key is you need a judge that is not the same worker. One agent writes. Another reviews. CI runs. A policy gate checks scope. If the patch touches auth, billing, infra, secrets, database migrations, whatever matters in your world, it gets a higher bar. JOSH: So the eval isn't "can it code." It's "can it survive our process." ERIK: That's it. Coding alone is not the job. Shipping safely is the job. A model that writes beautiful code but ignores a test failure is a liability with syntax highlighting. [pause] JOSH: Microsoft Flint is next. A visualization language for AI agents sounds very enterprise. What do you see there? ERIK: The phrase sounds enterprise, yeah. But the need is real. Agent systems are hard to debug because the important stuff happens between steps. Prompt goes in, tool call comes out, state changes, another agent reacts, maybe a retry happens, then nobody knows why the final answer looks haunted. JOSH: What should a good agent visualization show? ERIK: It should show the graph. Inputs, outputs, tools, decisions, retries, memory reads, memory writes, and stop conditions. Not just happy path. Happy path diagrams are decorations. Show me where it can loop. Show me where it can spend money. Show me where it can write to disk. Show me where a human gets pulled in. JOSH: That sounds more like network automation than chatbot stuff. ERIK: Duuude, it is network automation. Same pattern. In NSO, Terraform, Kubernetes, NATS, whatever, the question is always state. What did you think the state was, what did you change, what happened after, and who approved it? [beat] ERIK: Agents are just louder about it because they speak English while making questionable decisions. JOSH: Dry but fair. ERIK: In the Echo and Neo lab, I don't want vibes. I want traces. There are 117 services running on the production server right now, and 137 distinct projects have emitted telemetry to PrimeBus. If something breaks, I need to know which service screamed, which agent touched it, what message hit the bus, and what changed after. JOSH: That's where Flint could fit? ERIK: If it makes agent flows visible and shareable, yes. Think about onboarding. A new engineer asks, "What happens when a test fails?" Instead of a doc that gets stale in eleven minutes, you show the actual graph: test failure event, PrimeBus message, Claude fix attempt, variant A and B, CI run, Gandalf review, merge or block. JOSH: That's a lot cleaner than a wall of logs. ERIK: Logs are still the source of truth. But diagrams help humans build the mental model. The trap is when teams use diagrams instead of telemetry. A static diagram of an agent is a bedtime story. A diagram backed by real events is useful. JOSH: What would you not want from a tool like that? ERIK: No pretty boxes that hide the hard parts. Don't show me "agent decides." Show me the decision inputs. Don't show me "tool executes." Show me the tool name, permission boundary, timeout, retry count, and what data came back. Don't show me a green check unless the test actually ran. [beat] JOSH: You sound like you've been burned. ERIK: Not burned. Educated by computers being computers. Every small issue becomes a guardrail. That's how you build this stuff. Start small in the lab. Prove the win. Add telemetry. Add a stop condition. Then let it touch more. JOSH: And that connects back to the OpenAI benchmark story. ERIK: Same root problem. You can't trust what you can't inspect. Benchmarks need inspectable tasks. Agents need inspectable flows. Automation needs inspectable outcomes. If any part of that chain is hidden, you're back to faith-based engineering. [pause] JOSH: The John Deere right-to-repair story feels different. Why put that next to AI agents and coding evals? ERIK: Because it's the same fight wearing overalls. JOSH: That's a sentence. ERIK: Look, the tractor owner bought the machine. But the repair path got trapped behind software, diagnostics, firmware, and vendor access. That is not just a farm issue. That's every modern system. Cars, routers, phones, medical gear, cloud appliances. The hardware is yours until the software says it isn't. JOSH: What's the builder angle? ERIK: Builders should pay attention because access is power. If you can't inspect logs, export config, run diagnostics, replace a part, or patch your own system, you don't own it. You're renting permission. JOSH: That hits infrastructure too. ERIK: Constantly. Network teams know this pain. A vendor box throws an error, the CLI gives you half a clue, the real diagnostic is hidden, and support says upload a bundle into the portal. Cool. Meanwhile the outage timer is laughing at you. [beat] ERIK: That's why I like open tooling where it makes sense. Postgres. Linux. NATS. Terraform state, even when it's annoying. You can inspect it. You can build around it. You can recover without asking a permission server for emotional support. JOSH: How does that compare to the John Deere case? ERIK: Same principle. Independent repair shops need diagnostic access. Owners need parts and software paths. In infrastructure, internal teams need runbooks, source control, test environments, and automation hooks. If the only person who can fix a system is the vendor, that's not resilience. That's a subscription to waiting. JOSH: Does AI make that better or worse? ERIK: Both. AI can make repair easier because it can read logs, compare configs, generate runbook steps, and suggest fixes fast. But if the platform blocks access, the AI is stuck outside the fence with a clipboard. JOSH: That's bleak. ERIK: It's practical. Give an agent real telemetry and it can help. Give it screenshots and vibes, and now you're paying tokens to squint. [pause] JOSH: You mentioned ScanBrief earlier. It pulled today's signals out of a pretty noisy feed, right? ERIK: Yeah. ScanBrief scored 97 items across 54 sources today. That's exactly why I built it. I don't need another feed of everything. I need the stuff that maps to builders: AI systems, evals, languages, infra, security, ownership. JOSH: Like TypeScript 7 and Postgres in Rust? ERIK: Yep. TypeScript 7 is worth watching because developer tooling speed matters. If your type checker gets faster, your agents get faster too, because the feedback loop is shorter. Postgres rewritten in Rust passing regression tests is interesting for a different reason. Not because everyone should rewrite databases tomorrow. Please don't. But because compatibility is the hard part, not the language flex. JOSH: And Bun being rewritten in Rust? ERIK: Same bucket. Runtime projects are getting more serious about memory safety, performance, and maintainability. Rust keeps showing up because it gives systems people more control without C turning every Tuesday into a crime scene. [beat] JOSH: That's going on a mug. ERIK: Put a stack trace on the back. [pause] JOSH: Bring this home. What's the bigger pattern across these stories? ERIK: The pattern is proof. OpenAI is saying coding evals need better proof. Flint is about proving what agents did. John Deere is about proving owners and shops can inspect and repair what they bought. TypeScript, Rust, Postgres, Bun, all of that is proof through tests, compatibility, and repeatable builds. ERIK: Builders who win are going to build proof into the system. Not after. Not as a meeting. In the workflow. Every agent action leaves a trace. Every patch runs tests. Every merge has a reviewer. Every alert maps to a runbook. Every runbook can turn into automation. JOSH: That's the Dark Factory idea. ERIK: Exactly. Do one thing each day. Put events on a bus. Let agents subscribe. Add guardrails. Add telemetry. Keep the human for judgment, not button clicking. The last 7 days in PrimeBus had 294 auto-merger runs: 207 merged, 87 blocked, 0 escalated. That's not because the agents are perfect. It's because the system expects them not to be. JOSH: That's probably the most useful sentence today. ERIK: Agents are interns with rocket fuel. Give them work. Give them tests. Don't give them the keys without a fence. [pause] ERIK: This episode is sponsored by Prime Automation Solutions. If you're still doing it manually, we automate it. Also, special on a website — $250. primeautomationsolutions.com [pause] JOSH: Alright, what's the AI pro tip today? ERIK: Make your agent fixes compete against a baseline before they touch anything important. Here's the simple version. When an agent proposes a patch, run three checks. First, does the original failure now pass. Second, does the broader test suite still pass. Third, did the patch stay inside the expected files. ERIK: Then add one more step. Ask a separate review agent to argue against the patch. Not summarize it. Attack it. Look for hidden behavior changes, missing tests, weird dependency changes, and config drift. JOSH: So don't ask the reviewer to be nice. ERIK: Never. Nice reviewers ship bugs. Give the review agent permission to block. If it blocks, store the reason. Over time, those reasons become your local benchmark. That's how your system gets smarter without pretending the model got smarter. ERIK: Start with a shell script and one repo. You don't need a giant platform. You need repeatable proof. That's your tip. Use it. [pause] JOSH: Track your freedom score and net worth with the Freedom Blueprint app — free download, link in the show notes. [pause] JOSH: One more thing — we started a Discord for builders. If you're shipping AI, automation, or anything that makes a human obsolete — come hang out. Link at buildorbereplaced.dev. ERIK: Post what you built. We'll post what we're building. Real wins, real builds, no fluff. [pause] ERIK: Build or be replaced. JOSH: If you want these signals in your inbox every morning, scanbrief.dev. See you tomorrow.