Mon–Fri · 6 AM ET
← All Episodes
EP  • 00:13:30

Using “underdrawings” for accurate text and numbers | Build or Be Replaced

Today: Using “underdrawings” for accurate text and numbers | BYOMesh – New LoRa mesh radio offers 100x the bandwidth | DeepClaude – Claude Code agent loop with DeepSeek V4 Pro Episode date: 2026-05-04.

Download MP3 →

Transcript

JOSH: It's Monday, May 4. This is Build or Be Replaced — powered by ScanBrief.dev. I'm Josh, here with Erik Anderson.
ERIK: AI stops being cute when it starts making calls faster than the humans in the room.
JOSH: Stick around — Erik's got an AI pro tip at the end about separating planner models from executor models.
[pause]
JOSH: Three quick ones to start. First up, image models are finally getting text and numbers right with underdrawings. Is that a real shift or another demo trick?
[beat]
ERIK: Real shift. If the model can place the label where you asked for it, half the fake design magic turns into useful production work.
[beat]
JOSH: Next one. BYOMesh says its LoRa mesh radio pushes way more bandwidth than the old stuff. Worth paying attention to?
[beat]
ERIK: Yep. Better long-range bandwidth means better telemetry, better field sensors, and fewer remote boxes pretending they didn't see the alert.
[beat]
JOSH: Last one. Mercedes says physical buttons are coming back. Did touchscreens finally annoy enough people?
[beat]
ERIK: They lost the second somebody hid defrost in a submenu. If the user is driving, working fast, or wearing gloves, glass is the wrong answer.
[pause]
JOSH: One headline we're taking deeper. OpenAI's o1 reportedly beat triage doctors with limited patient data. That's not a toy benchmark.
[beat]
ERIK: No. That's judgment under pressure. That's where the money is and that's where people get replaced.
[pause]
JOSH: Start there. Model gets around sixty-seven percent, humans land more like fifty to fifty-five with thin data. What's your first read?
[beat]
ERIK: First read is always the same. What was the exact task, what were the inputs, and what did they count as correct. That said, the gap still matters. Thin context, time pressure, real consequences, and the model still came out ahead. That's not autocomplete. That's decision support getting teeth.
[beat]
JOSH: Why does that hit so hard outside medicine?
[beat]
ERIK: Because every serious operation has triage. NOC alerts. Failed deploys. Security noise. Routing changes. API errors. Somebody gets handed six ugly clues and has to decide what's real before the blast radius grows. That's the same pattern. Different room, different badge.
[beat]
JOSH: So when people hear doctor story, you're hearing operations story.
[beat]
ERIK: Exactly. The medical angle gets attention because the stakes are obvious. But the pattern is everywhere. If a model can rank urgency better than a tired human with partial context, then every team doing incident response should pay attention right now.
[beat]
JOSH: Put that into your world. What does that look like on your systems?
[beat]
ERIK: PrimeBus is already doing a rough version. Every app in the lab throws events onto the bus. NATS moves them around. A failing build, a broken Selenium selector, a bad cron run, a dependency issue, config drift, all of it shows up as signals. Then an agent loop scores the likely cause and decides whether it should take the shot or route it to a human. We've logged 214 auto-fixes at a 78 percent success rate across the mesh.
[beat]
JOSH: That's wild. Two hundred fourteen fixes with no human at the keyboard?
[beat]
ERIK: Correct. Not magic either. Tight scope. Small blast radius. Strong rollback path. That's the adult version of autonomy. People screw this up because they think the only modes are full self-driving or no automation at all. Wrong. The right mode is let the machine act where failure is cheap and recovery is obvious.
[beat]
JOSH: And when it isn't obvious?
[beat]
ERIK: Route it out. That's why HumanRail exists. Confidence low, context messy, too many side effects, hand it to a person. Nobody sane wants a model freehanding production because it feels lucky today.
[beat]
JOSH: What's the part the average team gets wrong when they try this?
[beat]
ERIK: Bad context first. Fake metrics second. Bad context means they dump raw logs, random screenshots, and a Jira title into the model and expect judgment. That's lazy. Fake metrics means they say "analyst productivity improved" without proving the decision got better. That's corporate wallpaper.
[beat]
JOSH: What's a real metric then?
[beat]
ERIK: False positives down. Mean time to recovery down. Escalations avoided. Rollbacks avoided. Repeat incidents cut. On PrimeBus, the only thing I care about is whether the system made the right move and whether the verification passed. A pretty explanation means nothing if the service is still broken.
[beat]
JOSH: That feels like the dividing line now. Talking versus doing.
[beat]
ERIK: Yep. Everybody has words now. Good judgment tied to an action loop is still rare. That's why most "AI strategy" decks are already stale. The model doesn't need to write you a poem about the outage. It needs to tell you the auth token expired, rotate it, rerun the job, and prove the queue drained.
[beat]
JOSH: The limited-data part still bothers me a little. Humans usually defend themselves with "well, I need more context."
[beat]
ERIK: Humans always want one more dashboard and one more meeting. Sometimes that's real. A lot of the time it's cover. What this result says is imperfect context plus a good reasoning loop is enough to win in narrow lanes today. That's uncomfortable if your whole job is making medium-quality calls from incomplete inputs.
[beat]
JOSH: Which jobs should feel that first?
[beat]
ERIK: Triage-heavy jobs. Junior SOC work. Level-one incident review. Support queue sorting. Log classification. Change review prep. Research assistants doing first-pass filtering. A chunk of that work turns into model-first, human-second pretty fast.
[beat]
JOSH: Does that mean fewer people or better people?
[beat]
ERIK: Both. Fewer people doing rote sorting. Better people handling edge cases and building the guardrails. Same thing happened everywhere else. Once the machine can do the repetitive judgment, the value moves to system design, exception handling, and proving the thing can be trusted.
[beat]
JOSH: That's a hard sentence if you're still doing inbox triage by hand.
[beat]
ERIK: Then stop doing inbox triage by hand. Harsh answer, but clean answer.
[pause]
JOSH: Second deep dive. DeepClaude. Claude Code agent loop with DeepSeek V4 Pro. On paper that sounds like stacking models because people enjoy complexity.
[beat]
ERIK: Sometimes it is. Half the internet is wiring together five models to avoid making one decision. But there is a smart version. Each model needs a job. Planner. Executor. Reviewer. Maybe classifier. If they don't have distinct roles, you've built a Rube Goldberg machine with a token bill.
[beat]
JOSH: So when does it make sense?
[beat]
ERIK: When the jobs are different enough that one model's strength covers another model's weakness. Claude is strong at code edits, structure, and not completely losing the plot in a long task. A reasoning-heavy model can be good at first-pass decomposition, weird edge cases, or ranking options before the code writer touches anything.
[beat]
JOSH: You're doing that already, aren't you?
[beat]
ERIK: In a simpler way, yeah. Three Claude instances on the mesh. Bob on Neo, Bill on the Mac M3, Homer on Morpheus. PrimeBus pushes work across them. One agent can classify the failure, another can generate A and B fixes, another can verify. That's not academic. That's live infrastructure in a two-server home lab with 64 projects and 229 cron jobs.
[beat]
JOSH: Why not just use the best model everywhere and call it a day?
[beat]
ERIK: Because "best" is fake if the task changes every step. The model that writes decent code isn't always the model that plans well. The model that's cheap enough to watch logs all day isn't the one you want burning tokens on deep analysis. Same reason you don't use a forklift to deliver mail.
[beat]
JOSH: What's the architecture you actually like?
[beat]
ERIK: Cheap model first for routing and classification. Strong reasoning model second if the problem is ambiguous. Strong coding model for the patch. Then a verifier loop that doesn't care who wrote the fix. Tests pass or they don't. Service healthy or not. That's the stack.
[beat]
JOSH: Sounds obvious when you say it. Why do people still get this wrong?
[beat]
ERIK: Because they start with the model instead of the workflow. They ask, "Which frontier model should run my company?" Wrong question. Start with, "What are the repeatable decisions, what context is needed, and what proves success?" Models come after that. If you reverse it, you get expensive demos.
[beat]
JOSH: Give me a concrete example.
[beat]
ERIK: PrimeBus catches a failed test run at two in the morning. First pass classifier reads the failure shape. If it's a flaky selector in Selenium, that's one lane. If it's a dependency conflict, that's another. Planner says what kind of fix should exist. Executor writes two variants. Verifier runs tests. Winner gets merged if the checks are clean. That's how you get 214 auto-fixes instead of 214 Slack threads.
[beat]
JOSH: And the failure cases?
[beat]
ERIK: Plenty. Wrong root cause. Patch passes local tests and still misses the real issue. Model edits too much. Fix works but introduces subtle drift. That's why verification matters more than the generated patch. No guardrails, no autonomy. Simple.
[beat]
JOSH: Where does cost enter the picture?
[beat]
ERIK: Immediately. If you put the expensive brain on the cheap task, you're lighting money on fire. ScanBrief is a good example. It scans 56 sources and scores 396 items every morning for a 6:30 delivery. No sane person wants premium reasoning on every single raw item. Most of that work is filter, dedupe, rank. Save the expensive thinking for the few places where nuance actually changes the outcome.
[beat]
JOSH: So planner and executor isn't just better quality. It's cost control too.
[beat]
ERIK: Correct. Quality, speed, and cost all move if the roles are clean. Everybody wants one model to be an employee. Better move is build a small team out of narrow roles and make them prove their work.
[beat]
JOSH: What should builders stop doing today?
[beat]
ERIK: Stop making one giant prompt that says "think hard, write the code, test it, explain it, and make it perfect." That's toddler architecture. Break the work apart. Route. Plan. Execute. Verify. Keep receipts.
[pause]
JOSH: Last one. The button story. Mercedes bringing physical controls back sounds small, but people reacted hard.
[beat]
ERIK: Because everybody knows it was dumb. Car companies spent years confusing "looks modern" with "works well." A slider on a screen is fine until your hands are cold, the road sucks, sunlight hits the panel, and you need defrost now.
[beat]
JOSH: This isn't really about cars, is it?
[beat]
ERIK: No. It's about interfaces built for demos instead of use. Same disease shows up in enterprise tools all the time. Teams hide important actions behind layers because the screen looks clean in a screenshot. Then an operator needs two extra clicks during an incident and suddenly "minimal" feels stupid.
[beat]
JOSH: How does that map to AI products?
[beat]
ERIK: Same exact problem. People are building shiny chat wrappers when the user really needs a button that says classify this alert, write the patch, compare these configs, check drift, rerun the failed job. Utility beats theater. Always.
[beat]
JOSH: So a physical button is kind of the interface version of a good runbook.
[beat]
ERIK: That's a good way to say it. Known action. Predictable result. Low cognitive load. Runbooks should feel like that. Agents should feel like that too. If the operator has to wonder what the system is about to do, you've already lost trust.
[beat]
JOSH: You think this swings back hard?
[beat]
ERIK: In serious tools, yes. Fast. The winners won't be the prettiest apps. It'll be the ones that let people act with confidence. In my lab, every useful tool got simpler over time. More automation underneath. Fewer choices on top. That's not sexy. It works.
[beat]
JOSH: Dry way to say the touchscreen lost.
[beat]
ERIK: Glass had a good run. Defrost won.
[pause]
ERIK: This episode is sponsored by Prime Automation Solutions. If you're still doing it manually, we automate it. Also, special on a website — $250. primeautomationsolutions.com
[pause]
JOSH: Alright, what's the AI pro tip today?
[beat]
ERIK: Split planner from executor. Always. Use one model to think about the shape of the problem and produce a tight plan. Then hand that plan to a second model whose only job is execution. Code, content, query, config, whatever. Don't ask the same model to brainstorm, decide, build, and judge in one pass unless the task is tiny.
[beat]
ERIK: Reason is simple. Planning and doing are different jobs. When you separate them, prompts get shorter, output gets cleaner, and failure is easier to debug. If the plan is bad, fix the planner. If the code is bad, fix the executor. On PrimeBus, that separation is a big chunk of why the auto-fix loop stays sane.
[beat]
ERIK: Practical version. Prompt one: "Classify the failure, identify the likely root cause, and return a five-step repair plan." Prompt two: "Execute only step three using this repo context and change the minimum needed." Prompt three: "Verify with tests and explain any mismatch." That's your tip. Use it.
[pause]
ERIK: That tip is straight out of The Autonomous Engineer — my book on building systems that run themselves. Grab it on Amazon.
[pause]
JOSH: One more thing — we started a Discord for builders. If you're shipping AI, automation, or anything that makes a human obsolete — come hang out. Link at buildorbereplaced.dev.
ERIK: Post what you built. We'll post what we're building. Real wins, real builds, no fluff.
[pause]
ERIK: Build or be replaced.
JOSH: If you want these signals in your inbox every morning, scanbrief.dev. See you tomorrow.