Mon–Fri · 6 AM ET
← All Episodes
EP  • 00:14:01

House with 15m underground tunnels for sale for 300k | Build or Be Replaced

Today: House with 15m underground tunnels for sale for 300k | “Math 2.0” will need to value mathematical progress more holistically | Show HN: Bigwords.page – Turn any screen into a sign. The URL is the app Episode date: 2026-10-08.

Download MP3 →

Transcript

JOSH: It's Thursday, October 8. This is Build or Be Replaced — powered by ScanBrief.dev. I'm Josh, here with Erik Anderson.
ERIK: Episode 123, and the theme today is simple. Cheap fast agents are coming for every boring workflow you still defend with a spreadsheet.
JOSH: Stick around — Erik's got an AI pro tip at the end about using small models as task routers.
[pause]
JOSH: First headline. Anthropic has Claude Haiku 5.5 showing up in ScanBrief today. Fast, cheaper, and apparently smarter. What matters there?
ERIK: Haiku-class models are the workhorses. Nobody should be using a giant model to classify log lines, score tickets, or route work. If Haiku 5.5 gives you lower token cost and adjustable effort, that means more agent runs for the same bill.
[beat]
JOSH: Second one. Docker Agent is on the board. Is this just another agent wrapper?
ERIK: Maybe. But Docker has the right surface area. Build, run, inspect, logs, containers, images. If they make agents container-native instead of chat-window-native, that gets interesting fast.
[beat]
JOSH: Third headline. Margaret Hamilton has died. Big name in software history.
ERIK: Huge name. Apollo guidance software, software engineering before people really respected software engineering. Every time somebody says software is secondary to the hardware, go read what happened when software had to land humans on the moon.
[pause]
JOSH: ScanBrief scored 97 items across 56 sources today. The Haiku item landed near the top. Why does a smaller model release matter more than a giant benchmark headline?
ERIK: Because most real automation isn't one giant genius thinking for five minutes. It's a lot of small decisions all day. Classify this event. Decide if this test failure is flaky. Pull the right runbook. Pick which agent gets the job. Summarize the diff. Check if the output smells wrong.
[beat]
ERIK: That's not trophy-model work. That's Haiku work. That's small model, low latency, low cost, high volume. People keep trying to build one big brain. Systems people build traffic control.
JOSH: Traffic control for agents?
ERIK: Exactly. PrimeBus is built around that idea. Events hit the bus. Agents subscribe. Some agents review code. Some agents inspect failures. Gandalf blocks unsafe merges. The auto-merger has run 2295 attempts since June 5. 1507 merged, 788 blocked by Gandalf, 0 escalated to me.
[beat]
ERIK: That blocked number is the story. People hear blocked and think failure. Wrong. That's the guardrail doing its job before a bad change hits prod.
JOSH: Wait, zero escalated to you?
ERIK: Yep. Zero escalated to me in that stat block. That's not because every change is perfect. It's because the system has boundaries. If confidence is bad, it blocks. If tests don't line up, it blocks. If the diff touches something sketchy, Gandalf says no.
JOSH: That's a pretty good name for the gatekeeper.
ERIK: It earned it.
[beat]
JOSH: So where does Haiku 5.5 fit into that kind of setup?
ERIK: Routing first. You don't need a frontier model to read an event payload and say, this is a test failure, this is a lint failure, this is a dependency issue, this is a stuck job. A small model can do that. Then you send only the hard part to Claude Sonnet or Opus or whatever the right heavier model is.
JOSH: So the small model is the dispatcher.
ERIK: Dispatcher, filter, first-pass reviewer, log summarizer. Pick your term. The point is it sits in front of the expensive work. That changes the math.
[beat]
ERIK: In network automation, same thing. If NSO throws a service error, I don't need a giant model to tell me the device timed out. I need something fast to classify the error, attach the runbook, check recent telemetry, and decide if this is safe to retry. Then, if the config intent is weird, bring in the bigger model.
JOSH: That sounds less magical than the demos.
ERIK: Good. Magical demos usually die in production. Boring systems live.
[pause]
JOSH: You said adjustable effort matters too. Why?
ERIK: Because not every task deserves the same amount of thinking. If I'm scoring ScanBrief items, I can run cheap and fast. Today it scored 97 items across 56 sources. Most of that work is ranking, dedupe, summarization, and signal detection. That's perfect for a smaller model.
[beat]
ERIK: But if the brief finds a model release that could change my pipeline economics, I want a deeper pass. Compare pricing, read the system card if it exists, test latency, run the model against a few real tasks. Adjustable effort lets you spend thought where it pays.
JOSH: Like QoS for reasoning.
ERIK: Exactly. That's the network guy answer. Not all packets are equal. Not all prompts are equal either.
JOSH: That's clean.
ERIK: It should be. We've had traffic engineering forever. AI people are rediscovering queues, priorities, retries, dead letters, and rate limits, then giving them new names.
[beat]
JOSH: Dry but fair.
ERIK: Very fair.
[pause]
JOSH: Let's move to Docker Agent. You said Docker has the right surface area. What does that mean?
ERIK: Agents need a place to work. They need a filesystem, tools, network rules, secrets policy, logs, rollback, and a clean way to die when they do something dumb. Containers already give you a lot of that.
JOSH: So instead of an agent poking around your laptop, it works inside a box.
ERIK: Yep. And that box has rules. That's the part people skip. They put an agent on a machine with broad access and then act shocked when it edits the wrong file or burns tokens in a loop.
[beat]
ERIK: Docker as an agent runtime makes sense because you can say, here's the repo, here's the test command, here's the env, here's the timeout, here's the network policy. Go solve this one thing. Then throw the container away.
JOSH: That feels like CI with a brain.
ERIK: That's basically what it is. CI with a brain and a credit card, which is why you need guardrails.
JOSH: There it is.
ERIK: Always. Agents are junior engineers with infinite patience and no common sense unless you give them structure.
[beat]
JOSH: How would you use Docker Agent in your own setup?
ERIK: First place would be isolated repair attempts. PrimeBus sees a failing job. It emits the event through NATS. The repair agent spins up a container with the repo at that commit, runs the failing test, asks Claude for two fix variants, applies each variant in separate containers, runs tests, compares results, and only then proposes a merge.
JOSH: That's very specific.
ERIK: Because vague agents are expensive screen savers.
[beat]
ERIK: The important piece is isolation. I don't want an agent fixing one repo while accidentally reading another repo's secrets or changing shared state. Containers make the work repeatable. If the fix passes once in a clean container, I trust it more.
JOSH: And if it fails?
ERIK: Emit telemetry, store the logs, block the merge, move on. No drama. PrimeSentinel can watch for stuck jobs. PrimeRestorer handles backups. The point isn't that nothing breaks. The point is the system notices faster than I do.
JOSH: That's the part people miss. The agent isn't the whole system.
ERIK: Right. The agent is a worker. The system is the bus, policies, tests, logs, rollbacks, metrics, secrets, and human approval where needed.
[pause]
JOSH: Does Docker becoming more agent-friendly make this easier for normal teams?
ERIK: Yes, if they treat it like infrastructure. No, if they treat it like a toy.
[beat]
ERIK: A normal team should start with one boring workflow. Don't say, our agent will manage the platform. Say, our agent will reproduce failed tests in a container and write a summary with the exact command, failing file, suspected cause, and suggested owner.
JOSH: That's small.
ERIK: Small is how you win. Once that works, let it draft a fix. Once that works, let it open a PR. Once that works, let it merge low-risk changes after tests pass. Step by step.
JOSH: Not straight to production auto-pilot.
ERIK: Please don't. That's how you end up in a meeting with legal and a very quiet Slack channel.
[beat]
JOSH: Painful image.
ERIK: Accurate image.
[pause]
JOSH: The third thread I want to hit is Margaret Hamilton and this Math 2.0 story. They feel different from model releases and Docker tools, but they're both about how we measure progress. Am I reaching?
ERIK: No, that's the connection. Margaret Hamilton proved that software discipline mattered before the industry had good language for it. Math 2.0 is arguing that progress shouldn't just be theorem count or publication count. Same pattern. The measurement system lags the actual work.
JOSH: Give me the engineering version.
ERIK: In software, people still measure output like it's 1998. Tickets closed. Lines changed. Meetings attended. Velocity charts. That's not engineering quality. That's office weather.
[beat]
ERIK: Real progress is whether the system gets safer, faster, more observable, easier to recover, and less dependent on one person being awake.
JOSH: That hits.
ERIK: It should. I have 146 services running on the production server right now. If my metric is, Erik worked hard today, that's useless. I need to know which services emitted telemetry, which jobs failed, which fixes merged, which ones Gandalf blocked, which backups completed, and which signals ScanBrief ranked high enough to care about.
JOSH: So Math 2.0, but for builders.
ERIK: Exactly. Builder metrics need to value durable progress. Did you add a test that catches a real failure? Did you move a manual runbook into code? Did you put events on a bus? Did you add a dead-letter queue? Did you make restore boring?
[beat]
ERIK: PrimeRestorer is not glamorous. Nobody claps for backup orchestration until the database is gone. Then it's suddenly everyone's favorite project.
JOSH: That's always how backups work.
ERIK: Yep. Ignored until sacred.
[pause]
JOSH: Margaret Hamilton's work was life-or-death software. Most builders aren't landing spacecraft. What should they take from that?
ERIK: Respect the boring parts. Interfaces. Error handling. Priority scheduling. Logs. Human-readable state. The Apollo software had to handle overload and still protect the mission. That's the lesson.
[beat]
ERIK: Your SaaS app isn't Apollo. Your Kubernetes cluster isn't Apollo. But the principle holds. When the system is under stress, what gets dropped, what gets preserved, and who gets paged?
JOSH: That's a better question than, does it work in the demo?
ERIK: Demos are sunny-day driving. Production is rain, night, bad tires, and somebody changed DNS.
JOSH: That's too real.
ERIK: Networking gives you scars.
[beat]
JOSH: How do you measure that in your own work?
ERIK: I care about closed loops. Event happens, system sees it, system classifies it, system acts or blocks, system records the outcome. That's the loop.
[beat]
ERIK: GovContractScanner is a good example. It doesn't just scrape opportunities and make a pile. It tracks primes, watches new federal opportunities, and turns that into something I can act on. A pile of data is not a system. A scored, routed, auditable flow is a system.
JOSH: That's the difference between collecting and building.
ERIK: Exactly. Collecting is easy. Building means the next action is obvious.
[pause]
JOSH: I want to go back to something you said. Cheap models make more agent runs possible. Is there a danger there?
ERIK: Absolutely. Cheap bad automation is still bad automation. Lower cost means people will run more junk. More junk means more noise, more false confidence, more random PRs, more alerts nobody reads.
JOSH: So cost going down doesn't remove discipline.
ERIK: It increases the need for discipline. When inference was expensive, the bill slowed you down. When inference gets cheap, architecture has to slow the dumb stuff down.
[beat]
ERIK: That's why I like blocked counts. PrimeBus having 788 blocked by Gandalf tells me the guardrail is active. If everything merges, I don't trust it. Perfect pass rates are suspicious.
JOSH: That's counterintuitive.
ERIK: Only if you haven't operated systems. A healthy filter rejects things. Spam filters reject mail. Firewalls drop traffic. CI fails builds. A merge guardrail should block work that doesn't meet the bar.
JOSH: And if humans complain?
ERIK: Good. Then you tune the policy with evidence. Show me the blocked diff, the test result, the risk label, and the reason. Don't bring vibes to an engineering argument.
[pause]
JOSH: For a team listening today, what's the practical move from all of this?
ERIK: Pick one workflow where decisions repeat. Ticket triage. Failed tests. Security issue summaries. Device config audits. Contract opportunity scoring. Anything repetitive with clear inputs and outputs.
[beat]
ERIK: Put the inputs in a queue. Use a small model to classify. Use a bigger model only when needed. Run the action in a container. Log every decision. Add a block condition. Then review the blocked work every Friday.
JOSH: That's very doable.
ERIK: It is. People make this mysterious because mystery sells better. The actual work is plumbing, tests, and logs.
JOSH: The glamorous stuff.
ERIK: Listen, NATS subjects and clean JSON payloads are beautiful. You either understand that or you haven't been hurt enough yet.
[pause]
ERIK: This episode is sponsored by Prime Automation Solutions. If you're still doing it manually, we automate it. Also, special on a website — $250. primeautomationsolutions.com
[pause]
JOSH: Alright, what's the AI pro tip today?
ERIK: Use a small model as a router before you use a big model as a solver.
[beat]
ERIK: Make a simple schema. Task type, risk level, required tools, confidence, next model. Feed it the event, not your whole repo. Let the small model decide whether this is summarize, classify, ignore, escalate, or solve.
[beat]
ERIK: Then enforce the decision in code. If risk is high, no write tools. If confidence is low, human review. If task type is log summary, don't call the expensive model. If task type is code repair, run it in a container with tests.
[beat]
ERIK: This saves money, but the bigger win is control. Your AI system stops being one giant prompt blob and starts acting like infrastructure. That's your tip. Use it.
[pause]
JOSH: Track your freedom score and net worth with the Freedom Blueprint app — free download, link in the show notes.
[pause]
JOSH: One more thing — we started a Discord for builders. If you're shipping AI, automation, or anything that makes a human obsolete — come hang out. Link at buildorbereplaced.dev.
ERIK: Post what you built. We'll post what we're building. Real wins, real builds, no fluff.
[pause]
ERIK: Build or be replaced.
JOSH: If you want these signals in your inbox every morning, scanbrief.dev. See you tomorrow.