Transcript
JOSH: It's Friday, August 21. This is Build or Be Replaced — powered by ScanBrief.dev. I'm Josh, here with Erik Anderson.
ERIK: Episode 93, ScanBrief scored 90 items across 56 sources, and the signal today is simple: the stack is getting smarter, but the blast radius is getting bigger.
JOSH: Stick around — Erik's got an AI pro tip at the end about using a second model as your reviewer.
[pause]
JOSH: First headline. Hacker News had a project called Huzzah, a new way to code with AI. Is that actually new, or just another wrapper around chat?
[beat]
ERIK: The wrapper matters if it changes the feedback loop. Chat is fine for asking questions. Building needs state, diffs, tests, and memory of what already broke.
JOSH: So the story isn't the UI.
ERIK: Correct. The story is that AI coding tools are moving from autocomplete to working sessions. That's where builders win or get replaced by somebody who knows how to drive the agent.
[pause]
JOSH: Second headline. A malicious Rust crate called Arrayref ran a build-time payload. That's pretty ugly.
[beat]
ERIK: Build-time code execution is the nightmare hiding in plain sight. You install a dependency, the package runs before your app even exists, and now your laptop is part of the supply chain story.
JOSH: Rust people don't like hearing that.
ERIK: Nobody likes hearing it. Doesn't matter if it's Rust, npm, PyPI, or a shell script named install.sh. Trust is not a package manager feature.
[pause]
JOSH: Third headline. HTML can do more natively now. Popovers, dialogs, accordions, command buttons. Why does that matter?
[beat]
ERIK: Because half the JavaScript on the web is duct tape for things the browser should have handled years ago. Native dialog and popover behavior means fewer dependencies, fewer bugs, and less code for AI to misunderstand.
JOSH: Less code sounds like the theme.
ERIK: Less unnecessary code. Big difference.
[pause]
JOSH: Let's start with the AI coding tools. Huzzah got attention, and there was also a story about cleaning up Claude 5 token output with a separate LLM. What's the bigger signal?
ERIK: The bigger signal is that people are finally accepting that one model call is not a workflow. It's one move on the board. Useful, but not the system.
[beat]
ERIK: Coding with AI gets interesting when you stop treating the model like a smart intern in a chat window and start treating it like a worker on a bus. It gets a task, it checks context, it makes a change, it runs tests, it reports evidence, and another worker reviews it.
JOSH: That's basically how you run PrimeBus, right?
ERIK: Yeah. PrimeBus is NATS underneath, event-driven, with agents subscribed to different kinds of messages. Test failures, warnings, deploy events, merge candidates, all of it goes on the bus. Since June 5, the PrimeBus auto-merger has run 2295 attempts. 1507 merged, 788 blocked by Gandalf, 0 escalated to me. That's the number that matters.
JOSH: Zero escalated is the part that jumps out.
ERIK: Exactly. The blocked count is not a failure. That's the guardrail doing its job. Gandalf sees something it doesn't like, blocks the merge, and the system keeps moving without waking me up. That's how you build confidence. Not by pretending the AI is perfect. By assuming it isn't.
[beat]
JOSH: Where does Huzzah fit into that?
ERIK: Any tool that makes the coding session more structured is moving in the right direction. The old flow was: ask model, copy code, paste code, pray. That's not engineering. That's gambling with better syntax highlighting.
JOSH: And the newer flow?
ERIK: The newer flow is: give the agent a bounded job, let it inspect the repo, make the smallest change, run the check, and explain the result. Then you decide whether it earned the next task.
JOSH: That sounds slower than just asking for the whole feature.
ERIK: It feels slower until the first giant AI patch breaks authentication, billing, and CSS in the same commit. Then suddenly small steps feel pretty good.
[beat]
JOSH: What about that "Vomit" story? Cleaning Claude output with another LLM sounds funny, but is it useful?
ERIK: Funny name, useful idea. Model output has waste in it. Extra prose, half-formed reasoning, tool chatter, formatting junk. If you're building an agent pipeline, you don't want every downstream step reading a novel.
JOSH: So one model writes, another cleans?
ERIK: Sometimes. Or one model proposes, another extracts the patch. One model summarizes logs, another decides if the logs are clean. Different jobs need different prompts and sometimes different models. That's normal systems design.
JOSH: Wait, really? You're saying the answer to AI mess is more AI?
ERIK: More control. Not more chaos. A second model can be a parser, critic, classifier, or judge. The mistake is asking one giant prompt to do planning, coding, testing, review, and release notes all at once. That's how you get confident garbage.
[beat]
ERIK: ScanBrief works the same way. It doesn't just pull links and dump them into a newsletter. It collects, scores, ranks, dedupes, summarizes, and then builds the brief. Today it scored 90 items across 56 sources. Each step is boring on purpose. Boring steps are reliable steps.
JOSH: That's a good line.
ERIK: It's also how real infrastructure survives Monday morning.
[pause]
JOSH: Okay, second deep story. Malicious Rust crate. People hear "dependency attack" and think, yeah, that happens somewhere else. What's the real risk?
ERIK: The real risk is that modern software builds are basically remote code execution parties with nice logos.
[beat]
ERIK: A package manager pulls code from the internet. The package runs build scripts. Those scripts can read environment variables, touch files, phone home, grab tokens, modify generated code. If your CI has secrets, that package may get a shot at them.
JOSH: That's bleak.
ERIK: It's accurate. Bleak is optional.
JOSH: What would you do differently?
ERIK: First, lock dependencies. No floating versions in serious systems. Second, separate install from test from deploy. Third, build in an environment where secrets are not present until they're needed. Fourth, log weird network calls during builds.
[beat]
JOSH: That sounds like a lot for a small project.
ERIK: Small projects become real projects the day somebody pays you or depends on you. PrimeProspector can scan leads and follow up automatically. Cool. But if the build system leaks an API key, now the automation is mailing strangers on behalf of an attacker. That's not a feature I want.
JOSH: Fair.
ERIK: Same with PrimeTrader. TradingView webhooks come in, logic fires, decisions get made. Any system that touches money, identity, customer messages, or infrastructure needs a tighter supply chain than "seems fine from GitHub stars."
JOSH: How do AI agents make this worse?
ERIK: Agents install things fast. That's the problem. A human might pause on a weird dependency. An agent sees an import error and says, "I can fix that," then adds a package. Helpful little disaster.
[beat]
JOSH: So block installs?
ERIK: Not all installs. Gate them. Agent can propose a dependency, but the system should ask: who owns it, when was it published, does it run build scripts, does it have native code, does it touch postinstall, does it match a known package name, did it appear this week.
JOSH: That's the Gandalf role again.
ERIK: Yep. Gandalf doesn't need to be dramatic. It just needs to say no at the right time. That is the most valuable word in automation.
JOSH: No as a feature.
ERIK: Best feature in production.
[pause]
JOSH: Third deep story. HTML can do that. This sounds smaller than malicious packages and AI agents, but you got more interested than I expected.
ERIK: Because boring platform features remove entire classes of bugs.
[beat]
ERIK: Native popovers, dialogs, exclusive accordions, command buttons, these are not flashy. Nobody is raising a funding round for "button opens thing." But every custom dropdown, modal, focus trap, and click-outside handler is code that can break.
JOSH: And code AI has to read.
ERIK: Exactly. AI is better when the system is simpler. If your app uses native HTML behavior, the model can reason about it. If your app has four custom modal libraries, two state managers, and a div pretending to be a button, the model starts guessing.
JOSH: Frontend people are going to feel attacked.
ERIK: Good. Use buttons.
[beat]
JOSH: What's a real example?
ERIK: Take a settings panel. Old way: React state for open or closed, event listener on document, escape key handler, focus management, outside click handler, aria attributes, cleanup on unmount, maybe a portal. Newer way: native dialog or popover where the browser handles a chunk of it.
JOSH: That's less clever.
ERIK: That's the point. Clever code is expensive. Simple code ships.
JOSH: How does that show up in your systems?
ERIK: HumanDesignApp has an iMessage state machine. The hard part is not drawing the screen. The hard part is keeping user state sane across messages, retries, and interpretation steps. If the UI can be plain and dependable, more attention goes to the actual state machine.
[beat]
ERIK: Same with internal tools. A dashboard for 145 services running on the production server right now does not need a circus. It needs search, status, filters, logs, restart buttons, and confirmation dialogs that work every time. Native primitives help.
JOSH: That's a lot of services.
ERIK: It is. Which is why I don't want a modal bug eating my morning.
[pause]
JOSH: There's also a thread today about intermediate tokens not really being reasoning traces. Why does that keep coming up?
ERIK: Because people want the model's scratchpad to be a legal deposition. It isn't. Tokens are computation artifacts. Useful sometimes, misleading other times.
JOSH: So don't treat "thinking" text as truth?
ERIK: Correct. Treat outputs as claims. Verify them. If an agent says tests passed, make it show the command and result. If it says it checked a file, make sure the file was actually read. If it says a dependency is safe, ask what evidence it used.
[beat]
JOSH: That changes how people should prompt.
ERIK: It changes how people should build. Prompting is not enough. You need tools, logs, permissions, test gates, review gates, and rollback. The model is a component. Not the whole machine.
JOSH: That's probably the most practical AI take today.
ERIK: Duuude, it's the only one that matters. Stop worshiping the text. Build the harness.
[pause]
JOSH: Quick check. If someone listening wants to improve their AI coding setup this weekend, what's the first move?
ERIK: Add a review step that the coding agent cannot skip. Doesn't have to be fancy. One prompt writes the patch. Second prompt reviews only the diff and the test output. Different context. Different job. If the reviewer finds risk, block the merge.
JOSH: Even for solo builders?
ERIK: Especially for solo builders. Solo builders are the easiest people to fool because nobody is there to say, "why did you add a package named arrayref at 1:13 in the morning?"
[beat]
JOSH: That one feels personal.
ERIK: Every security story is personal if you have enough cron jobs.
[pause]
ERIK: This episode is sponsored by Prime Automation Solutions. If you're still doing it manually, we automate it. Also, special on a website — $250. primeautomationsolutions.com
[pause]
JOSH: Alright, what's the AI pro tip today?
ERIK: Use a second model pass as a diff-only reviewer.
[beat]
ERIK: Here's the exact pattern. First agent gets the task and edits the code. Then freeze the patch. Second agent gets only three things: the user request, the diff, and the verification output. Tell it: find bugs, security risk, missing tests, and behavior changes. Tell it not to rewrite anything. Review only.
JOSH: Why diff-only?
ERIK: Because full context makes reviewers wander. Diff-only keeps it honest. The reviewer has to judge what changed. If it needs more context, it asks for a specific file. That's how you avoid the model writing a book report instead of finding the bad line.
[beat]
ERIK: Use a cheaper fast model for routine review, and a stronger model when the change touches auth, billing, deploy, secrets, or data loss. Put that rule in your script. Don't make it a mood decision.
JOSH: Builder can do that today.
ERIK: Today. Git diff, test output, second model, block on risk. That's your tip. Use it.
[pause]
JOSH: Binge all five episodes this weekend plus our YouTube shorts — links at buildorbereplaced.dev.
[pause]
JOSH: One more thing — we started a Discord for builders. If you're shipping AI, automation, or anything that makes a human obsolete — come hang out. Link at buildorbereplaced.dev.
ERIK: Post what you built. We'll post what we're building. Real wins, real builds, no fluff.
[pause]
ERIK: Build or be replaced.
JOSH: If you want these signals in your inbox every morning, scanbrief.dev. See you tomorrow.