Mon–Fri · 6 AM ET
← All Episodes
EP  • 00:12:00

Show HN: Audiomass – a free, open-source multitrack audio editor for the web | Build or Be Replaced

Today: Show HN: Audiomass – a free, open-source multitrack audio editor for the web | DeepSeek reasonix, DeepSeek native coding agent with high caching and low cost | Migrating from Go to Rust Episode date: 2026-05-25.

Download MP3 →

Transcript

JOSH: It's Monday, May 25. This is Build or Be Replaced — powered by ScanBrief.dev. I'm Josh, here with Erik Anderson.
ERIK: Cheap agents are coming, memory is eating the hardware budget, and backend code is still where weak prompts go to die.
JOSH: Stick around — Erik's got an AI pro tip at the end about making constraints agents can't quietly forget.
JOSH: [pause]
JOSH: First headline. DeepSeek has a native coding agent called Reasonix, and builders are watching the price. What's the signal?
ERIK: Cost changes behavior. If a coding agent is cheap enough and the cache is good enough, you don't run it once. You run it constantly, against small problems, before they become big ones.
JOSH: Second. High-bandwidth memory is reportedly around 63 percent of AI chip component cost. Why should a software person care?
ERIK: Because your AI bill is not magic. The model is expensive because the hardware is expensive. Memory is becoming the toll booth.
JOSH: Third. There's new work on constraint decay in backend code generation. That's a very academic way to say agents forget stuff, right?
ERIK: Yep. Give an agent ten requirements and a long enough task, and by the end it may satisfy the pretty five and forget the dangerous five. That's not a prompt problem. That's an engineering problem.
JOSH: [pause]
JOSH: Reasonix first. A cheaper coding agent sounds like the thing everyone has been waiting for.
ERIK: It might be. The headline isn't "another model can write code." We've had that. The headline is run cost.
JOSH: Why does cost matter that much?
ERIK: Because one run is weak. One run is a suggestion. Real agentic work is attempts, tests, review, comparison, rollback, retry, and then maybe a merge.
ERIK: That's how I run my stuff. PrimeBus moves the events. Gandalf reviews the code. PrimeSentinel watches for stuck jobs. Nothing gets trusted because it sounds confident.
JOSH: So cheaper agents don't mean fewer guardrails.
ERIK: No. They mean more chances to use the guardrails. That's the whole point.
JOSH: [beat]
JOSH: Give me a concrete example.
ERIK: Say a service throws a test failure after a dependency update. The basic version is a human sees the failed build, opens the logs, asks Claude for help, copies a patch, runs tests, fixes the fix, then pushes.
ERIK: The better version is PrimeBus sees the failure event. It includes the repo, branch, failed command, stack trace, changed files, and owner. An agent subscribes to that event, proposes a fix, tests it, and Gandalf reviews the diff. If it can't prove the change, it stops.
JOSH: That's the factory idea.
ERIK: Duuude, exactly. Not one heroic chat window. A bunch of small workers reacting to events. NATS is boring. That's why it works. Boring pipes, smart edges.
JOSH: Where would Reasonix fit in that stack?
ERIK: First-pass repair. Test fixes. Dependency bumps. Small refactors. Documentation drift. Maybe code search plus patch generation. I would not hand it a scary Terraform migration and go make lunch.
JOSH: What gets the stronger model?
ERIK: Architecture. Reviews. Anything with money, auth, customer data, network reachability, or blast radius. Stronger models are for judgment. Cheaper models are for repetition.
JOSH: That's a clean split.
ERIK: It has to be. Teams get weird about this. They pick one model like it's a religion. That's not engineering. That's merch.
JOSH: Wait, merch?
ERIK: Yeah. If the hoodie says the model name, people start making procurement decisions with feelings. Use the right tool. Claude for deep code reasoning. Codex where it fits. A cheaper agent for boring patches. Selenium for browser verification. NSO for network intent. Terraform for state. Kubernetes for workloads. Don't make a screwdriver feel included.
JOSH: That's going on a mug.
ERIK: It would be an ugly mug, but accurate.
JOSH: [beat]
JOSH: What's the risk with cheaper agents?
ERIK: Teams will make bad code faster. That's always the first risk.
JOSH: That sounds bleak.
ERIK: It's honest. Cheap compute doesn't fix weak process. If your current process is "a developer vibes a patch and nobody reads it," agents will make that worse. Faster junk is still junk.
JOSH: So what's the correct move?
ERIK: Put the agent inside a system. The agent is not the system. The bus is the system. The tests are the system. The reviewer is the system. The policy gate is the system. The rollback path is the system.
ERIK: ScanBrief works the same way. It doesn't trust one source. It pulls, scores, ranks, dedupes, and then the brief comes out. The value isn't one summary. It's the pipeline.
JOSH: That matters for news and code.
ERIK: Same pattern. Inputs are messy. Models are useful. Outputs need checks. If you remember nothing else, remember that.
JOSH: [pause]
JOSH: Let's move to hardware. HBM being 63 percent of component cost sounds like inside baseball. Is it?
ERIK: No. It's the reason your agent pricing looks the way it does.
JOSH: How so?
ERIK: AI workloads are hungry for memory bandwidth. Not just memory capacity. Bandwidth. Those chips need to move huge tensors fast. HBM sits close to the accelerator and feeds it. If the memory can't keep up, the expensive compute sits there waiting.
JOSH: So the chip is not the whole story.
ERIK: Correct. Everyone talks about GPUs like they're magic rocks. The package matters. The memory matters. The interconnect matters. Power matters. Cooling matters. The datacenter matters. Then somebody in finance sees the invoice and needs a chair.
JOSH: That's where software builders feel it.
ERIK: Yep. Model vendors don't price from vibes. They have hardware costs, utilization problems, supply limits, and customers slamming APIs all day. If HBM is eating that much of the component cost, cheap inference has a physical ceiling.
JOSH: Does that make local AI more attractive?
ERIK: For some jobs, yes. Not everything. Local models are great when privacy matters, latency matters, or the task is small and repeatable. But don't pretend a home lab replaces a frontier model for hard reasoning.
JOSH: You run a lab. Where do you draw that line?
ERIK: Local and small for routing, classification, extraction, watchdog logic, and quick summaries. Strong remote models for code review, architecture, bigger writing, and anything where the cost of being wrong is high.
JOSH: That's less romantic than "run everything local."
ERIK: Romance is expensive in infrastructure.
JOSH: Fair.
ERIK: I love self-hosting. I run my own systems because I want control. PrimeDash shows me what's alive in the lab. PrimeSentinel tells me when a job gets stuck. But if Claude is better for a hard coding pass, I use Claude. The goal is shipping working systems, not winning a purity contest.
JOSH: How does memory cost change what builders should do today?
ERIK: Stop wasting context. That's the boring answer and it saves real money.
JOSH: Meaning?
ERIK: Don't paste the entire repo into every request. Don't ask the agent to reread files it doesn't need. Don't send screenshots when a DOM query would do. Don't make every task a novel.
ERIK: Build context packs. Small files that tell the agent the architecture, commands, test rules, and project conventions. Use retrieval when you need it. Cache repeatable context. Keep prompts tight.
JOSH: That's basically engineering your AI bill.
ERIK: Exactly. And it helps quality. Smaller, cleaner context means fewer distractions. Agents don't need your whole life story to fix a failing unit test.
JOSH: [beat]
JOSH: Is this where teams get surprised?
ERIK: They get surprised because chat hides cost. It feels like typing. But every token has a path through hardware somebody paid for. When you automate agents, that cost becomes a line item fast.
JOSH: What do you track?
ERIK: For my systems, I care about run reason, model used, repo, outcome, test result, review result, and whether a human had to step in. The money matters, but outcome matters more. A cheap failed run can be more expensive than one good expensive run.
JOSH: That's a good way to put it.
ERIK: Builders need cost per accepted change, not cost per prompt. Cost per useful result. That's the metric.
JOSH: [pause]
JOSH: The third story is constraint decay. Agents forgetting backend requirements. This one feels like it hits closer to production.
ERIK: It does. Backend work is full of invisible rules. Idempotency. Auth. Rate limits. Transactions. Cleanup. Secret handling. Audit logs. Retry safety. Schema migration order. Cache invalidation. Everybody's favorite little haunted basement.
JOSH: And the agent drops those as the task gets longer?
ERIK: It can. The longer the task, the more chances the model starts solving the local problem and loses the original contract. It may produce code that looks reasonable and still violates the one thing that mattered.
JOSH: Like what?
ERIK: "Add a webhook endpoint, but verify signatures, reject replays, write an audit record, and never log the raw payload." The agent might build the endpoint, parse JSON, return 200, and write a test for the happy path. Looks good. Also terrible.
JOSH: Because the dangerous requirements were not checked.
ERIK: Right. The prompt said them. The system didn't enforce them.
JOSH: That's the line.
ERIK: That's the whole episode. If the requirement matters, it needs to exist outside the prompt.
JOSH: Where?
ERIK: Tests. Types. Policy checks. CI gates. Review rubrics. Runtime monitors. Contract tests. Static analysis. Even a checklist file in the repo is better than a requirement trapped in chat history.
JOSH: So the fix is not a longer prompt.
ERIK: Longer prompts help until they become fog. Agents need crisp constraints. The task should say, "Do this. These files are in scope. These tests must pass. These behaviors are mandatory. Do not change these contracts." Then the repo should back that up.
JOSH: How do you do that with Gandalf?
ERIK: Gandalf reviews against rules. It looks for risky changes, missing tests, bad patterns, and whether the patch actually matches the ask. The key is the review is a separate step. The author agent doesn't grade its own homework.
JOSH: Humans are also bad at that.
ERIK: Exactly. That's why code review exists. Agents didn't invent overconfidence. They just made it available through an API.
JOSH: Nice.
ERIK: PrimeBus helps because every event carries structured context. Not a giant blob of "please fix this." It sends what happened, where it happened, what changed, what command failed, and what decision is needed. Agents do better when the job is shaped.
JOSH: What's the beginner version for someone without your whole setup?
ERIK: Create an AGENTS file or a repo rules file. Put commands in it. Put forbidden patterns in it. Put testing expectations in it. Then make your agent read it before work and again before final review.
JOSH: Read it twice?
ERIK: Yep. First for direction. Last for compliance. Constraint decay is real, so refresh the constraints near the decision point.
JOSH: That sounds almost too simple.
ERIK: Simple is fine. I like simple. Simple runs at 2 AM.
JOSH: [beat]
JOSH: What should they not do?
ERIK: Don't ask the same agent, "Did you follow all the rules?" and accept "yes." That's theater. Make it show evidence. Test output. Diff summary. Files changed. Risks. Anything it couldn't verify.
JOSH: Evidence over confidence.
ERIK: Every time. Confidence is not a build artifact.
JOSH: That's probably the episode title.
ERIK: Put it on the ugly mug with the screwdriver line.
JOSH: [pause]
ERIK: This episode is sponsored by Prime Automation Solutions. If you're still doing it manually, we automate it. Also, special on a website — $250. primeautomationsolutions.com
JOSH: [pause]
JOSH: Alright, what's the AI pro tip today?
ERIK: Put your non-negotiable constraints in a file the agent must reread before it finishes.
JOSH: What goes in that file?
ERIK: Keep it short. Commands to run. Security rules. Files it cannot touch. Required tests. Output format. Review checklist. Call it AGENTS.md, RULES.md, whatever your tools already support.
ERIK: Then use a two-pass prompt. First line: "Read the rules file before making changes." Last line: "Before final, reread the rules file and list evidence for each constraint."
JOSH: Why does that work?
ERIK: Because it fights drift at the end, where drift hurts most. The agent may start strong and wander. Make it reload the contract right before it declares victory.
JOSH: That's usable today.
ERIK: Yep. No platform migration. No fancy orchestration. One file, one habit, better agent behavior. That's your tip. Use it.
JOSH: [pause]
ERIK: That tip is straight out of The Autonomous Engineer — my book on building systems that run themselves. Grab it on Amazon.
JOSH: [pause]
JOSH: One more thing — we started a Discord for builders. If you're shipping AI, automation, or anything that makes a human obsolete — come hang out. Link at buildorbereplaced.dev.
ERIK: Post what you built. We'll post what we're building. Real wins, real builds, no fluff.
JOSH: [pause]
ERIK: Build or be replaced.
JOSH: If you want these signals in your inbox every morning, scanbrief.dev. See you tomorrow.