DeepSeek v4 says it can drop into OpenAI and Anthropic workflows without a… | Build or Be Replaced
AI news and automation insights for 2026-04-24. Episode date: 2026-04-24.
Download MP3 →
Build or Replaced
AI news and automation insights for 2026-04-24. Episode date: 2026-04-24.
Download MP3 →JOSH: It's Friday, April 24. This is Build or Be Replaced powered by ScanBrief.dev. I'm Josh, here with Erik Anderson. ERIK: Three stories today. One is useful, one is dangerous, and one is people confusing rumor with product. JOSH: Stick around — Erik's got an AI pro tip at the end about getting better results by splitting thinking from doing. [pause] JOSH: First headline. DeepSeek v4 says it can drop into OpenAI and Anthropic workflows without a rewrite. Real story or marketing line? ERIK: Real enough to matter. If your SDK calls keep working and the model swap is mostly config, teams test faster, and faster testing means cheaper model churn. [beat] JOSH: Next one. Anthropic had to address Claude Code quality reports. What happened? ERIK: They had three separate issues, fixed them by April 20, reset limits, and tuned defaults. Translation: if your agent felt weird this week, it probably wasn't your prompt, it was the stack. [beat] JOSH: Third headline. Bitwarden CLI got pulled into a Checkmarx supply chain mess. Why should builders care? ERIK: Because everybody trusts the security tool in the pipeline until the security tool becomes the threat. One poisoned dependency can turn your CI into a delivery truck for malware. [pause] JOSH: And the ones we’re not going deep on? ERIK: GPT-5.5 is rumor bait until there’s an actual release. Meta cutting 10 percent tells you big tech is still trimming for margin while pretending it’s strategy. Ubuntu 26.04 looks solid, especially TPM-backed disk encryption, but that’s a weekend migration story, not a morning panic story. [pause] JOSH: Start with DeepSeek. Why did that one jump out to you? ERIK: Because compatibility wins. Always. Nobody wakes up excited to rewrite every wrapper, every retry handler, every tool call, every model selector just to test one new model. If DeepSeek v4 can sit behind an OpenAI-style or Anthropic-style interface and behave close enough, builders will try it by lunch. [beat] JOSH: Close enough is doing a lot of work there. ERIK: It is. API compatibility is not behavior compatibility. Big difference. You can match the endpoint, the auth pattern, the SDK shape, and still break workflows because the model thinks differently, formats differently, calls tools differently, or times out in different places. [beat] JOSH: So why are people excited? ERIK: Because switching cost is where experiments go to die. In my lab I’ve got 64 projects running across two servers, neo and morpheus. I’m not rebuilding every pipeline because a model company says their vibes are better. I want a swap layer. PrimeBus pushes messages around the mesh, my agents subscribe, and I can change the model behind a service without detonating the rest of the system. [beat] JOSH: That’s how you keep shipping daily? ERIK: Exactly. Loose coupling. Boring plumbing. Everybody wants the magical agent. Nobody wants to talk about the adapter layer that keeps the magical agent from setting the house on fire. [beat] JOSH: What does DeepSeek changing old model names tell you? ERIK: It tells you they’re consolidating around modes instead of brands. Non-thinking and thinking modes under one family is cleaner operationally. Also dangerous if you’re sloppy. A lot of teams hard-code model names all over the place. Then the vendor renames one thing and half the scripts start throwing errors at 2:13 a.m. [beat] JOSH: You’ve seen that in the wild? ERIK: Constantly. Different vendors, same movie. One of the reasons PrimeBus exists is to stop model naming from leaking into every corner of the system. My apps post intent. Summarize this. Classify this. Generate fix candidates. Review diff. The service that fulfills that intent chooses the model. Then if I want Claude for code review and another model for headline scoring, I change one service, not 19 apps. [beat] JOSH: What’s the practical takeaway for a small team? ERIK: Stop binding business logic directly to a model string. Put a policy layer in front of it. Even if it’s a dumb JSON config on day one. `task_type: code_fix -> preferred_model: claude`, `task_type: short_summary -> preferred_model: cheaper_model`. That one move saves you later. [beat] JOSH: And if someone wants to try DeepSeek v4 this afternoon? ERIK: Do it in a contained path. Pick one workflow. Something measurable. Don’t say, “we’re migrating the whole company.” That’s how grown adults create outages. Take one task like support ticket classification or first-pass summarization. Run the same 100 inputs across your current model and DeepSeek. Compare latency, token cost, formatting consistency, and how often a human has to fix the output. [beat] JOSH: Real numbers matter there. ERIK: Always. In ScanBrief I score 396 items daily from 56 sources and deliver the brief at 6:30 a.m. If a model saves me money but misses the format window, it failed. If it’s brilliant but inconsistent with structured output, it failed. If it needs hand-holding every few calls, it failed. Cheap wrong answers are still expensive because now a human is doing cleanup. [pause] JOSH: Where does this land for people who feel the model market changes every five minutes? ERIK: That feeling is correct. Vendors are sprinting. Your job is not loyalty. Your job is optionality. Build the lane change before you need the lane change. [pause] JOSH: Anthropic next. There were quality reports around Claude Code, the Agent SDK, and Cowork. For people who felt that but couldn’t prove it, what happened? ERIK: The useful part is they acknowledged separate issues instead of one vague “we had an incident.” That matters. Different layers can fail at the same time. UI freezes, degraded code quality, weird behavior under high reasoning effort. They fixed the issues by April 20 and reset usage limits. That’s the operational story. [beat] JOSH: What’s the engineering story? ERIK: Agentic systems are sensitive to defaults. People underestimate that. Change the default reasoning effort on a coding agent and you change latency, cost, completion behavior, tool usage, even whether the human thinks the model is “smart” or “broken.” Sometimes the model is fine and the orchestration is the problem. Sometimes the orchestration is fine and the default turned every task into a philosophical essay. [beat] JOSH: Wait, really? ERIK: Really. I’ve watched agents go from useful to annoying because they started overthinking a file rename. Nobody asked for a dissertation on a test fixture. We asked for a patch. [beat] JOSH: How do you guard against that in your own setup? ERIK: By separating roles. PrimeBus routes work based on task type. One Claude instance is better at code generation. Another is better at review. Another handles triage. Bob, Bill, Homer. Three Claude instances on the mesh. Same family, different jobs. If one starts acting strange, I can isolate the lane instead of wondering why the whole lab feels haunted. [beat] JOSH: That sounds like overkill for most people. ERIK: It sounds like overkill until your single-agent setup quietly burns three hours and 400,000 tokens trying to solve the wrong problem. Then suddenly a little structure sounds very reasonable. [beat] JOSH: What do you mean by isolate the lane? ERIK: Example. PrimeBus has 214 auto-fixes so far with a 78 percent success rate. Those are not random miracles. The system sees a failure, classifies it, generates A/B fix variants, runs tests, and only merges the winner. If code quality drops, I can inspect where. Was the diagnosis bad? Was the patch bad? Did the reviewer get too lenient? Did high-effort reasoning make it slower and not better? You want observability around the agent, not faith. [beat] JOSH: So when users said Claude Code felt off, you don’t treat that like whining. ERIK: No. I treat it like a production report. Users notice degradation before dashboards do. Not always accurately, but fast. If ten builders say “it got weird,” pay attention. They live in the tool all day. They know the feel of it. [beat] JOSH: What should teams change after seeing this? ERIK: Three things. First, pin important settings. Don’t let a vendor default decide your whole week. Second, log request class, model, effort level, latency, and outcome so you can compare before and after a provider change. Third, keep a fallback model path for critical work. If your coding agent goes sideways during a release, “we’ll wait for the vendor to sort it out” is not a plan. [beat] JOSH: That fallback path can be another provider? ERIK: Yep. Or another mode. Or even another prompt strategy with the same provider. I don’t care what religion somebody has around models. What matters is whether work keeps moving. [beat] JOSH: Any good example from your systems? ERIK: ScanBrief is simple on purpose. Gather. Dedup. Score. Summarize. Publish. If one summarizer degrades, I can still ship because the upstream pipeline is clean and the downstream formatting is rigid. Same with InkEngine. Writing a 55,000-word book in two days sounds flashy. What actually matters is checkpointing, chapter segmentation, review passes, and keeping the model from drifting off into fake confidence. Structure beats hype. [pause] JOSH: Last deep dive. Bitwarden CLI in a supply chain campaign. This one feels ugly. ERIK: It is ugly. Supply chain attacks are ugly because they exploit trust, and trust is the whole point of the toolchain. People install a CLI like Bitwarden because it’s supposed to make secrets handling safer. Then a compromised dependency path or poisoned image turns that trust into exposure. [beat] JOSH: Break that down for the audience that doesn’t live in CI pipelines. ERIK: Your build process pulls tools, containers, packages, extensions, scanners. You assume they’re clean because they’re common, or security-branded, or came from a source you used last month. An attacker only needs one path into that chain. Once they’re in, they don’t need to smash the front door. They ride your automation. Quietly. [beat] JOSH: Why does this keep happening? ERIK: Because modern software factories are giant graphs of borrowed code and borrowed infrastructure. Fast teams automate everything, which is good. Fast teams also trust too much by default, which is bad. Same knife. Different direction. [beat] JOSH: What’s the first fix? ERIK: Reduce blind trust. Pin versions. Verify signatures where possible. Mirror critical dependencies. Restrict outbound network access in CI. Don’t let every build container talk to the whole internet like it’s on vacation. [beat] JOSH: That sounds expensive. ERIK: Less expensive than incident response. People will spend months building clever agents and zero days hardening the pipeline those agents depend on. That’s backwards. If your code assistant helps you write faster but your build chain is porous, you’ve automated the path to compromise. [beat] JOSH: You’ve got dry words for that, I assume. ERIK: Fancy self-own. [beat] JOSH: That’s wild. ERIK: It is. And it gets worse with AI wrappers because people install random extensions, random model CLIs, random helper scripts, random Docker images from accounts they discovered 14 minutes ago. Then they pipe secrets through them. That’s not experimentation. That’s volunteering. [beat] JOSH: What should a small shop actually do today? ERIK: Inventory the tools in the path from commit to deploy. Not the ones you remember. The actual ones. Every action, every image, every package, every extension. Then rank them by blast radius. Secrets tools, build runners, deployment bots, artifact publishers. Those get the first hardening pass. [beat] JOSH: Any examples from your lab? ERIK: I run 229 cron jobs. That means I’ve got 229 chances to do something dumb on schedule. So I keep secrets scoped tight, isolate jobs, and use guardrails around anything that can mutate state. PrimeTrader watches 109 tickers in three batches every market hour. If that pipeline breaks, I lose alerts. Annoying. If a secrets pipeline breaks, that’s a different category. You separate “inconvenient” from “catastrophic” fast. [beat] JOSH: How does HumanRail fit into this? ERIK: HumanRail exists for exactly this kind of boundary. If confidence is low, or the action is risky, route to a human. People keep trying to remove humans from every step. Wrong target. Remove humans from the boring steps. Keep humans at the high-blast-radius junctions until you’ve earned enough confidence to automate more. [beat] JOSH: So the theme across all three stories is not “AI is magic.” ERIK: Correct. The theme is interfaces matter, defaults matter, and trust boundaries matter. That’s the work. Fancy demos get headlines. Adapters, telemetry, and guardrails keep the business alive. [pause] ERIK: This episode is sponsored by Prime Automation Solutions. If you're still doing it manually, we automate it. Also, special on a website — $250. primeautomationsolutions.com [pause] JOSH: Alright, what's the AI pro tip today? ERIK: Split thinking from doing. Don’t ask one model call to research, reason, code, review, and summarize in the same breath. That’s how you get mush. Make stage one produce a plan or a diagnosis. Make stage two execute one narrow task. Make stage three review against a checklist. [beat] ERIK: Same model or different models, doesn’t matter. What matters is the contract between stages. In my systems I’ll have one pass classify the failure, another generate two fixes, and a final pass judge the diff and test output. That’s a big reason PrimeBus got to 214 auto-fixes with a 78 percent success rate. The agent isn’t “smarter.” The workflow is cleaner. [beat] ERIK: If you’re building today, start small. Take one messy prompt and break it into three messages with explicit outputs. JSON if you can. Then log which stage failed. Once you can see the failure point, you can actually improve it. That’s your tip. Use it. [pause] JOSH: Binge all five episodes this weekend plus our YouTube shorts — links at buildorbereplaced.dev. [pause] JOSH: One more thing — we started a Discord for builders. If you're shipping AI, automation, or anything that makes a human obsolete — come hang out. Link at buildorbereplaced.dev. ERIK: Post what you built. We'll post what we're building. Real wins, real builds, no fluff. [pause] ERIK: Build or be replaced. JOSH: If you want these signals in your inbox every morning, scanbrief.dev. See you tomorrow.