Mon–Fri · 6 AM ET
← All Episodes
EP  • 00:08:16

Does Reddit have an astroturfing problem? What the data suggests | Build or Be Replaced

Today: Does Reddit have an astroturfing problem? What the data suggests | Sonnet 5.5 | It's Time to Investigate the AI Labs Episode date: 2026-09-29.

Download MP3 →

Transcript

ERIK: It's Tuesday, September 29th. This is Build or Be Replaced — powered by ScanBrief.dev. I'm Erik Anderson, here with Josh.
JOSH: It's Tuesday, September 29th. This is Build or Be Replaced — powered by ScanBrief.dev. I'm Josh, here with Erik Anderson.
ERIK: Today's theme is who's watching the machines. Turns out even Nvidia thinks your AI agent needs a babysitter.
JOSH: Stick around — Erik's got an AI pro tip at the end about picking the right size model for the job.
[pause]
JOSH: Let's start with headlines. First one — a Hacker News thread is asking if Reddit has an astroturfing problem, and the data kind of says yes.
ERIK: Not shocking. Anytime a platform gets big enough to move opinion, someone's going to try to buy the opinion. The interesting part isn't that it happens — it's that people finally have the tooling to detect it at scale.
[beat]
JOSH: Next — somebody hijacked a PS5's RTMP stream.
ERIK: Sony's console is basically a Linux box wearing a costume. Anytime you expose a streaming pipeline, you've opened a door. I've got RTMP running on three of my own boxes — you patch that stuff or you get to be the headline.
[beat]
JOSH: California wine grape growers are struggling — demand's down over twenty percent in five years.
ERIK: That one's just a supply chain lagging a taste shift. Same thing happens in tech. You don't feel the demand curve turn until you're the one holding the inventory.
[beat]
JOSH: Last one — a new survey says more than half of professional developers now ship AI-generated code without a full manual review.
ERIK: That number doesn't scare me, it just tells me most of those teams don't have a Gandalf. If a human isn't reviewing every line, something else has to be. Otherwise you're just hoping.
[pause]
JOSH: Alright, deep dive time. The big one — Sonnet 5.5. Anthropic dropped a system card for it. Why does that matter to you specifically?
ERIK: A system card means the release is close, and it means I need to plan a migration test before it lands, not after. I've got three Claude instances running the Bobaverse — Neo, Bill, Homer — plus Gandalf doing review. When a new Sonnet drops, I don't just flip the switch. I run it against my eval suite first.
JOSH: What's actually in that eval suite?
ERIK: Real prompts from real jobs. Code review outputs, fix generation, routing decisions. If the new model doesn't beat the old one on my own workloads, it doesn't get promoted. PrimeBus has run 2,295 auto-merge attempts since June — 1,507 merged, 788 blocked by Gandalf, zero escalated to me. That 788 number is the guardrail working, not the system failing. A model swap that quietly drops that merge rate is a regression I'd never see unless I measured it.
[beat]
JOSH: So you don't just trust the vendor's benchmark.
ERIK: Never. Their benchmark is marketing. My benchmark is 154 services on my production server that have to keep running tonight.
JOSH: How long does a migration test like that actually take you?
ERIK: A few hours if nothing's weird. I stand the new model up in a shadow lane first — it sees the same traffic, makes the same calls, but nothing it does actually ships. I compare its decisions against what the current model did. Only once the numbers line up do I even think about promoting it into the real chain.
JOSH: And if it doesn't line up?
ERIK: Then it sits in the shadow lane until it does, or it doesn't ship at all. There's no partial credit in production.
[pause]
JOSH: Second story — this one's smaller but caught your eye. "Jeff" — Jev-compatible 0.8B decision models, trained at home, running in about 30 milliseconds.
ERIK: This is the one I actually care about most today. Everybody's chasing bigger models. This guy trained a sub-billion-parameter model at home that makes a decision in 30 milliseconds. That's not a chatbot — that's a router.
JOSH: What's the difference?
ERIK: A chatbot has a conversation. A decision model looks at an input and picks a lane — yes or no, this backend or that one, escalate or don't. I don't need frontier-class reasoning to decide if a lead is worth a callback. I need something fast and cheap that's right the vast majority of the time. That's exactly the shape of thing I'd point at Real Estate Automation — scoring inbound leads doesn't need a big model, it needs a decision in 30 milliseconds so the callback goes out before the lead goes cold.
JOSH: Would you actually run something like that in prod?
ERIK: I'd test it in the lab first, same as always. But the pattern's right. Tiny model, narrow job, ridiculous speed. I've got dozens of projects' worth of telemetry hitting PrimeBus every day — most of those decisions don't need a big brain, they need a fast one.
JOSH: Is that basically the same idea as running a small model locally versus calling out to the cloud?
ERIK: Same family of idea, yeah. Local and small means no network round trip, no per-call bill, and no dependency on somebody else's uptime. When the decision is narrow enough, that trade is an easy yes.
[pause]
JOSH: Last deep dive — Nvidia wants to put a watchdog chip next to every AI agent.
ERIK: This is the headline I've basically already built. A watchdog that sits outside the agent and checks what it's doing before it acts — that's Gandalf. That's my whole review layer. Nvidia's just proposing it in silicon instead of software.
JOSH: Why would you want it in hardware instead of just code, like you're doing?
ERIK: Speed and trust boundary. If the watchdog runs in the same process as the agent, a bad agent can potentially talk its way around it or just crash it. Put it on separate silicon and the agent literally cannot touch it. I don't have that luxury at my scale, so I do it with process isolation instead — the reviewer reads the diff fresh, it doesn't share memory with whatever generated the fix.
JOSH: Does that actually stop bad merges?
ERIK: That's what the 788 blocked number is. Almost eight hundred changes that looked fine to the model that wrote them and got stopped by the model reviewing them. That's the entire argument for a watchdog, proven out in my own pipeline every single day. Nvidia's just realizing everybody running agents at scale is going to need one. I'd have told them that a year ago.
[beat]
JOSH: So the theme really is nobody trusts the agent alone anymore.
ERIK: Nobody should. Mine included. I don't trust my own fleet unsupervised — that's not pessimism, that's twelve agents and a hundred and fifty-four services. You supervise or you get surprised.
JOSH: Is there a version of this where the watchdog itself gets it wrong?
ERIK: Sure, and that happens. Gandalf oscillates sometimes on a borderline diff, flags something twice, changes its mind. That's why I still read the actual code path before I trust a blocked verdict. The watchdog isn't infallible, it's just a second set of eyes that never gets tired.
[pause]
ERIK: This episode is sponsored by Prime Automation Solutions. If you're still doing it manually, we automate it. Also, special on a website — $250. primeautomationsolutions.com
[pause]
JOSH: Alright, what's the AI pro tip today?
ERIK: Stop defaulting to the biggest model for every task. Before you write a prompt, ask what the job actually needs — a decision, a summary, a full reasoning chain. If it's a decision, route it to something small and fast, like that 0.8B model we just talked about. If it's a summary, a mid-size model is plenty. If it's real reasoning — architecture, debugging, anything with a lot of ambiguity — send it to your frontier model. I run that exact split across my own backends. Every request gets classified first, then routed to the smallest model that can actually do the job right. That saves money and it's faster for the caller, because a 30-millisecond decision beats a multi-second reasoning call every time the job doesn't need reasoning.
JOSH: So bigger isn't automatically better.
ERIK: Bigger is a tool, not a default. Match the tool to the job. That's your tip. Use it.
JOSH: We also drop daily market picks and automation tips on YouTube — search Build or Be Replaced.
JOSH: One more thing — we started a Discord for builders. If you're shipping AI, automation, or anything that makes a human obsolete — come hang out. Link at buildorbereplaced.dev.
ERIK: Post what you built. We'll post what we're building. Real wins, real builds, no fluff.
ERIK: Build or be replaced.
JOSH: If you want these signals in your inbox every morning, scanbrief.dev. See you tomorrow.