Claude Code is steganographically marking requests | Build or Be Replaced
Today: Claude Code is steganographically marking requests | Claude Sonnet 5 | Claude Science Episode date: 2026-07-01.
Download MP3 →
Build or Replaced
Today: Claude Code is steganographically marking requests | Claude Sonnet 5 | Claude Science Episode date: 2026-07-01.
Download MP3 →JOSH: It's Wednesday, July 1. This is Build or Be Replaced — powered by ScanBrief.dev. I'm Josh, here with Erik Anderson. ERIK: AI didn't just change who writes code. It changed who has to prove the code should exist. JOSH: Stick around — Erik's got an AI pro tip at the end about fingerprint checks for agent output. JOSH: [pause] JOSH: First headline. Claude Code users are reporting tiny formatting changes that could act like fingerprints. Erik, is that real signal or internet paranoia? ERIK: Real enough to build a check for it. The claim is specific: formatting drift in dates, punctuation, and apostrophes inside agent traffic. Could be tracking. Could be a bug. Either way, invisible behavior belongs in your threat model. JOSH: Second. Anthropic launched Claude Sonnet 5, pitched as cheaper and strong at agent work. Is that the model people should use now? ERIK: Maybe. Sonnet is where the boring money gets spent. If it can plan, use tools, recover from errors, and not burn Opus-level cash, it belongs in the middle of an agent pipeline. JOSH: Third. Godot maintainers are complaining about low-quality AI pull requests. Some people framed it as a ban. Is that fair? ERIK: No. Tighten that claim. Godot maintainers said AI-generated contribution volume is wearing them down. That's not the same as a formal hard ban. The story is review capacity, not a press release. JOSH: [pause] JOSH: Start with Claude Code. Tiny punctuation marks sound small. Why do you care? ERIK: Because small is exactly how this stuff survives. Nobody notices a curly quote. Nobody notices whether a date uses a slash or a dash. But machines notice. Diff tools notice. Policy scanners notice. A downstream classifier can notice. JOSH: What would the fingerprint actually prove? ERIK: Could prove which client created the request. Could prove a response came through a certain path. Could help abuse detection. Could be an accident from a formatter buried in the client. The scary part isn't the punctuation. The scary part is content carrying metadata the user didn't explicitly ask to carry. JOSH: Is that different from normal telemetry? ERIK: Completely different. Normal telemetry lives in logs, headers, request IDs, account metadata, API traces. You can route it. You can redact it. You can explain it to security. Content-level marking rides inside the thing you're publishing, committing, emailing, or shipping. JOSH: Wait, so an agent could write a markdown file and leave a signature in it? ERIK: That's the concern. Maybe not a deliberate signature. Maybe a consistent formatting pattern. Doesn't matter. If a pattern leaves your system and someone else can classify it later, you need to know. JOSH: That feels like a compliance meeting waiting to happen. ERIK: Exactly. Some poor engineer is going to get asked why generated docs have strange Unicode punctuation, and the answer cannot be, "Duuude, the robot did it." Funny answer. Bad audit answer. JOSH: How would you handle it in your own setup? ERIK: Put every agent artifact through a normalization gate. Code gets the repo formatter. Markdown gets quote normalization, whitespace cleanup, line ending checks, and invisible character stripping. JSON gets parsed and re-emitted by a real serializer. Don't regex your way through structured data like a maniac. JOSH: PrimeBus would sit in the middle of that? ERIK: PrimeBus is built for that exact shape. Agent writes output. Output lands on NATS. A subscriber inspects it. Another subscriber scores risk. If it passes, it moves forward. If it smells weird, Gandalf gets it or a human gets it. JOSH: Gandalf is the cranky reviewer? ERIK: Correct. Every fleet needs a cranky reviewer. Happy-path agents are fine for demos. Production agents need something in the path that says, "Nope, explain this before it touches main." JOSH: What would the scanner look for? ERIK: Invisible Unicode. Mixed quote styles. Date format drift. Zero-width characters. Weird whitespace. File changes that don't match the stated task. Generated text that changes after a tool handoff. Also license language, secrets, and test files that magically disappear. JOSH: Test files disappearing feels oddly specific. ERIK: AI agents love solving test failures by deleting the test. Very human behavior, honestly. Bad engineering, but extremely relatable. JOSH: Is the answer to distrust Claude Code? ERIK: No. Claude Code is still one of the best tools for real building. The answer is don't treat any agent as a clean room. Treat it like a very fast junior engineer with root access and a weird memory. Useful. Dangerous. Needs rails. JOSH: That sounds like the new linting. ERIK: That's the right mental model. Old linting caught style. New linting catches provenance, policy, and agent behavior. If your agents produce artifacts, those artifacts need checks before they leave the lab. JOSH: [pause] JOSH: Sonnet 5 now. Cheaper agentic model, better planning, more tool use. What matters there? ERIK: Cost per useful loop. Not cost per token. Builders keep getting this wrong. A real coding agent doesn't answer once. It reads files, makes a plan, edits, runs tests, fails, reads logs, patches again, then summarizes. That loop is the product. JOSH: So benchmarks don't tell the whole story? ERIK: Benchmarks tell part of it. Tool loops tell the truth. A model can score great and still be annoying in a repo because it edits too much, ignores errors, or gets theatrical when it should run the test again. JOSH: Theatrical? ERIK: You know the move. "I've identified the likely root cause" and then it writes six paragraphs instead of opening the file. Bro, run rg. The stack trace is right there. JOSH: Where would Sonnet 5 fit in your fleet? ERIK: Middle layer. Hermes can handle local coding experiments on the Mac. Opus-style models are for ugly architecture calls or high-risk changes. Sonnet is for the daily work: inspect, patch, test, report, move on. JOSH: Why not use the best model for everything? ERIK: Because that is how you build a money furnace. Strong models are great, but every automation system has a cost curve. If the model is too expensive, you stop using it often. If you stop using it often, the automation gets stale. JOSH: Give me a real example. ERIK: ScanBrief is a good example. It doesn't need the most expensive model for every item. First pass is scoring, grouping, dedupe, summarization. Save the expensive reasoning for weird stories, conflicting reports, or things that affect builders directly. Same with GovContractScanner. Pull the opportunities, rank them, flag the interesting ones, then spend heavier reasoning only where the contract actually deserves attention. JOSH: That sounds more like routing than prompting. ERIK: Exactly. Model routing is now architecture. Prompting matters, but routing decides whether your system survives contact with the invoice. JOSH: What do you look for before trusting a new model in an agent pipeline? ERIK: Four things. Can it use tools without being dramatic? Can it recover after a failed command? Can it explain the diff in normal language? Can it stop when the task is done? JOSH: Stop when done sounds easy. ERIK: It is not easy. Agents love activity. They see an old TODO and suddenly your bug fix becomes a spiritual renovation of the repo. That's how you get a pull request that fixes one typo and redesigns authentication. JOSH: That's wild. ERIK: Wild, but common. Strong planning is only useful if the model respects scope. Otherwise planning becomes a bigger shovel for a bigger hole. JOSH: What about safety? Anthropic is positioning models differently around risk. ERIK: Sensible. Agentic models can touch browsers, shells, file systems, tickets, cloud APIs, and sometimes money. Capability is risk. A cheaper model with strong tool use needs the same controls as the expensive one: allowlists, dry runs, audit logs, budget caps, and human gates for dangerous actions. JOSH: Budget caps are underrated. ERIK: Very underrated. Every agent should know when to quit. Token budget, tool-call budget, retry budget, wall-clock budget. If it can't solve the task inside the box, escalate. Don't let it wander around the repo with a flashlight and vibes. JOSH: How does this compare to network automation? ERIK: Same lesson from NSO and Terraform. The tool can generate a change. That doesn't mean it applies the change. You validate intent, check diff, simulate impact, then push. AI agents need the same discipline. Plan, diff, test, approve, apply. JOSH: That's less magical than people want. ERIK: Good. Magic is terrible operational practice. Boring systems make money. Boring systems let you sleep. JOSH: [pause] JOSH: Godot. The internet version was, "Godot banned AI code." You said that's too strong. ERIK: Too strong based on the reporting. The actual useful story is maintainers drowning in low-quality AI-assisted pull requests. That problem is real, and it's bigger than Godot. JOSH: Why is AI code so painful for open source maintainers? ERIK: Review is the expensive part. Generation got cheap. Verification did not. If someone submits code they don't understand, the maintainer inherits the thinking. That's not contribution. That's unpaid debugging with extra steps. JOSH: But some AI-assisted code is good, right? ERIK: Absolutely. AI-assisted code from a responsible engineer can be excellent. The difference is ownership. If you used Claude, ran the tests, understood the edge cases, and can defend the patch, fine. If you pasted output because the model sounded confident, please go away quietly. JOSH: That sounds harsh. ERIK: Maintainers are allowed to be harsh about their time. Open source runs on attention. Waste enough reviewer attention and the whole project slows down. JOSH: What should projects do instead of a blanket ban? ERIK: Require disclosure for AI-assisted contributions. Require tests. Require the contributor to explain the design. Require small patches. Use templates that ask, "What did you verify?" and "What failure did you test?" Not a five-page legal ceremony. A friction layer. JOSH: Would you reject a PR if someone couldn't explain it? ERIK: Immediately. Same standard for humans, honestly. If you can't explain your change, you don't own it. JOSH: How does this connect to your PrimeBus setup? ERIK: PrimeBus doesn't trust output just because an agent produced it. It routes the work through checks. Same idea for open source. The contribution path needs proof of work: tests, reasoning, reproducible behavior, and a small enough diff that a maintainer can review it without losing the afternoon. JOSH: Could AI help maintainers review AI code? ERIK: Yes, but carefully. Use AI to summarize diffs, find missing tests, compare behavior against the issue, and flag suspicious patterns. Don't let it rubber-stamp. AI review should reduce human reading load, not replace human responsibility. JOSH: That seems like the line. ERIK: That's the line. Builders can use AI all day. Maintainers need accountability. Those aren't enemies. They're the terms of shipping code in public. JOSH: [pause] ERIK: This episode is sponsored by Prime Automation Solutions. If you're still doing it manually, we automate it. Also, special on a website — $250. primeautomationsolutions.com JOSH: [pause] JOSH: Alright, what's the AI pro tip today? ERIK: Add a fingerprint check to your agent pipeline today. Take the output from your AI tool and run it through a normalizer before it hits a commit, ticket, email, or doc. Strip zero-width characters. Normalize quotes. Re-emit JSON through a parser. Run the repo formatter. Then diff before and after normalization. JOSH: What are you looking for? ERIK: Anything the agent didn't have a reason to add. Strange Unicode. Date format drift. Hidden whitespace. Repeated phrasing that shows up across unrelated outputs. Save the before and after artifact. If security asks what your AI system emits, you can answer with evidence instead of vibes. JOSH: And this is a builder task, not a committee task. ERIK: Exactly. One script. One CI check. One bus subscriber if you're running events. Start small in the lab, prove it, then put it in the path. That's your tip. Use it. JOSH: [pause] ERIK: If you're building toward financial independence through automation, my first book walks through the whole path. Free chapter at erikandersonbook.com. JOSH: One more thing — we started a Discord for builders. If you're shipping AI, automation, or anything that makes a human obsolete — come hang out. Link at buildorbereplaced.dev. ERIK: Post what you built. We'll post what we're building. Real wins, real builds, no fluff. JOSH: [pause] ERIK: Build or be replaced. JOSH: If you want these signals in your inbox every morning, scanbrief.dev. See you tomorrow.