Transcript
JOSH: It's Thursday, May 21. This is Build or Be Replaced — powered by ScanBrief.dev. I'm Josh, here with Erik Anderson.
ERIK: The developer toolchain is now an attack surface with a nice icon and five-star reviews.
JOSH: Stick around — Erik's got an AI pro tip at the end about making agents prove their work before they touch production.
[pause]
JOSH: Headlines. First up, GitHub confirmed a breach tied to a malicious VS Code extension. Erik, that's every developer's favorite nightmare, right?
[beat]
ERIK: Yep. Trusted tool, trusted editor, bad payload. GitHub says 3,800 repositories were exposed, and the lesson is simple: your IDE is part of your production environment now.
JOSH: Second, Anthropic is reportedly expanding into Colossus2 and using GB200 chips. What's the signal there?
ERIK: Compute is the product. Everyone talks about model intelligence, but the companies with chips, power, cooling, and scheduling win the next round.
JOSH: Third, an OpenAI model reportedly disproved a central conjecture in discrete geometry.
ERIK: That's not a chatbot writing a poem. That's machine reasoning pushing into math. If that holds up, the AI conversation gets a lot less cute and a lot more serious.
[pause]
JOSH: Let's start with the GitHub breach. A malicious VS Code extension gets installed, thousands of repos get exposed. How bad is that?
ERIK: Bad in the way that matters. Not dramatic movie bad. Boring infrastructure bad. Somebody installs an extension because it looks useful. Maybe it formats code. Maybe it helps with Git. Maybe it has a logo that looks like twelve other tools. Then it gets access to tokens, workspace files, terminal sessions, environment variables, maybe secrets in plain text because people still do that. Now your private repos aren't private anymore.
[beat]
ERIK: That's the part people miss. A repo breach isn't just source code. It's architecture. It's comments. It's test fixtures. It's CI files. It's deployment patterns. It's old credentials somebody forgot in a sample config. It's the map of the building.
JOSH: Wait, really? Source code alone is that useful?
ERIK: Source code is the least interesting part sometimes. If I can see your Terraform, your Kubernetes manifests, your GitHub Actions, your service names, your S3 bucket naming, your internal package names, I'm building a plan. I know where to poke. I know what assumptions your team makes. And if your developers have extensions running with broad access, that's a really nice doorway.
JOSH: How do you defend against that without making developers miserable?
ERIK: You treat extensions like dependencies. Approved list. Version pinning where you can. Permission review. Separate workspaces for sensitive repos. No random extension gets to walk into the bank vault because it has dark mode.
[beat]
ERIK: And you watch behavior. PrimeBus processed 2231 automation events across 10 projects today. That matters because events tell you when something changes. New process. New outbound call. New repo clone. New secret access. New extension install. If you don't have telemetry, you're asking humans to remember what normal feels like. Humans are bad at that. We have meetings.
JOSH: So this isn't just a GitHub problem.
ERIK: Nope. GitHub is the headline because it has the brand name. The real issue is developer supply chain. npm packages, browser extensions, VS Code plugins, GitHub Actions, Docker images, AI coding tools, MCP servers, all of it. Developers are now running little automation agents all day. Some are useful. Some are a tire fire with a README.
JOSH: Where do AI coding tools fit into that?
ERIK: Same bucket. If Claude, Cursor, Copilot, Windsurf, or any agent can read your repo and run commands, it needs boundaries. Not vibes. Boundaries. Read access is one thing. Write access is another. Shell access is another. Production deploy access is a completely different animal.
[beat]
ERIK: In my setup, Gandalf reviewed and merged 8 code changes to production overnight. That's not because I trust automation blindly. It's because the pipeline has checks, tests, diff review, and blocking rules. The PrimeBus auto-merger has run 1800 attempts since 2026-04-17: 714 merged, 1074 blocked by Gandalf, 12 escalated to Erik. That blocked number is the story. Guardrails did their job.
JOSH: That's the opposite of just letting an agent rip through a repo.
ERIK: Exactly. Agents should be powerful, but they should be treated like junior engineers with root access and unlimited coffee. Useful, fast, occasionally terrifying. You make them show diffs. You make them run tests. You make them explain risk. You don't give them a production token because they asked nicely.
[pause]
JOSH: The second deep story is Anthropic expanding to Colossus2 and using GB200. Everyone hears chip names and data center names and tunes out. Why should builders care?
ERIK: Because model access is going to feel like bandwidth. Some days fast. Some days expensive. Some days rate-limited. Some days the model you want is not the model you get. The companies building massive AI infrastructure are trying to make that less painful for themselves first, and then maybe for us.
JOSH: GB200 is Nvidia's big AI platform, right?
ERIK: Right. Grace Blackwell. CPU and GPU tied together for AI workloads. High memory bandwidth, fast interconnect, built for giant model training and inference. The details matter to hardware people. For builders, the useful translation is this: bigger clusters mean bigger models, longer context, faster inference, and more competition around price.
JOSH: Does that mean everything just gets cheaper?
ERIK: Not automatically. Compute never gets cheap in a straight line. Demand eats every improvement. Faster chips show up, then model labs train bigger models. More efficient inference shows up, then people run agents all day instead of one prompt. You give developers more tokens, they fill the bucket by lunch.
[beat]
ERIK: PrimeRouter exists because I don't want one provider deciding whether my systems work today. It routes LLM calls by priority tier, tracks fleet telemetry, and fails over across providers. That's not fancy. That's plumbing. But plumbing is what keeps the house from becoming a swamp.
JOSH: So your take is, don't bet your automation on one model?
ERIK: Correct. Don't bet the business on one model, one endpoint, one context window, one vendor account, or one pricing page. Use Claude where it wins. Use GPT where it wins. Use smaller models when the task is boring. Use local or cheap models for classification if the quality is good enough. Builders need routing, evals, and budgets. Not brand loyalty.
JOSH: How does token speed fit into that? ScanBrief had a story today about how fast tokens per second actually feels.
ERIK: That story is underrated. People say 500 tokens per second like it's a clean number. But tokens aren't words. Code is dense. JSON is dense. Prose feels different. A model generating 500 tokens per second on a short summary feels instant. The same rate on a code patch with tests and logs can still feel like you're watching paint apply for a permit.
JOSH: That's painfully specific.
ERIK: Because it's real. In agent systems, latency is not just one model response. It's plan, read files, call tools, edit, run tests, parse errors, retry, summarize, wait on CI. If each step is small, token speed matters less than orchestration. If each step is huge, token speed matters a lot.
[beat]
ERIK: That's why I care about event buses. NATS, PrimeBus, queues, workers, telemetry. You don't want one giant agent doing everything in one giant context. You want small jobs with receipts. ScanBrief scored 51 items across 54 sources today. That works because the pipeline breaks the job apart: fetch, dedupe, score, summarize, rank. Each piece is measurable. If one piece gets weird, I can see it.
JOSH: That's the pattern listeners should steal?
ERIK: Yes. Break the work into events. Give every event an ID. Log inputs and outputs. Make retries visible. Keep human escalation as a first-class path. That's how you build automation that survives bad days.
JOSH: And the AI labs are racing to give us more compute for that.
ERIK: They are. But compute doesn't fix sloppy architecture. Bigger models make bad systems fail with better grammar.
[pause]
JOSH: Third deep story. An OpenAI model disproves a central conjecture in discrete geometry. This one sounds academic, but you said it changes the AI conversation.
ERIK: It does if the result holds up. A model helping disprove a conjecture means we're not just using AI for summarizing emails and writing unit tests. We're using it to search a hard problem space and find something humans missed.
JOSH: Is that reasoning, or is it just brute force?
ERIK: Sometimes brute force is reasoning with better shoes. Math discovery has always used computation. The interesting part is whether the model can guide the search, form useful structures, reject dead ends, and produce something mathematicians can verify. It doesn't need to be mystical. It needs to be productive.
JOSH: What does that mean for engineers?
ERIK: It means the ceiling is moving. Today it's geometry. Tomorrow it's protocol design, compiler passes, network failure models, test generation, formal verification, routing policy analysis. The stuff that used to require a very expensive human staring at a whiteboard for two weeks can start with an AI generating candidate proofs, counterexamples, or simulations.
[beat]
ERIK: In networking, this matters. I've spent years around Cisco NSO, service models, rollback behavior, config drift, device templates, all the fun stuff that makes your eyes do pushups. AI doesn't have to replace the senior engineer in one shot. It can generate the test cases the senior engineer would've written after three cups of coffee and mild regret.
JOSH: That's where the "build or be replaced" part gets uncomfortable.
ERIK: It should be uncomfortable. If your job is only to type what someone else already decided, you're in trouble. If your job is to define the problem, design the system, set the guardrails, verify the output, and ship, AI makes you harder to replace.
JOSH: So the builder's advantage is judgment?
ERIK: Judgment plus systems. People say "learn AI" and then they mean prompts. Prompts are fine. But the real advantage is turning AI into a worker inside a controlled system. Inputs, tools, permissions, tests, rollback, alerts. The model is not the product. The system around it is the product.
JOSH: That sounds like your home lab talking.
ERIK: It is. The Bobaverse fleet has 5 agents: Neo, Homer, Bill, Echo, Gandalf. Claude plus GPT. That's not because I enjoy naming machines, although I clearly do. It's because different agents do different jobs. One reviews. One builds. One watches telemetry. One handles routing. Gandalf blocks bad merges. Very on brand.
[beat]
JOSH: What happens when the model is wrong?
ERIK: Same thing that happens when a human is wrong. The system catches it or you have a bad day. That's why I don't care about demos where an agent writes a todo app. Cool. Can it handle secrets? Can it roll back? Can it explain why it touched a file? Can it stop when confidence is low? Can it ask a human instead of pretending?
JOSH: That's the HumanRail idea, right? Human when the model isn't confident?
ERIK: Exactly. Humans shouldn't be in every loop. That's how you build a slow robot. Humans should be in the right loop. Weird permission change. Failing test with ambiguous cause. Security-sensitive diff. Production database migration. The boring stuff can pass. The sharp stuff gets a checkpoint.
JOSH: Does the math story make you more aggressive about using AI?
ERIK: More serious, not reckless. If models can contribute to math, they can absolutely contribute to infrastructure. But infrastructure has blast radius. A wrong proof annoys mathematicians. A wrong automation deletes a database. PrimeRestorer exists because I like sleeping. Backups, system state, S3 offsite. Not glamorous. Very useful when somebody's clever agent gets clever in the wrong directory.
[pause]
JOSH: Quick thread tying these together. Malicious extensions, giant AI compute, models doing math. What's the builder's take?
ERIK: The builder's take is simple. Tools are getting stronger on both sides. Attackers get better automation. Builders get better automation. Model labs get more compute. Your editor gets more powerful. Your agents get more access. The gap is governance. Not corporate theater governance. Real engineering governance. What can run? What can read? What can write? What gets logged? What gets blocked? Who gets paged?
[beat]
ERIK: If you're a solo builder, do the small version. Separate tokens. Use least privilege. Keep secrets out of repos. Review extensions. Run agents in a sandbox. Require tests before merge. Log every automation action. If you're a company, do the same thing with fewer excuses and more money.
JOSH: And if someone says that's too much overhead?
ERIK: Then they haven't been breached yet.
[pause]
ERIK: This episode is sponsored by Prime Automation Solutions. If you're still doing it manually, we automate it. Also, special on a website — $250. primeautomationsolutions.com
[pause]
JOSH: Alright, what's the AI pro tip today?
ERIK: Make your agents produce evidence, not just output. Every agent task should end with three things: what changed, how it verified the change, and what it refused to do. If it edited code, make it show the files and test results. If it read logs, make it cite the exact error. If it skipped something, make it say why.
[beat]
ERIK: This is how you stop fake confidence from entering your system. Add a final instruction: "Before finishing, list the commands you ran, the checks that passed, the checks that failed, and any risky assumptions." Then wire that into your review flow. Don't accept "done." Done is a claim. Evidence is the receipt.
[beat]
ERIK: That's your tip. Use it.
[pause]
JOSH: Track your freedom score and net worth with the Freedom Blueprint app — free download, link in the show notes.
[pause]
JOSH: One more thing — we started a Discord for builders. If you're shipping AI, automation, or anything that makes a human obsolete — come hang out. Link at buildorbereplaced.dev.
ERIK: Post what you built. We'll post what we're building. Real wins, real builds, no fluff.
[pause]
ERIK: Build or be replaced.
JOSH: If you want these signals in your inbox every morning, scanbrief.dev. See you tomorrow.