Mon–Fri · 6 AM ET
← All Episodes
EP  • 00:12:58

1-Click GitHub Token Stealing via a VSCode Bug | Build or Be Replaced

Today: 1-Click GitHub Token Stealing via a VSCode Bug | Use your Nvidia GPU's VRAM as swap space on Linux | AI outperforms law professors in Stanford Law study Episode date: 2026-06-03.

Download MP3 →

Transcript

JOSH: It's Wednesday, June 3. This is Build or Be Replaced — powered by ScanBrief.dev. I'm Josh, here with Erik Anderson.
ERIK: Today's theme is simple: your tools are now part of your attack surface, your memory hierarchy, and your legal department. That's a weird Wednesday.
JOSH: Stick around — Erik's got an AI pro tip at the end about making agents prove their work before they touch production.
[pause]
JOSH: ScanBrief scored 80 items across 54 sources today. First headline: a VSCode bug that can steal GitHub tokens with one click. How bad is that?
ERIK: Bad in the boring way, which is the dangerous way. If github.dev hands a broad OAuth token to a browser-based IDE, and a bug lets a page grab it, that's not a popup problem. That's repo write access.
JOSH: Next one: people are using Nvidia VRAM as swap space on Linux. Clever hack or terrible idea?
ERIK: Both. VRAM is fast, but it isn't RAM, and once you start pretending memory tiers are all the same, Linux will teach you manners. For the right workload, though, it's a fun trick.
JOSH: Third headline: AI beat human answers in a Stanford Law study, at least in professor preference. Are lawyers cooked?
ERIK: Lawyers are not cooked. Bad legal writing is cooked. There's a difference, and a lot of people are about to learn it expensively.
[pause]
JOSH: Let's start with the VSCode token story. A lot of people hear “web IDE bug” and think, okay, patch the browser and move on. Is that the wrong read?
[beat]
ERIK: Yeah. The bug is only the door handle. The real story is token design.
ERIK: When you open github.dev or Codespaces-style tooling, the browser experience needs access to your repos. That's the whole point. But if the token is broad, and it can read and write across private repos, now one click can become supply chain access.
JOSH: One click meaning the user doesn't have to paste a token anywhere?
ERIK: Exactly. That's what makes it nasty. The user thinks they're opening a file or a link. The browser environment already has trust. The attacker doesn't need to socially engineer a PAT out of you. They ride the trust that's already there.
JOSH: That feels worse than a normal phishing link.
ERIK: It is worse for builders. A normal phishing link steals a login if you're sloppy. This steals the thing your tools use after login. And developers are swimming in those tokens. GitHub, npm, cloud providers, CI runners, package registries, LLM gateways. It's keys all the way down.
JOSH: What's the practical fix? Because “don't use web IDEs” doesn't sound realistic.
ERIK: Don't make any one token too useful. That's the fix.
ERIK: Repo-specific tokens. Short lifetimes. Read-only by default. Separate browser tokens from automation tokens. And for anything that can write, put a human or a policy engine between request and merge.
ERIK: PrimeBus is built around that idea. It processed 2392 automation events across 11 projects today, but the events don't equal permission. An agent can ask. Gandalf reviews. A policy decides. Then production gets touched.
JOSH: You had overnight merges too, right?
ERIK: 13 code changes were automatically reviewed by Gandalf and merged to production overnight. That's the part people get excited about. But the more important number is the block path.
ERIK: PrimeBus auto-merger has run 4193 attempts since 2026-04-17: 1196 merged, 2985 blocked by Gandalf, 12 escalated to Erik. The merge rate is 28.5%. The rest were caught by guardrails, not failures.
JOSH: Wait, really? So most attempts don't merge?
ERIK: Correct. That's the point. A good automation system says no constantly. People keep trying to measure agents by how much they do. Wrong metric. Measure how well they refuse to do dumb things.
JOSH: How would that apply to this GitHub token issue?
ERIK: Assume the token leaks. That's the mindset.
ERIK: If a browser token leaks, what can it write? If it writes, what catches the diff? If it opens a PR, what checks run? If it changes a workflow file, who reviews that? If it touches secrets, does the system scream?
ERIK: You don't secure this by trusting the IDE. You secure it by making every identity small and every write path annoying enough to survive a bad Wednesday.
JOSH: Annoying on purpose.
ERIK: Yep. Friction is not always the enemy. Sometimes friction is the adult in the room.
[pause]
JOSH: The second story is weirder. Nvidia VRAM as Linux swap. Help me understand why anyone would do this.
ERIK: Because VRAM is fast and nerds are undefeated.
ERIK: Normal swap uses disk. Even NVMe is much slower than main memory. GPU VRAM has huge bandwidth. So the idea is: if your machine is memory constrained, maybe use unused VRAM as a high-speed spill area instead of crawling on disk swap.
JOSH: That sounds useful for AI workloads.
ERIK: It can be. But you need to be precise about the workload. If you're running local models, image generation, build jobs, big browsers, maybe a database, and you're tight on RAM, this can keep the machine from falling over.
ERIK: It doesn't magically turn VRAM into normal system RAM. Access paths matter. Latency matters. Driver behavior matters. Page movement matters. The CPU still has to get to that memory through a different path.
JOSH: So it's not “install this and your laptop is a supercomputer.”
ERIK: No. Please don't put that on a thumbnail.
ERIK: It's more like adding a weird, fast, limited overflow bucket. Good if you understand the bucket. Bad if you forget it's weird.
JOSH: Where would you actually use it?
ERIK: Lab machines. Build boxes. AI workstations. Places where you can test failure modes without making customers sad.
ERIK: My home lab is two servers, but the principle is the same. Neo and Morpheus run a lot of local services, agents, cron jobs, and telemetry. If I add a memory tier, I want metrics before I trust it. Page in rate. Page out rate. GPU contention. Process latency. Driver resets. All of it.
JOSH: That sounds like a PrimeRouter problem too.
ERIK: Yep. PrimeRouter is the LLM gateway, so it's already thinking in tiers. Priority-tier routing, fleet OTEL, provider failover. If a request is high priority, it gets a better route. If a provider is degraded, it moves. If telemetry says the route is getting weird, it backs off.
ERIK: Memory should be treated the same way. RAM first. VRAM swap maybe. Disk swap last. But you need observability or you're just vibes with a kernel module.
JOSH: Is there a risk that AI builders jump on hacks like this because they want more local capacity?
ERIK: Absolutely. Local AI made everyone care about memory again. Model weights, context windows, embeddings, vector indexes, browser automation. It piles up fast.
ERIK: The lesson is not “use VRAM as swap.” The lesson is “your bottleneck moved.” For years people cared about CPU and disk. Now builders care about memory bandwidth, GPU utilization, token routing, and queue backpressure.
JOSH: Queue backpressure sounds less fun than GPU VRAM.
ERIK: It's less fun until your system doesn't fall over.
ERIK: ScanBrief is a good example. It scored 80 items across 54 sources today. That sounds small compared to internet-scale nonsense, but it still has stages. Fetch. Dedupe. Score. Summarize. Rank. Publish. If scoring stalls, the queue needs to absorb it. If summarization is slow, PrimeRouter needs to route smarter. If a source returns junk, it shouldn't poison the whole brief.
JOSH: So the real builder lesson is measure the pipe.
ERIK: Measure the pipe, then make the pipe boring. Fancy memory hacks are fine. But boring queues, boring retries, boring limits, boring dashboards. That's what keeps the thing alive.
[beat]
JOSH: “Make the pipe boring” is going on a sticker.
ERIK: Put it next to “your token leaked.”
[pause]
JOSH: Third deep story: AI outperforming law professors in a Stanford Law study. The headline sounds huge. What's the real signal?
ERIK: The signal is that expert readers are starting to prefer machine-generated work in domains we used to call too nuanced.
ERIK: Legal reasoning has structure. Facts. Rules. Exceptions. Precedent. Jurisdiction. The writing needs to be clear, complete, and careful. LLMs are good at producing that shape when prompted well and checked by someone competent.
JOSH: But legal work has risk. Hallucinated cases, wrong citations, bad advice.
ERIK: Yep. And that's why the conclusion is not “replace every lawyer with a chatbot.” The conclusion is “the first draft is dead.”
ERIK: Same thing happened in code. The blank file used to be expensive. Now the blank file is cheap. The expensive part is knowing what should exist, what can break, and what should never ship.
JOSH: How does that compare to network automation?
ERIK: Very similar. Claude can write NSO service code. It can write Terraform. It can generate Selenium tests. But if you don't understand the network, you're just generating confident outage material.
ERIK: In networking, the model can draft the change. The engineer owns the blast radius.
JOSH: That's a clean line.
ERIK: It's the only line that matters.
ERIK: Same in law. The model drafts. The lawyer owns the client, the ethics, the filing, and the consequences. If they don't use AI, they'll be slower. If they use AI without review, they'll be dangerous. The winners are the ones who build a review loop.
JOSH: That's basically your Gandalf pattern again.
ERIK: Exactly. Agents need reviewers. Reviewers need rules. Rules need telemetry.
ERIK: In the last 7 days, PrimeBus had 360 auto-merger runs: 228 merged, 132 blocked, 0 escalated. That tells me the guardrails are catching things before I become the bottleneck. That's the goal.
JOSH: Zero escalated is the fun part.
ERIK: It's fun because it means the system knows when it's allowed to act. Not because it's magic. Because I built the policy path.
JOSH: What would a legal team copy from that?
ERIK: Make the model produce structured outputs. Ask for claims, citations, uncertainty, and required human checks as separate fields. Don't accept a beautiful paragraph as the artifact. Beautiful paragraphs hide sins.
JOSH: That's good.
ERIK: Then run a verifier. Different model, different prompt, maybe a rules engine. Have it attack the answer. Missing authority. Weak analogy. Wrong jurisdiction. Unsupported claim. Then route low-confidence items to a human.
ERIK: That's HumanRail thinking. When the model isn't confident, or the policy says the impact is high, push it to a person. Don't pretend the robot became a partner because it used semicolons correctly.
JOSH: Dry but fair.
ERIK: Legal writing has a lot of semicolons. I don't make the rules.
JOSH: Does this change how people should learn?
ERIK: Yes. Stop learning as if memory is the product. Memory still matters, but judgment is the product.
ERIK: Builders should learn how to decompose work, design checks, read logs, spot weak assumptions, and test outputs. If you're a junior lawyer, junior engineer, junior analyst, whatever, your job is not to compete with the model at typing. Your job is to become the person who can tell when the model is wrong.
JOSH: That's harder.
ERIK: It is. But it's also better work.
ERIK: The people in trouble are the ones whose whole value was “I can produce a passable first draft.” AI just made passable first drafts cheap. That's not mean. That's math.
[pause]
JOSH: There's one more headline I want your quick take on: Microsoft showing MAI-Code-1-Flash. Another coding model. Do we need another one?
ERIK: Yes, if it gives builders more routing options. No, if it's just a benchmark trophy.
ERIK: I don't care what the launch post says. I care how it behaves in PrimeRouter against real tasks. Small patch. Big refactor. Test repair. Config migration. Bad repo. Good repo. That's where the truth lives.
JOSH: So model choice becomes infrastructure.
ERIK: Correct. Model choice is no longer a settings dropdown. It's routing policy.
ERIK: One model for cheap classification. One for code review. One for long-context architecture. One for nasty debugging. Use the right tool, log the result, kill the route if it gets dumb.
[pause]
ERIK: This episode is sponsored by Prime Automation Solutions. If you're still doing it manually, we automate it. Also, special on a website — $250. primeautomationsolutions.com
[pause]
JOSH: Alright, what's the AI pro tip today?
ERIK: Make your agent return evidence before action.
ERIK: Not a summary. Evidence.
ERIK: If it's fixing a bug, make it return the failing test, the file it changed, the exact reason the change should fix it, and the command it ran after. If it's writing Terraform, make it return the plan risk. If it's reviewing a PR, make it return the rule that blocked or approved it.
JOSH: So the prompt forces accountability.
ERIK: Yes. Use a two-step pattern.
ERIK: Step one: “Inspect only. Return findings, affected files, risks, and confidence. Do not edit.”
ERIK: Step two: “Apply the smallest change that addresses finding number one. Then run the narrowest useful test and report evidence.”
ERIK: That stops the model from wandering around your repo like it pays rent. It also gives you logs you can feed into a reviewer model or a policy gate.
ERIK: That's your tip. Use it.
[pause]
ERIK: If you're building toward financial independence through automation, my first book walks through the whole path. Free chapter at erikandersonbook.com.
[pause]
JOSH: One more thing — we started a Discord for builders. If you're shipping AI, automation, or anything that makes a human obsolete — come hang out. Link at buildorbereplaced.dev.
ERIK: Post what you built. We'll post what we're building. Real wins, real builds, no fluff.
[pause]
ERIK: Build or be replaced.
JOSH: If you want these signals in your inbox every morning, scanbrief.dev. See you tomorrow.