Mon–Fri · 6 AM ET
← All Episodes
EP  • 00:13:17

They’re made out of weights | Build or Be Replaced

Today: They’re made out of weights | Failing grades soar with AI usage, dwindling math skills in Berkeley CS classes | Elixir v1.20: Now a gradually typed language Episode date: 2026-06-04.

Download MP3 →

Transcript

JOSH: It's Thursday, June 04. This is Build or Be Replaced — powered by ScanBrief.dev. I'm Josh, here with Erik Anderson.
ERIK: AI isn't replacing builders. It's replacing people who treat the model like a vending machine.
JOSH: Stick around — Erik's got an AI pro tip at the end about making agents prove their work before they touch production.
[pause]
JOSH: ScanBrief scored 73 items across 54 sources today. First headline: Elixir just got gradually typed. Why should people care?
ERIK: Because mature languages are moving toward stronger feedback loops without making everyone rewrite their code. Elixir v1.20 adding gradual typing means old systems can start catching dead code and runtime type issues before they bite you.
[beat]
JOSH: Next one: Gemma 4 12B is being described as a unified, encoder-free multimodal model. Is that a big deal or just model-card poetry?
ERIK: It matters if it reduces moving parts. Text, image, and maybe more under one architecture means fewer glue systems around the model. Builders should watch that, because the boring integration work is where half the time disappears.
[beat]
JOSH: Third headline: Let's Encrypt is talking post-quantum certificates. Too early?
ERIK: No. Certificates have long tails. If your infra team waits until quantum attacks are practical, you're already behind. This is the kind of migration you test in the lab years before procurement discovers the acronym.
[pause]
JOSH: The top ScanBrief story today was a weird one: they're made out of weights. That's the whole argument. LLMs are just weights doing math.
ERIK: Yeah, and I love that because it's both obvious and offensive to people who want magic.
[beat]
ERIK: These models are giant piles of floating-point numbers. Matrix multiplication. Layers. Attention. More matrix multiplication. Out comes code, poetry, configs, bad advice, good advice, and occasionally something that looks like a junior engineer with sleep debt.
[beat]
ERIK: The point is not that the model is simple in behavior. The point is that the mechanism is brutally simple. Reasoning isn't in a separate module. There's no little symbolic engine hiding behind the curtain saying, "now I shall think." The capability is in the weights.
JOSH: That feels unsettling. If there's no explicit reasoning module, how do you trust it?
ERIK: You don't trust it. You test it.
[beat]
ERIK: That's the whole builder answer. People keep asking whether the model understands. Cool philosophy question. My production question is, can it produce an artifact that passes tests, matches constraints, and survives review?
[beat]
ERIK: PrimeBus processed 396 automation events across 7 projects today. None of those systems care if Claude has an inner life. They care if the event payload is valid, if the command is allowed, if Gandalf approves the diff, and if the deploy checks pass.
JOSH: So the lesson is not "AI is fake." It's "treat AI like an unreliable worker with receipts."
ERIK: Exactly. Give it tasks where success can be measured. Code patch. Terraform plan. NSO service change. Kubernetes manifest. Selenium scrape. Database migration. Anything where the output can be checked.
[beat]
ERIK: The mistake is asking the model to be right in prose. That's weak. Make it write the test. Make it run the test. Make a second agent review the test. Then make the system decide whether the change moves forward.
JOSH: Wait, really? You'd trust that more than a human review?
ERIK: Depends on the human.
[beat]
ERIK: A tired human rubber-stamping pull requests at 11 PM is not a safety system. A model with no guardrails is also not a safety system. But a pipeline where Claude generates a fix, Gandalf reviews it, tests run, and PrimeBus records the event trail? That's a system.
[beat]
ERIK: Overnight, 11 code changes were automatically reviewed by Gandalf and merged to production. That's not me hoping the model is smart. That's automation with gates.
JOSH: And when the gates say no?
ERIK: Then it doesn't merge. PrimeBus auto-merger has run 4280 attempts since 2026-04-17: 1240 merged, 3028 blocked by Gandalf, 12 escalated to Erik. That's the story. Not the merge rate. The blocked count is the value.
[beat]
ERIK: People hear blocked and think failure. Wrong. Blocked means the guardrail worked. That's how I sleep.
[pause]
JOSH: That connects to another story today: Ted Chiang saying artificial intelligence is not conscious. Does that matter to builders?
ERIK: Not much.
[beat]
ERIK: I respect the argument. I get why people want clean language around consciousness, agency, intent, all of that. But if you're running infrastructure, the question is not whether Claude is conscious. The question is whether Claude can produce a bad change faster than your controls can stop it.
JOSH: That's darker than I expected.
ERIK: It's practical.
[beat]
ERIK: A forklift doesn't need consciousness to put a hole in a wall. A model doesn't need consciousness to delete the wrong table, rotate the wrong key, or ship a broken route-policy to a router. You don't secure tools based on their feelings. You secure them based on blast radius.
JOSH: So containment matters more than the debate.
ERIK: Yes. And that was another ScanBrief item today: the ways people contain Claude across products. That's the right conversation. Tool permissions. Filesystem boundaries. Network access. Human approval. Audit trails. Sandboxes. Rate limits. Dry runs.
[beat]
ERIK: Same stuff we've done in network automation forever. You don't let a script push to the core because it printed "looks good." You do pre-checks, diffs, maintenance windows, rollback plans, post-checks, and logs. AI doesn't remove any of that. It makes it more important.
JOSH: How does that show up in your lab?
ERIK: Echo and Neo run the Bobaverse fleet. Five agents in the fleet right now: Neo, Homer, Bill, Echo, Gandalf. Claude plus GPT. They don't all get the same permissions.
[beat]
ERIK: Gandalf is a reviewer. PrimeSentinel watches for stuck jobs. PrimeBus carries the events. Some agents can suggest. Some can write. Some can merge only after checks pass. That separation matters.
JOSH: That's more like a small engineering org than a chatbot.
ERIK: Duuude, exactly. That's the mental shift.
[beat]
ERIK: Stop thinking "chat window." Think "worker on a bus." The worker subscribes to events. It gets a bounded task. It emits a result. Another worker checks it. The bus records it. If something smells wrong, it escalates.
[beat]
ERIK: That's how you get useful work out of weights. Not by begging the model to be wise. By putting it inside a system where dumb output gets caught and good output moves.
JOSH: What does that mean for someone who doesn't have your setup?
ERIK: Start smaller.
[beat]
ERIK: Take one annoying thing. Maybe a flaky test. Maybe a weekly report. Maybe a router config audit. Make the model produce a patch or report. Then add one check. Then add a second check. Then make the result land somewhere automatically.
[beat]
ERIK: Don't start with "AI runs my company." Start with "AI opened a pull request, and the pull request had evidence."
JOSH: Evidence meaning logs, tests, diffs?
ERIK: Yes. Evidence beats vibes every day.
[pause]
JOSH: The Berkeley story got a lot of attention today. Failing grades soared in CS 10 and CS 61A, and instructors are pointing at AI use, cheating, and weaker math skills. What's your read?
ERIK: The bill came due.
[beat]
ERIK: Students used AI like an answer machine. Some probably got through assignments without building the mental model. Then the exam shows up, or the project changes shape, and suddenly the model isn't there to carry the whole load.
JOSH: That's the fear people have at work too, right?
ERIK: Yeah. And it's valid.
[beat]
ERIK: If you use Claude to avoid learning, you're renting competence. If you use Claude to compress feedback loops, you're building competence faster. Huge difference.
JOSH: Give me the line between those two.
ERIK: If you can't explain the output, you're not done.
[beat]
ERIK: That's it. Claude can write the function. Fine. But can you explain the edge cases? Can you change it when the API moves? Can you debug it when the test fails? Can you tell me why that SQL query is safe? If not, you didn't build. You copied.
JOSH: That's blunt.
ERIK: Good.
[beat]
ERIK: This matters because AI makes weak foundations look fine until they don't. Same with networking. Someone can paste an NSO service template they don't understand. It might deploy. But when the device state drifts, or the rollback hits a weird corner, now they own a problem they can't reason about.
JOSH: So how should students or junior engineers use it?
ERIK: Make it your coach, not your ghostwriter.
[beat]
ERIK: Ask it to quiz you. Ask it to generate three broken examples. Ask it why your solution fails. Ask it to explain the same code as if you're debugging production at 2 AM. Then close the model and do it yourself.
[beat]
ERIK: For math, same thing. Don't ask for the final answer first. Ask for hints. Ask for a similar problem. Ask it to grade your work. If you skip the struggle completely, you skip the wiring in your own head.
JOSH: That sounds slower.
ERIK: Short term, yes. Long term, no.
[beat]
ERIK: Builders need taste and judgment. You don't get that from letting the model make every decision. You get it by comparing outputs, running tests, reading failures, and fixing the thing.
JOSH: How do you keep yourself honest with that?
ERIK: I make the system embarrass me in writing.
[beat]
ERIK: PrimeBus logs the event. Gandalf says why something should be blocked. PrimeSentinel tells me when a job is stuck. ScanBrief ranks sources and relevance instead of letting me cherry-pick whatever headline sounds cool.
[beat]
ERIK: ScanBrief scored 73 items across 54 sources today with an average relevance score of 54.2. That keeps me honest. I don't wake up and manually pick stories that flatter my priors. The system gives me a stack, then I argue with it.
JOSH: That's a good phrase: argue with it.
ERIK: That's the job now.
[beat]
ERIK: Don't worship the model. Don't dismiss it either. Argue with it. Make it show work. Make it run the command. Make it cite the file. Make it produce the diff. Make it survive a second model.
[pause]
JOSH: The Uber AI pricing story is also interesting. A $1,500 a month AI limit is a useful signal. What signal?
ERIK: That serious AI use is becoming a line item, not a toy subscription.
[beat]
ERIK: People got trained on twenty bucks a month. That's cute for chat. It is not the number for agents doing real work all day. If AI is replacing contractor hours, analyst hours, QA time, support time, or engineering review time, the pricing will move toward the value of that work.
JOSH: So builders should expect higher AI bills?
ERIK: Yes, and they should measure them like infrastructure.
[beat]
ERIK: Tokens are not magic dust. They're compute. If your agent is burning model calls to summarize the same file twenty times, fix the workflow. Cache context. Use smaller models for boring classification. Use Claude Sonnet for code and reasoning where it earns the spend. Use local checks before model checks.
JOSH: Where do people waste the most?
ERIK: Letting the big model do tiny jobs.
[beat]
ERIK: Classification. Routing. Deduping. Formatting. Basic extraction. Those don't always need the expensive brain. ScanBrief can score and rank a lot before the final summary step. PrimeBus can route events without asking a model what a queue is. Please don't pay premium token rates for "is this JSON valid."
JOSH: That's painfully specific.
ERIK: Because I've done dumb things too.
[beat]
ERIK: Every builder has. You wire something up, it works, then the bill arrives and now your genius automation has the financial discipline of a teenager with a debit card.
JOSH: What's the fix?
ERIK: Budget by workflow.
[beat]
ERIK: Not by model. Workflow. How much is this task worth? What happens if it fails? Does it need the best model? Can a cheaper model do the first pass? Can deterministic code do half the job before AI enters?
[beat]
ERIK: That's how you decide. If InkEngine is drafting book chapters, model quality matters. If PrimeSentinel is checking whether a process is stuck, use code. If Gandalf is reviewing a production diff, spend the tokens. Cheap review is expensive when it misses the bad change.
JOSH: That's the cleanest AI pricing framework I've heard.
ERIK: It's just engineering.
[beat]
ERIK: Treat models like any other dependency. Latency, cost, reliability, permissions, logs. Pick the right one for the job. Then measure it.
[pause]
ERIK: This episode is sponsored by Prime Automation Solutions. If you're still doing it manually, we automate it. Also, special on a website — $250. primeautomationsolutions.com
[pause]
JOSH: Alright, what's the AI pro tip today?
ERIK: Make your agent produce a proof bundle before it can act.
[beat]
ERIK: Simple version: every AI-generated change must include four things. The diff. The reason. The test or command it ran. The rollback plan.
[beat]
ERIK: Put that in your prompt. Better, put it in your pipeline. If any piece is missing, the agent doesn't get to continue.
[beat]
ERIK: Example: "Before opening the pull request, write a summary of files changed, list the exact verification command, include the output, and describe how to revert this change." Then have another model or script check that all four exist.
[beat]
ERIK: This works today with Claude, Cursor, GitHub Actions, Jenkins, whatever. It turns AI from "trust me bro" into "here's the receipt."
[beat]
ERIK: That's your tip. Use it.
[pause]
JOSH: Track your freedom score and net worth with the Freedom Blueprint app — free download, link in the show notes.
[pause]
JOSH: One more thing — we started a Discord for builders. If you're shipping AI, automation, or anything that makes a human obsolete — come hang out. Link at buildorbereplaced.dev.
ERIK: Post what you built. We'll post what we're building. Real wins, real builds, no fluff.
[pause]
ERIK: Build or be replaced.
JOSH: If you want these signals in your inbox every morning, scanbrief.dev. See you tomorrow.