The lamps in my house | Build or Be Replaced
Today: The lamps in my house | ChatGPT is adding real cartoonists' signatures to fake New Yorker cartoons | Beam: Reflection's 501B open-weight model Episode date: 2026-10-06.
Download MP3 →
Build or Replaced
Today: The lamps in my house | ChatGPT is adding real cartoonists' signatures to fake New Yorker cartoons | Beam: Reflection's 501B open-weight model Episode date: 2026-10-06.
Download MP3 →JOSH: It's Tuesday, October 6. This is Build or Be Replaced — powered by ScanBrief.dev. I'm Josh, here with Erik Anderson. ERIK: Today's theme is simple: AI is getting bigger, weirder, and way less theoretical. The people treating it like a toy are going to have a bad quarter. JOSH: Stick around — Erik's got an AI pro tip at the end about making agents explain their next move before they touch anything. [pause] JOSH: First headline. Reflection released Beam, a 501-billion-parameter open-weight model. Big number. Real signal? ERIK: Yeah. It’s sparse Mixture-of-Experts, so only 23 billion parameters are active at a time. That matters because inference cost is the whole ballgame now. Big brains are cool. Big brains you can afford to run are better. JOSH: Next, ChatGPT is apparently adding real cartoonists' signatures to fake New Yorker-style cartoons. That sounds messy. ERIK: It’s messy because signatures are identity, not style. If a model imitates a joke format, fine, argue about taste. If it drops a real artist’s mark on fake work, now you’re in provenance and trust territory. JOSH: Third headline. Opus 5.5 agents found two room-temperature magnetic semiconductor candidates. Wait, really? ERIK: Really, and that’s the interesting part. Agentic discovery is moving from “write me code” into “go search a design space humans can’t manually cover.” That’s where this gets uncomfortable for entire research workflows. [pause] JOSH: Beam is the big model story today. What actually happened there? ERIK: Reflection released Beam, a 501-billion-parameter sparse MoE model with 23 billion active parameters. Trained on 23.8 trillion tokens, then hit with high-compute reinforcement learning. Translation: they didn’t just make a big text predictor. They pushed it toward coding, reasoning, and agent behavior. [beat] ERIK: The important part is the sparse design. Everyone hears 501 billion and thinks, “cool, I need a data center.” But if only 23 billion are active for a token, the runtime story changes. You still need serious hardware, but you’re not paying the full dense-model tax on every request. JOSH: So this is more about economics than bragging rights? ERIK: Exactly. Model size used to be the headline. Now usable cost is the headline. If a model can reason well, code well, and run cheaper per useful answer, that’s where builders start caring. JOSH: How does that compare to what you run? ERIK: In my world, model choice is a routing problem. PrimeBus doesn’t care about brand names. It cares about job type, cost, latency, and failure mode. I’ve got 12 agents in the Bobaverse fleet right now across Neo, Homer, Bill, Echo, Gandalf, Claude, and GPT. The question is never “what’s the best model?” The question is “which model should touch this task, and what guardrails sit around it?” JOSH: Give me an example. ERIK: Code review is different from summarizing ScanBrief. A review agent needs to inspect diffs, reason about tests, and understand blast radius. A brief agent needs to rank signals and not hallucinate what happened. ScanBrief scored 100 items across 56 sources today. That job is mostly retrieval, scoring, dedupe, and summarization. I don’t need the most expensive brain for every item. [beat] ERIK: But if PrimeBus sees a production failure and spins up a fix attempt, now I want a stronger model or a multi-agent path. First agent diagnoses. Second proposes a patch. Gandalf checks the blast radius. Tests run. Then maybe the auto-merger touches it. JOSH: You mentioned Gandalf. That’s the guardrail agent? ERIK: Yeah. Gandalf is the one standing there saying, “Nope, you don’t get to ship that.” PrimeBus auto-merger has run 2295 attempts since June 5: 1507 merged, 788 blocked by Gandalf, 0 escalated to me. That blocked number is the story. That’s not failure. That’s the system refusing bad changes before I have to care. JOSH: That’s wild. ERIK: It’s what happens when you stop treating AI like a chat window and start treating it like an unreliable contractor with API access. You don’t yell at it when it gets something wrong. You put it in a workflow where wrong answers get caught. [pause] JOSH: Does Beam change the open model side of that? ERIK: Potentially. Open weights matter because teams can inspect, tune, host, and route without waiting on a vendor dashboard to bless them. For network automation, government work, client data, or lab systems, that matters. Sometimes you can’t send the payload to a hosted model. Sometimes you just don’t want to. JOSH: But open-weight doesn’t mean easy, right? ERIK: Not at all. A 501-billion-parameter sparse model is not “download it on your laptop and go make coffee.” People confuse open with cheap. Open gives you control. It does not magically pay your GPU bill. [beat] ERIK: The builder takeaway is this: design your systems so models are replaceable. Don’t wire your whole product around one provider, one model name, and one prompt. Wrap it. Measure it. Log every decision. Put the model behind a bus if you can. NATS, queues, events, whatever. The shape matters more than the brand. JOSH: That’s very PrimeBus. ERIK: Yeah, because it works. I don’t want 140 distinct projects emitting telemetry to PrimeBus because I like dashboards. I want every system speaking the same event language so agents can subscribe, react, fix, block, or tell me something changed. That’s infrastructure. The model is just one worker in the shop. [pause] JOSH: The cartoon signature story feels completely different, but maybe it’s the same trust problem. ERIK: It is. ChatGPT generating fake New Yorker-style cartoons and attaching real cartoonists' signatures is a provenance problem. That’s not just “AI art is weird.” That’s a system creating false attribution. JOSH: Why is the signature line different from style imitation? ERIK: Style is fuzzy. You can argue forever about influence, parody, training data, all of it. A signature is a claim. It says this person made this. If the model puts a real artist’s signature on generated work, it’s forging context. [beat] ERIK: Same issue shows up in code. If an agent writes a patch and commits it under a human’s name without metadata, that’s bad provenance. If it changes a Terraform file and nobody can trace the prompt, test result, review status, and merge decision, you’re asking for pain. JOSH: What does good provenance look like in an AI system? ERIK: Every output needs a chain of custody. Input source, model used, prompt version, tool calls, tests, reviewer, final actor. Not because compliance people love paperwork. Because when something breaks, you need to know which part lied. JOSH: That sounds heavier than most people’s AI setup. ERIK: Most people’s AI setup is a browser tab and vibes. That’s fine for a poem. It’s not fine for changing a firewall policy, billing a client, or sending mail on behalf of a business. [pause] JOSH: Where does this show up in your own builds? ERIK: email-reactor is a good one. It reads client website change requests from email and can turn them into work. That sounds simple until you realize email is chaos wearing a tie. Clients say “move the thing up a little,” attach the wrong screenshot, reply to an old thread, or ask for three changes in one sentence. [beat] ERIK: I don’t let the agent just edit production because it sounded confident. It parses the request, ties it to the client, checks the site context, creates a proposed action, and records what happened. If it’s not confident, it doesn’t pretend. It routes to a human or asks for clarification. JOSH: So the signature problem is really an audit trail problem? ERIK: Yep. The cartoon story is a consumer version of the same failure. The model produced something that looked official but wasn’t. In enterprise systems, that’s how you get mystery changes, fake approvals, bad invoices, and logs nobody trusts. JOSH: That’s a fun list. ERIK: Yeah, real party. Bring snacks and an incident bridge. [beat] JOSH: How do builders avoid it? ERIK: Put identity boundaries around generated output. Label generated artifacts. Keep model metadata. Don’t let agents impersonate people. If an agent sends an email, it should be clear that automation prepared it or sent it under an approved workflow. If an agent commits code, the commit should say what agent touched it and what test evidence exists. JOSH: Does that slow things down? ERIK: A little. But bad speed is expensive. PrimeBus has 146 services running on the production server right now. If I let agents move fast with no identity, no telemetry, and no rollback path, I’d spend my life cleaning up my own cleverness. No thanks. [pause] JOSH: Third deep dive. AI agents finding room-temperature magnetic semiconductor candidates. This feels like science fiction, but it was in the feed today. ERIK: This one is huge because the agent isn’t just answering a question. It’s searching a scientific design space. Opus 5.5 agents identified two room-temperature magnetic semiconductor candidates, including one newly designed compound and one material that had been synthesized back in 1999. JOSH: Why does that matter? ERIK: Materials science is brutal. There are too many combinations, too many constraints, and too much literature for humans to grind through manually. An agent can read papers, form hypotheses, check known materials, propose candidates, and loop through reasoning steps faster than a human team doing the first pass. JOSH: Are we talking replacement of scientists? ERIK: Not clean replacement. More like compression. The boring search work gets eaten first. The scientist still has to validate, synthesize, measure, and challenge the result. But the front end of discovery changes. You don’t start with “what should we try?” You start with “the agent found 80 candidates, ranked them, showed its work, and flagged two that look worth lab time.” [beat] JOSH: That’s a big change. ERIK: It is. Same pattern as network automation. Years ago, engineers logged into boxes and typed commands. Then we got NSO, templates, service models, CI pipelines, Terraform, Kubernetes, all that. The work didn’t disappear. The manual part got squeezed. The people who understood systems moved up the stack. The people who only typed commands got exposed. [pause] JOSH: How would you connect that to builders listening today? ERIK: Stop asking, “Can AI do my whole job?” That’s the wrong question. Ask, “Which part of my job is search, comparison, translation, validation, or repetition?” That’s where agents land first. [beat] ERIK: GovContractScanner is a good example from my side. Federal contract opportunities are not hard because one record is impossible to read. They’re hard because there’s too much noise, bad timing, and matching logic. The scanner watches, filters, tracks primes, and surfaces what’s worth attention. That’s not magic. It’s a machine doing the boring first pass every day without getting tired. JOSH: And the human still decides what to pursue. ERIK: Correct. But now the human starts at the short list. That’s the win. The same thing happens in research. The same thing happens in sales. Same thing in incident response. Agents narrow the field, propose actions, and collect evidence. Humans approve the weird stuff until the weird stuff becomes normal. JOSH: Where do people mess this up? ERIK: They give the agent a vague goal and too much authority. “Go find opportunities” is weak. “Watch these sources, score against this profile, reject anything outside these NAICS codes, summarize the top matches, and open a review item when confidence is above this threshold” is an actual workflow. JOSH: That’s much less glamorous. ERIK: Correct. Glamour is usually where systems go to die. [pause] JOSH: What about the smaller stories today? The flattest route in San Francisco was interesting. ERIK: Loved that one. It computes distance versus elevation using street segments, lidar elevation data, and map graphs. That’s a great reminder that useful software is often just good data plus a clear constraint. Shortest route is not always best route. Anyone who has walked up a San Francisco hill with a laptop bag knows this. JOSH: Example.com getting a redesign also made the list. That feels oddly emotional for developers. ERIK: Example.com is internet furniture. When stuff like that changes, people notice because boring stable things are rare now. There’s a lesson there. If your tool is trusted, don’t redesign it into a slot machine. JOSH: And Common Lisp being called the best programming language? ERIK: Every few months, someone rediscovers Lisp and becomes spiritually unavailable for a week. Good language. Powerful ideas. Also, your production team still has to maintain the thing after your enlightenment. [pause] ERIK: This episode is sponsored by Prime Automation Solutions. If you're still doing it manually, we automate it. Also, special on a website — $250. primeautomationsolutions.com [pause] JOSH: Alright, what's the AI pro tip today? ERIK: Before you let an agent use a tool, make it write a one-sentence action plan and a stop condition. [beat] ERIK: Not a giant essay. One sentence. “I’m going to inspect the failing test, patch the parser, and stop if the change touches authentication.” That little step catches a lot of nonsense before it happens. JOSH: Why does that work? ERIK: It forces the model to declare intent. Then your wrapper can check it. If the plan says parser and the tool call edits billing code, block it. If the stop condition is triggered, halt and ask for review. This is cheap, easy, and you can add it today in a shell script, a GitHub Action, or a PrimeBus-style event flow. [beat] ERIK: Don’t ask the model to be trustworthy. Make it observable. That’s your tip. Use it. [pause] JOSH: We also drop daily market picks and automation tips on YouTube — search Build or Be Replaced. [pause] JOSH: One more thing — we started a Discord for builders. If you're shipping AI, automation, or anything that makes a human obsolete — come hang out. Link at buildorbereplaced.dev. ERIK: Post what you built. We'll post what we're building. Real wins, real builds, no fluff. [pause] ERIK: Build or be replaced. JOSH: If you want these signals in your inbox every morning, scanbrief.dev. See you tomorrow.