Mon–Fri · 6 AM ET
← All Episodes
EP  • 00:12:42

The newest Instagram “exploit” is the goofiest I've seen | Build or Be Replaced

Today: The newest Instagram “exploit” is the goofiest I've seen | Can the stockmarket swallow Anthropic, SpaceX and OpenAI? | macOS needs its grid back Episode date: 2026-06-02.

Download MP3 →

Transcript

JOSH: It's Tuesday, June 2. This is Build or Be Replaced — powered by ScanBrief.dev. I'm Josh, here with Erik Anderson.
ERIK: Automated support agents are now part of the attack surface. If you didn't design for that, somebody else is testing it for you.
JOSH: Stick around — Erik's got an AI pro tip at the end about making agents prove their work before they touch anything important.
JOSH: First headline. Instagram support AI got tied to a password reset exploit. What's the real story?
ERIK: The real story is trust boundaries. If a support bot can influence account recovery, that bot is now part of identity infrastructure. Treating it like a help widget is how you get smoked.
JOSH: Second one. Public markets are asking whether they can absorb companies like Anthropic, SpaceX, and OpenAI. Finance noise, or real signal?
ERIK: Real signal. These companies are being valued like cloud platforms, defense contractors, and energy consumers at the same time. The question isn't hype anymore. It's who pays the compute bill and what margin is left after the GPUs eat first.
JOSH: Third headline. Stanford posted AI agent rules for CS336. Why should builders care about a class policy?
ERIK: Because it's the same governance problem companies are avoiding. Where can an agent act? What has to be cited? What needs human review? That's not school policy. That's production policy in a hoodie.
JOSH: Let's start with Instagram. The phrase "support AI password reset exploit" sounds like a sentence from a security team's bad dream.
ERIK: It is. Password reset is not a support task. It's account ownership. That's the spare key to the house. If an automated flow can be convinced to send verification codes somewhere the attacker controls, the failure isn't that an AI said the wrong thing. The failure is that the system let a conversation cross into identity authority.
JOSH: What's the line there?
ERIK: The line is simple. A support agent can gather context. It can explain steps. It can open a ticket. It can summarize the user's issue for a human. It should not decide who owns the account. That decision needs hard proof outside the model.
JOSH: Wait, really? Even if the user gives a bunch of correct details?
ERIK: Correct details are not ownership. A breached email inbox can have correct details. A scraped profile can have correct details. A clever attacker can sound more organized than the real owner. Identity needs verified contact control, session history, device binding, hardware keys, recovery codes, or human escalation. A model saying "this looks right" is not one of those things.
JOSH: That feels like the new part. The exploit isn't a classic code bug. It's social engineering the machine.
ERIK: Exactly. People think the model is the product. It isn't. The workflow around the model is the product. The attacker doesn't need to break Claude or GPT. They need to find the part of the workflow where the model's output is treated like authorization. That's the weak spot.
JOSH: How would you design it?
ERIK: Start with permission tiers. Tier one, the agent can talk. Tier two, it can draft. Tier three, it can open internal tickets. Tier four, it can request a state change, but another system decides. Tier five, human review. Account recovery, billing changes, permissions, legal changes, network changes, production deploys, all of that lives above the model.
JOSH: So the model can recommend, but it doesn't get the keys.
ERIK: Right. That's the whole thing. In PrimeBus, agents can observe and propose actions. They can produce fix candidates. They can compare logs. But the bus is not a free-for-all. Events have source, type, handler, and policy. If something touches a risky path, it gets checked before anything moves.
JOSH: That sounds very network-engineer.
ERIK: It is. In Cisco NSO, you don't let random text mutate the network because it sounded confident. You validate the service model. You check device state. You dry run. You compare intended config to real config. Then you commit. AI needs the same discipline. The model is not a change-control board with nicer punctuation.
JOSH: What should teams check today if they already have AI support live?
ERIK: Pull every support workflow where the AI can touch account ownership, payments, refunds, permissions, security settings, or data export. Then ask one question. What proof is required before the change happens? If the answer is "the model decides," shut that path down and rebuild it.
JOSH: That's blunt.
ERIK: It should be. Support AI is fine. I like support AI. It can kill a lot of ticket noise. But identity needs deterministic controls. Verified email. Existing session. Known device. Hardware key. Human review. Audit trail. Those are controls. "The user sounded sad and urgent" is not a control.
JOSH: Let's move to the money story. Markets are looking at Anthropic, OpenAI, SpaceX, and wondering if there is enough capital to absorb all of this. What are they really asking?
ERIK: They're asking whether AI is a software margin story or an infrastructure story. Software margins are beautiful. Infrastructure margins are expensive. AI is both. That's why the market is confused.
JOSH: How so?
ERIK: A normal SaaS company sells seats and pays cloud bills. Big AI labs sell intelligence by the token and pay for chips, power, networking, datacenters, research, safety teams, enterprise sales, and model training runs that can cost stupid money before one customer pays an invoice. That's a different animal.
JOSH: So when people talk about valuations, the hidden question is compute?
ERIK: Yep. Compute is the factory. Tokens are the product. GPUs are the machines. Power is the fuel. Networking is the conveyor belt. If you don't understand that, you look at the revenue and miss the cost structure underneath it.
JOSH: Does that change how a builder should think about choosing models?
ERIK: Absolutely. Don't use the smartest model for every task. That's lazy architecture. Use the right model for the job. I use Claude when I need reasoning, code review, or careful system work. But for classification, routing, scoring, and boring text cleanup, smaller models or rules can be enough. ScanBrief doesn't need a PhD to know a press release is fluff.
JOSH: That's the pro tip teaser showing up early.
ERIK: It all connects. Cost discipline matters. If your agentic system sends every tiny decision to the most expensive model, you're building a slot machine with a repo attached.
JOSH: How do you think about that in your lab?
ERIK: Echo and Neo are where I test the system behavior. PrimeDash shows me what's running, what failed, what got noisy, and what needs a human. The point is visibility. If an AI workflow costs money, changes state, or wakes me up, I want telemetry on it.
JOSH: Most teams don't have that, do they?
ERIK: A lot don't. They have prompts in a product and vibes in a spreadsheet. That's not enough anymore. If you're using agents, you need traces. Which model ran? What tool did it call? What prompt version? What input? What output? What action did it request? Who approved it? If you can't answer that, you don't have automation. You have a mystery box with billing.
JOSH: That's a little harsh.
ERIK: It's true. And it doesn't mean you need a giant platform. Start small. Log every tool call. Store the prompt version. Make the agent return structured output. Put a policy check between recommendation and action. Then measure what gets approved, blocked, retried, and escalated. You learn fast when the system has receipts.
JOSH: Where does PrimeTrader fit into that thinking?
ERIK: PrimeTrader is a good example because trading is unforgiving. TradingView webhooks come in, signals get evaluated, and the system has to be boring on purpose. You don't want a creative agent in the middle of a money path. You want rules, risk checks, clear state, and logs. AI can help read market context or summarize movement, but execution has to be controlled.
JOSH: Same pattern again. Agent reads, rules decide.
ERIK: Exactly. Agents are great at interpretation. They're not magic permission engines. That's the distinction people keep missing.
JOSH: Let's hit Stanford's CS336 agent rules. Why did that stand out to you?
ERIK: Because they made the invisible rules visible. Students can use agents, but there are boundaries. You have to know what work is yours. You have to attribute help. You can't outsource the whole assignment and pretend you learned it. That maps directly to companies.
JOSH: In what way?
ERIK: Companies are already using agents to write code, write tickets, summarize incidents, draft customer emails, and generate configs. Fine. But who owns the output? What needs review? What can ship without a human? What must cite sources? What happens when the agent fabricates a dependency, a command, or a legal claim?
JOSH: That's the part people don't want to write down.
ERIK: They need to. A policy that lives in someone's head is not a policy. Write it down as system rules, code checks, and workflow gates. Don't just tell engineers "use AI responsibly." That's corporate oatmeal. Give them exact lines.
JOSH: Like what?
ERIK: "AI may draft code, but tests must pass before merge." "AI may summarize incidents, but cannot close customer-impacting tickets without owner review." "AI may generate Terraform, but plan output must be reviewed before apply." "AI may classify support requests, but cannot change identity data." That's useful. That's something people can follow.
JOSH: What about attribution? Does that matter outside school?
ERIK: It matters a lot. Not because I care about giving a chatbot a gold star. I care because attribution creates traceability. If a customer email was drafted by AI, say that internally. If a code change came from Claude, tag it. If a runbook step was generated, mark it and verify it. When something breaks, you need to know where it came from.
JOSH: Is this where teams get nervous and ban everything?
ERIK: Some will. That's the wrong move. Banning AI because you're afraid of bad output is like banning automation because a bad script can break a router. The answer is not fear. The answer is controls. Start in the lab. Prove wins. Add guardrails. Move to production when the system earns it.
JOSH: That sounds like your whole automation philosophy.
ERIK: It is. I left school at 16 and learned by building systems that either worked or didn't. Later I got the degree because I wanted to close that loop, but the engineering lesson was already there. Small tests. Real feedback. More guardrails. Then production. That's how you avoid dramatic failure stories. You catch the small stuff before it becomes the big stuff.
JOSH: What does this mean for engineers listening today?
ERIK: It means your job is shifting. The best engineer in the room won't be the person who types the most code. It'll be the person who can design the system around the agents. Tool boundaries. Policy checks. Observability. Rollback. Cost control. Security. That's where the value is.
JOSH: And the people who ignore it?
ERIK: They become the manual approval step somebody else automates. That's not a threat. That's the map.
JOSH: Give me the practical version. One thing to build this week.
ERIK: Build an agent approval gate. Pick one workflow. Could be code review, support replies, Terraform plans, or incident summaries. Make the agent produce three fields: proposed action, evidence, and risk. Then force a policy check before anything happens. Low risk can pass. Medium risk needs review. High risk stops. That's how you start turning a chatbot into infrastructure.
JOSH: Evidence is the key word there.
ERIK: Yep. Make it show receipts. Logs, links, tests, config diffs, source articles, ticket history. Whatever applies. If the agent can't point to evidence, it doesn't get to act. That's not anti-AI. That's how you make AI useful.
JOSH: This episode is sponsored by Prime Automation Solutions. If you're still doing it manually, we automate it. Also, special on a website — $250. primeautomationsolutions.com
JOSH: Alright, what's the AI pro tip today?
ERIK: Build a model router before you build another agent. Seriously. Take your workflow and split tasks into buckets. Reasoning, coding, summarization, classification, extraction, search, and action. Then assign the cheapest reliable method to each bucket. Rules beat models when the rule is obvious. Small models beat big models when the task is boring. Big models get used when judgment matters.
JOSH: What's the action step?
ERIK: Add one field to every AI task: required_model_reason. Make the agent or workflow explain why it needs that model. If it can't justify Claude Opus-level thinking, don't give it Claude Opus-level spend. Route it down. Test the output. Track failures. Your bill drops and your system gets easier to reason about.
JOSH: That's specific.
ERIK: It matters. Don't use the smartest model as a personality trait. Use it where it wins. That's your tip. Use it.
JOSH: We also drop daily market picks and automation tips on YouTube — search Build or Be Replaced.
JOSH: One more thing — we started a Discord for builders. If you're shipping AI, automation, or anything that makes a human obsolete — come hang out. Link at buildorbereplaced.dev.
ERIK: Post what you built. We'll post what we're building. Real wins, real builds, no fluff.
ERIK: Build or be replaced.
JOSH: If you want these signals in your inbox every morning, scanbrief.dev. See you tomorrow.