Transcript
JOSH: It's Monday, June 29. This is Build or Be Replaced — powered by ScanBrief.dev. I'm Josh, here with Erik Anderson.
ERIK: AI is now judging resumes, scanning code, and checking IDs. Cool. Also terrifying if nobody can replay what it did.
JOSH: Stick around — Erik's got an AI pro tip at the end about using a critic agent before you trust a model's answer.
JOSH: [pause]
JOSH: First headline. HackerRank open sourced an AI hiring tool, and people are already finding weird resume scores. What happened?
ERIK: Same resume, different scores. That's the problem. If a hiring system gives you a 66 once and a 99 later, that isn't intelligence. That's a slot machine with a calendar invite.
JOSH: [beat]
JOSH: Second. GLM 5.2 beat Claude Code in one security benchmark. Should Claude users panic?
ERIK: No. Benchmarks are useful, but only if you understand the harness. GLM did well on IDOR detection. The real story is that the model is only one part of the security system.
JOSH: [beat]
JOSH: Third. Age verification is back in the policy feed. Why does that keep ranking so high?
ERIK: Because age checks don't stay age checks. Once the internet gets identity gates, the next question is who gets the keys and what else they check.
JOSH: [pause]
JOSH: Start with the resume scoring story. Automated hiring has been around forever. Why does this one feel different?
ERIK: Public code changes the whole conversation. Before, a vendor could say, trust us, the system is fair. Now builders can run it, break it, and show receipts.
ERIK: HackerRank open sourced Hiring Agent, and people started testing it. Same resume, different outcomes. That sounds like a minor bug until you remember the output can decide whether a human ever sees the candidate.
JOSH: Wait, really? Same resume, same job, different score?
ERIK: Yep. And anybody who has built with LLMs should not be shocked. Models infer. They summarize. They react to formatting. They can treat a bullet order like it means competence.
ERIK: That's fine when you're drafting a note. It's not fine when you're ranking humans. Hiring needs repeatability, logs, replay, model versioning, prompt versioning, and an appeals path.
JOSH: That's the part people miss. It has to be explainable.
ERIK: Exactly. The AI output can't be the final authority. It should be evidence attached to a decision. Show me the skill match. Show me the job requirement. Show me the sentence in the resume that triggered the score.
ERIK: If the score moves and the facts didn't move, the system is not ready. That's harsh, but it's true.
JOSH: How would you build it?
ERIK: First, separate extraction from judgment. Use strict parsing for facts: dates, titles, employers, skills, certifications, projects. Then score those facts in a visible rules layer.
ERIK: If AI is involved, make it cite the exact evidence. Not a fluffy explanation. Actual lines. This requirement maps to this evidence. This gap exists because no evidence was found.
ERIK: Second, freeze the run. Same resume, same job, same model, same prompt, same settings, same output. If it changes every time, you don't have a hiring system. You have confetti.
JOSH: That's a very network automation answer.
ERIK: It is. Cisco NSO taught us this forever ago. If I push the same service template twice, I expect the same device config. If it changes because the template felt different today, that's not automation. That's a ticket generator.
ERIK: Same with PrimeBus. I don't let an agent merge code because it had a good feeling. Events have structure. Reviews have gates. Failures have routes. The system has to prove why it took an action.
JOSH: So the blocked action matters as much as the successful one.
ERIK: More, honestly. Everybody wants the machine to say yes. The money is in making it say no correctly. No merge. No send. No delete. No candidate rejection without evidence.
ERIK: That's where builders need to grow up a little. The cool demo is easy. The boring guardrail is the product.
JOSH: What does this mean for job seekers?
ERIK: Plain text wins. Boring formatting wins. Honest keyword matching wins. Don't get cute with columns, graphics, hidden text, or some weird PDF layout that looks good to a human and parses like soup.
ERIK: Take the job description, take your resume, and run a test loop. Ask Claude to extract the hard requirements. Then ask it to map your resume evidence line by line. Then ask a second model to attack the match.
JOSH: Attack it how?
ERIK: Ask: where would an ATS misread this? Where is the evidence weak? Which requirement sounds satisfied but isn't? What line would a recruiter not understand in five seconds?
ERIK: Don't ask AI, is this good. That's how you get a greeting card. Ask where it breaks.
JOSH: That is immediately useful.
ERIK: Yep. And if you're building a hiring tool, do the same thing at system level. Run adversarial resumes. Run formatting tests. Run replay tests. Build dashboards for drift. If the score changes, you should know why.
JOSH: [pause]
JOSH: Next story. GLM 5.2 beat Claude Code in an IDOR benchmark. For normal people, IDOR is what?
ERIK: Insecure direct object reference. Fancy name for, can user A access user B's stuff by changing an ID in the request.
ERIK: Account slash 123 becomes account slash 124, and suddenly you're looking at somebody else's invoice. It's quiet bad. Not movie hacker bad. Worse. Boring lawsuit bad.
JOSH: And GLM did better than Claude Code?
ERIK: In that benchmark, yes. That's interesting. It does not mean everybody should throw out Claude. It means open models are getting strong in narrow security tasks, and that changes how teams should think about toolchains.
JOSH: What's the catch?
ERIK: The harness. Always the harness. A model inside a weak workflow loses to a smaller model inside a good workflow.
ERIK: Security scanning is not just reading code. You need repo context, route maps, auth boundaries, test data, request examples, and a way to confirm the finding. Otherwise the model is guessing with confidence.
JOSH: So the wrapper matters.
ERIK: The wrapper is the product. Semgrep does well because it has structure. Rules. Code paths. Known patterns. It doesn't just vibe-read the repo and say, this feels insecure.
ERIK: LLMs are useful when they sit next to deterministic tools. Give me Semgrep, unit tests, dependency scanning, runtime traces, and then let Claude or GLM reason over the weird parts.
JOSH: Where would you use GLM after seeing that?
ERIK: I'd put it in a second-opinion lane. Claude writes or reviews. GLM checks security-sensitive patterns. A static analyzer catches known issues. Then a policy gate decides what happens next.
ERIK: That's basically how I think about PrimeRouter. Don't marry one model. Route work by task. Some requests need a big reasoning model. Some need cheap fast classification. Some need a provider fallback because the first one is having a bad day.
JOSH: That's not as fun as model fan clubs.
ERIK: Model fan clubs are expensive. Also boring. Pick the tool that wins the task. Then measure it.
ERIK: For IDOR specifically, I want the model asking ugly questions. Where is object ownership checked? Is the tenant ID enforced server side? Can the client supply user ID? Are list endpoints filtered by auth context?
ERIK: Those are boring questions. They save companies.
JOSH: How does this compare to network automation?
ERIK: Same pattern. In networking, you don't trust a config because it looks right. You validate intent. You check state. You compare pre and post. You make sure the route exists, the policy hit, the device accepted the config, and traffic still flows.
ERIK: Code security needs the same discipline. Don't stop at model says vulnerable. Prove the path. Reproduce it. Then fix it and add a test so it doesn't come back.
JOSH: What should a builder do today?
ERIK: Add a security critic step to your AI coding workflow. Not a vague one. A specific one.
ERIK: After Claude writes code, ask another pass to inspect only authorization boundaries, user-controlled IDs, file paths, redirects, and secrets. Tell it to ignore style. Tell it to produce test cases.
ERIK: Then run the tests. If the model can't produce a failing test, treat the finding as suspicious until proven.
JOSH: That's a good distinction.
ERIK: Huge distinction. Models are great at sounding certain. Production cares about proof.
JOSH: [pause]
JOSH: Third deep one. Age verification. This feels like a policy story, but you're treating it like an infrastructure story.
ERIK: Because it is. Policy becomes architecture. Architecture becomes defaults. Defaults become hard to remove.
ERIK: If a site has to verify age, it needs some signal. ID scan, face estimate, credit card, mobile carrier check, government token, third-party identity provider. Pick your poison.
JOSH: That sounds messy.
ERIK: It is messy. And every option creates a database, a trust relationship, or a tracking surface. Sometimes all three.
ERIK: The goal sounds simple: keep kids away from certain content. Fine. But the implementation can turn into identity middleware for the whole internet.
JOSH: That's the slippery part.
ERIK: Yep. Once the gate exists, people will ask for more gates. Age today. Location tomorrow. Real name next year. Political content after that. Nobody has to be cartoon evil for the system to drift.
JOSH: How should engineers think about it?
ERIK: Minimize data. Prove only what you need. If the question is, is this user over a threshold, don't store their birthday, address, ID photo, and shoe size.
ERIK: Use privacy-preserving tokens where possible. Short-lived proofs. Local checks. Clear deletion. No giant pile of documents sitting in a vendor bucket named production-final-really-final.
JOSH: That's painfully believable.
ERIK: Because it happens. The data you collect becomes your liability. Every field is a future incident report.
ERIK: This is where builders need to push back early. Not in a dramatic way. In a design review way. What do we store? Why? For how long? Who can query it? What happens if the vendor disappears?
JOSH: How does that connect to automation?
ERIK: Automation makes bad policy fast. If you wire an identity check into every request, your system can deny millions of actions consistently. That's great if the rule is right. Brutal if the rule is wrong.
ERIK: Same reason I care about PrimeBus guardrails. An agent that can act across systems needs boundaries. Not vibes. Boundaries.
JOSH: So the theme today is trust, but verified.
ERIK: Exactly. Hiring scores, security scans, age checks. Same core issue. Automated decisions need evidence, replay, and limits.
ERIK: If you can't explain it, replay it, or stop it, you shouldn't put it in the critical path.
JOSH: That's the episode title.
ERIK: Probably too useful. We'll workshop it for six weeks and pick something worse.
JOSH: [pause]
ERIK: This episode is sponsored by Prime Automation Solutions. If you're still doing it manually, we automate it. Also, special on a website — $250. primeautomationsolutions.com
JOSH: [pause]
JOSH: Alright, what's the AI pro tip today?
ERIK: Add a critic agent to every important AI workflow.
ERIK: Not a second chat where you ask, thoughts? Build a real critic step with a narrow job.
ERIK: Example: Claude writes the change. Critic agent reviews only risk. It checks auth, data loss, retries, secrets, migrations, and tests. It is not allowed to compliment the code. It is not allowed to rewrite everything. It has to produce blockers and non-blockers.
JOSH: Why narrow it that much?
ERIK: Because broad critics become annoying interns with a thesaurus. Narrow critics catch real issues.
ERIK: Use this prompt: review this output for failure modes only. Group findings as blocker, warning, or nit. For every blocker, include the exact evidence and a test that would catch it. If there is no evidence, say no finding.
ERIK: Then make your automation treat blockers differently. A blocker stops the run. A warning routes to a human or gets logged. A nit gets ignored unless you're doing cleanup.
ERIK: That's how you get useful AI review instead of AI theater. That's your tip. Use it.
JOSH: [pause]
ERIK: That tip is straight out of The Autonomous Engineer — my book on building systems that run themselves. Grab it on Amazon.
JOSH: [pause]
JOSH: One more thing — we started a Discord for builders. If you're shipping AI, automation, or anything that makes a human obsolete — come hang out. Link at buildorbereplaced.dev.
ERIK: Post what you built. We'll post what we're building. Real wins, real builds, no fluff.
JOSH: [pause]
ERIK: Build or be replaced.
JOSH: If you want these signals in your inbox every morning, scanbrief.dev. See you tomorrow.