← Writing

We let candidates use AI in interviews. It made evaluation harder, not easier.

Banning AI in a technical interview measures a skill nobody on our team uses anymore. Allowing it broke most of what we used to look for — here's what we look at now.

Our engineers spend a large part of their day directing coding agents. They decompose work, write specifications, review output, and run several tasks in parallel. Typing code from memory is a smaller and smaller fraction of the job.

So a closed-book interview measures a skill our team doesn't use. We stopped running them.

Candidates use whatever they use day to day. That decision was easy. What followed was not, because allowing AI quietly invalidates most of the signals a technical interview traditionally produces.

What stopped working

Whether they can produce working code. Everyone can now. A competent candidate with a competent agent will get to something that runs, on most reasonably scoped problems, in the time available. "It works" went from a discriminating signal to table stakes.

Speed. Related, and worse: speed now measures the tool more than the person. Somebody who has learned one agent well looks faster than somebody equally strong who hasn't. That's real information about tool fluency, but it isn't the information most interviewers think they're collecting.

The resume. This one broke first and hardest. Any resume can be rewritten to look AI-heavy in about ten minutes, and a lot of them have been. Reading them now, "built agentic workflows" is closer to a formatting choice than a claim. We treat resumes as a baseline filter and nothing more; what we actually read is the work.

What we look at instead

Where they stop and think. The most informative moments in these interviews are the pauses. A strong candidate stops before accepting a plan and asks whether it's the right shape. A weaker one takes the first plausible output and starts building on it. Same tools, opposite instincts — and this distinction is far more visible with AI in the room than it was without it.

What they do when the agent is confidently wrong. It will be, at some point, and how a candidate handles that is the single best predictor we've found. Do they notice? Do they debug the reasoning or just re-roll and hope? Can they tell the difference between "the model misunderstood me" and "I misunderstood the problem"?

How they decompose. Give someone an ambiguous task and watch how they break it up. This is now the core engineering skill: the quality of the pieces you hand to an agent determines the quality of what comes back, and the ability to see the seams in a problem doesn't come from any tool.

Whether they can defend the output. We ask candidates to explain a decision inside code they just produced with help. Not as a gotcha — because reviewing generated work is most of the job. Someone who can't explain what they shipped can't be trusted to review a fleet of agents shipping on their behalf.

Whether they ask about the user. Our engineers get deployed inside customer companies. Somebody who never asks who this is for, how often the problem happens, or what the person does today is going to build the wrong thing beautifully.

What they've built that nobody asked for. The strongest correlation in our hiring so far, imperfect and small-sample as it is: people who build things outside their assignments. Side tools, automations for their own annoyances, half-finished experiments. It's evidence of the trait we can't teach, which is pulling on a problem without being told to.

The uncomfortable part

Allowing AI made interviews harder to score, not easier. The old rubrics collapsed into pass/pass. What replaced them is more judgment-heavy, which means more room for interviewer bias, which means we have to be deliberate about structure: the same problems, the same probes, notes taken against specific behaviors instead of a general impression.

We're also aware of the trap we're setting for ourselves. It's easy to over-index on tool fluency and hire people who are excellent at driving today's agents and mediocre at engineering. The tools will turn over. Judgment, decomposition, and the instinct to check things won't.

So the bar we're actually trying to measure is this: can this person own an outcome in an ambiguous situation, with agents as a force multiplier rather than a crutch?

That's the job. It's what a forward-deployed engineer does inside a customer's company on week one, usually without a spec, frequently while discovering that the stated problem isn't the real one.

If that description sounds like your week rather than your fear, we're hiring.

Get started

Let's find the workflow worth fixing.

Bring one process that annoys everyone. We'll tell you honestly whether AI should touch it.

Book a call