Skip to content

Latest commit

 

History

History
241 lines (134 loc) · 26.8 KB

File metadata and controls

241 lines (134 loc) · 26.8 KB

Behavioural and STAR stories

Your existing STAR stories will not survive an AI leadership loop. Not because they are badly told, but because they were built for a different shape of work.

Classic STAR rewards a triumph. Situation, Task, Action, Result, and the Result is a win: latency down, revenue up, migration shipped, team scaled. That structure works when the work is deterministic. You knew roughly what building the thing would cost, you built it, it worked.

AI work does not produce that shape. The honest result of a serious AI project is frequently: we ran thirty experiments, twenty-eight failed, and we killed the programme at week six. Told inside a format designed for triumphs, that sounds like a career low point. So candidates do one of two things, and both lose the round. They either dress it up (inventing a win that the follow-up questions then dismantle) or they retreat into vagueness: "it was challenging, but we learned a lot, and the team grew from it."

An interviewer hearing "it was challenging but we learned a lot" hears nothing. Literally nothing. There is no information in that sentence to grade.

The fix is not better storytelling. The fix is to make the measurement the result. Compare the two sentences below and notice that they describe the same failed project.

"We tried an AI approach for document extraction. It was harder than expected. We eventually decided to move on, but the team learned a huge amount about the technology and we applied that later."

"We killed it at week six. The score plateaued at 71% exact-match against a bar of 90%, and the 90% bar and the six-week checkpoint were both agreed with the sponsor before we wrote any code. Killing it took one meeting because the decision rule was already signed."

The second candidate is a strong leader and the interviewer knows it inside twenty seconds. They set a bar before they had a result. They measured against it. They acted on the measurement without negotiating with themselves. That is the entire competency the round is testing, and the project failing is what let them demonstrate it.

This is the reframe that makes AI behavioural rounds tractable. In deterministic engineering, the result is the outcome. In AI engineering, the result is the quality of the decision you made given the evidence you had. A killed project with a clean decision rule beats a shipped project that nobody measured.


What an AI leadership story must contain

Four beats. If a story is missing any one of them, it will break under follow-up.

A number. What you measured and what it said. Not "accuracy improved" but "exact-match went from 71% to 88% on a 50-document golden set". The number does not have to be impressive. It has to be real, and you have to know where it came from, who built the set it was measured on, and what its weaknesses were. Interviewers do not probe numbers to check the number. They probe to find out whether you were close enough to the work to have one.

A decision. What you did with the evidence. The evidence is not the point; the action you took because of it is. The strongest version includes one decision you would now make differently, offered by you rather than extracted from you. "I would set the checkpoint at week three, not week six. Weeks four to six produced no new information." That single sentence does more for you than any success metric, because it proves you are still learning from the story rather than performing it.

A trade-off you actually made. Something real that you gave up. Latency for quality, cost for coverage, three weeks of eval build for a slower start, scope for a shippable date. If your story contains no sacrifice, the interviewer concludes either that the problem was easy or that you are editing.

The people. This is a management interview. If you are the hero of your own story, you have failed the round.

That last one is the most common failure in AI loops, and it is a strange one, because it usually happens to good candidates. Strong technical EMs tell technically excellent stories in which they personally did everything: "I built the eval harness, I found the retrieval bug, I ran the cost analysis, I made the call." Every sentence starts with "I". The story is true. The candidate did do those things. And the interviewer scores them as an individual contributor with a title, because nothing in twelve minutes of narration required a manager.

Fix it by naming people and being specific about what they contributed and what you did that only a manager could do. "Priya built the golden set with a claims adjuster called Martin who gave us two afternoons a week. My job was getting Martin's two afternoons out of an operations director who did not report to me, and protecting Priya's time from three inbound requests while she built it." That is a manager. The engineer built the artefact; you built the conditions in which it could exist.


The twelve prompts

Twelve prompts covering the competencies that AI leadership loops actually assess. The right-hand column is the load-bearing one: if you cannot fill it in, you do not have a story yet.

Every quoted line in the sections below is illustrative phrasing, invented to show the shape of an answer rather than reported from a real project. The figures in them carry no weight of their own; only your figures do.

# The prompt Competency it maps to The number your story needs
1 A time you killed or de-scoped an AI project Decision-making under uncertainty The score against the pre-agreed bar
2 A time an AI system you owned failed in production Ownership and incident leadership The failure rate and the blast radius
3 A time you disagreed with a stakeholder about what AI could do Influence and technical credibility What the evidence said
4 A time you changed your mind because of data Intellectual honesty The eval result that moved you
5 A time you set expectations with an executive who had seen a demo Stakeholder management The gap between demo and distribution
6 A time you built or fixed an evaluation process Engineering rigour Score before and after, or bugs it caught
7 A time you grew someone into AI work Talent development Time to independent contribution
8 A time you made a hiring decision that was hard, or wrong Hiring judgement What you missed, and how the loop changed
9 A time cost forced an engineering decision Commercial judgement Before and after unit cost
10 A time you said no to an AI initiative Strategic judgement What you did instead, and what it returned
11 A time you handled a safety, privacy, or fairness concern Risk judgement The control you built and what it caught
12 A time you led a team through uncertainty or a shift in what was possible People leadership under ambiguity Retention, delivery, or morale evidence

1. Killing or de-scoping an AI project

What they are testing: whether you can stop. Most organisations cannot, because sunk cost and executive sponsorship compound. They want evidence that you set an exit condition before you were emotionally invested in the answer.

The shape of a strong story: the bar and the checkpoint existed before the work started, agreed in writing with a named sponsor. You measured against it. The number did not clear the bar. You stopped, and stopping was quick because the argument had been had months earlier.

The trap: telling a story where you killed the project after nine months because it "clearly was not working". That is not judgement, it is exhaustion. The interviewer will ask when you first suspected, and the gap between that date and the kill date is your actual score.

2. An AI system failing in production

What they are testing: whether you treat probabilistic failure as an engineering problem rather than an act of nature. AI systems fail differently from services. They do not go down; they go subtly, confidently wrong, and nobody pages you.

The shape of a strong story: you can state the rate (what fraction of requests, over what window) and the blast radius (which users, what downstream harm, whether money or trust moved). You describe how you found out, and honestly whether a customer found out first. Then the control you added so the next one is caught by a monitor rather than a complaint.

The trap: describing a classic outage (a timeout, a bad deploy) that happened to involve a model. That is a services incident. They want the failure mode that only AI has.

3. Disagreeing with a stakeholder about what AI could do

What they are testing: technical credibility used as influence, not as a weapon. Can you be right without being insufferable, and can you make the disagreement cost less than being wrong would have?

The shape of a strong story: you did not win by arguing. You won by proposing a cheap test. "We disagreed about whether it could handle scanned documents, so we spent two days on forty scanned samples rather than two months on the argument." The evidence settled it, and the story is better if the evidence partly settled it against you.

The trap: the story where you were right, they were wrong, you told them, and they eventually saw sense. Nobody in the room believes it went like that, and nobody wants to work for the person telling it.

4. Changing your mind because of data

What they are testing: intellectual honesty, and whether you have any evidence-gathering habit at all. Someone who has never changed their mind either does not measure or does not notice.

The shape of a strong story: name the specific result. "I was sure fine-tuning was the answer. The eval showed the retrieval baseline within two points of the fine-tuned model at a fraction of the operating cost, so we shipped retrieval and I dropped it." Be clear what you believed, why it was reasonable to believe, and what specifically broke it.

The trap: a fake conversion where you changed your mind about something you never cared about. Pick a belief you held publicly and had staked something on. The cost of being wrong is what makes the story worth points.

5. Setting expectations after an executive sees a demo

What they are testing: whether you can manage the single most reliable source of AI project failure, which is a demo that worked on eight curated inputs meeting a budget conversation.

The shape of a strong story: you named the gap in terms the executive cared about. A demo is a sample of size eight from the head of the distribution. Production is the tail. You quantified it: "The demo cases were the 20% we handle well. Here are ten real tickets from last Tuesday; it gets four." Then you offered a scoped version that was real, rather than a lecture about why they were naive.

The trap: being the person who said no to the demo. You want to be the person who turned the demo into a plan with a date on it.

6. Building or fixing an evaluation process

What they are testing: engineering rigour, and whether you understand that the eval is the product's actual test suite. Many candidates have never built one and it shows within a minute.

The shape of a strong story: who built the golden set, how large, how the labels were agreed, what the inter-annotator disagreement taught you. Then the payoff: the score before and after, or a specific regression the harness caught before a customer did. "It caught a prompt change that quietly dropped date extraction by nine points."

The trap: describing eval as a dashboard. A dashboard is not an eval. If nothing was ever blocked from shipping by the thing you built, it did not exist.

7. Growing someone into AI work

What they are testing: whether you can build capability rather than buy it, because the market for AI engineers is unpleasant and every good EM needs an internal path. Related reading: Upskilling your existing team.

The shape of a strong story: a named person, their starting point, the specific first task you chose and why it was the right size, and the milestone. Time to independent contribution is the number: "Eleven weeks from first eval ticket to owning the retrieval quality workstream without me reviewing her decisions."

The trap: "I sent them on a course and they picked it up." That is not development, that is a budget line. What did you personally change about their work so that learning happened?

8. A hard or wrong hiring decision

What they are testing: hiring judgement, and specifically whether you learn from your own misses. Every EM has hired someone who did not work out. Candidates who claim otherwise have either not hired enough or are not being straight.

The shape of a strong story: what you missed, why your loop let you miss it, and the concrete change you made. "Our loop tested model knowledge and never tested whether someone could sit with an ambiguous eval result for a week. We added an exercise where the data is genuinely inconclusive." See Hiring guide.

The trap: blaming the hire. The moment you describe their failings rather than your process failings, the round is over.

9. Cost forcing an engineering decision

What they are testing: commercial judgement. AI features have a marginal cost per request that traditional software does not, and a lot of EMs have never had to think in unit economics.

The shape of a strong story: a before and after unit cost with the lever named. Caching, a smaller model for the easy 70% of traffic, shorter context, batching, moving a step out of the model entirely. "Cost per resolved ticket dropped by roughly two thirds when we routed the classification step to a small model and kept the large one for the escalation path." Say what you gave up. Something always got slightly worse.

The trap: talking about cost only as a percentage saving. Interviewers want to know whether you know the unit and whether it was ever compared against the value of the transaction. See Cost management.

10. Saying no to an AI initiative

What they are testing: strategic judgement, and whether you can defend the opportunity cost rather than just being conservative.

The shape of a strong story: the no is only half of it. The other half is what you did with the capacity instead and what that returned. "I declined the conversational interface and put those two engineers on search relevance. Search touched every user; the assistant would have touched the 3% who open the help menu." A no without an alternative is caution, not strategy.

The trap: a no motivated by discomfort with AI. If your reason was that the technology was unproven, you have described a preference. Give the interviewer the arithmetic.

11. A safety, privacy, or fairness concern

What they are testing: risk judgement in the specific way AI creates risk: not through a breach, but through a system doing exactly what it was built to do, at scale, to a group you did not think about.

The shape of a strong story: you found or anticipated something, you built a specific control, and the control caught something real. "We added a rule that any output containing an account number went to human review before send. It caught eleven cases in the first month, two of which were the wrong customer's number." Concrete controls, concrete catches. See Guardrails and safety and Responsible AI in practice.

The trap: describing a policy you wrote. Policies are not controls. What ran in the pipeline, and what did it stop?

12. Leading a team through uncertainty

What they are testing: people leadership when the ground moves. Between a capability shift and a strategy change, AI teams get their assumptions invalidated regularly, and morale is the first casualty.

The shape of a strong story: what specifically became uncertain, what you told the team and when, and what you deliberately held stable so they had something to stand on. The evidence is retention, delivery through the period, or something you can point to that shows the team was intact on the other side.

The trap: describing your own feelings about the uncertainty. The interviewer is buying the effect you had on other people, not your resilience.


The fully worked example

This is an illustrative example. The numbers, names, and company are invented to show the shape. Replace all of it with your own material; a borrowed story dies on the second follow-up.

Situation. I ran a team of nine in the claims side of an insurance business. We were spending a lot of manual effort on document intake. Adjusters were reading scanned loss reports and typing eleven fields into the claims system. It took about six minutes a document at a volume that justified two full-time roles. Our COO had seen a vendor demo of document extraction and wanted it live in the quarter.

Task. I owned the decision on whether we built it, and the delivery if we did. What I actually decided first was that we would not decide yet. I proposed a six-week POC with an exit condition agreed before we started: 90% exact-match on all eleven fields across a golden set, or we stop. I wrote that down, and the COO and I both signed off on the number and the date. That conversation was uncomfortable, because agreeing a kill condition in advance feels like planning to fail. I framed it as buying an option: six weeks to find out, and a cheap way out.

Action. The first thing we built was not the extractor. Priya, one of my engineers, spent the first ten days building a 50-document golden set with Martin, a senior adjuster. I did not build it; my job was getting Martin's two afternoons a week out of an operations director who did not report to me, and then defending Priya's time when three other requests landed on her. The golden set was the best thing we produced. Martin and Priya disagreed on the labels for six documents, and working through those disagreements told us something important: the field the COO cared most about, incident date, was genuinely ambiguous in about one document in eight, because handwritten reports often carry two dates and no rule about which one counts. Then we iterated. Baseline prompting, then better prompting, then a layout-aware pre-processing step, then retrieval against past claims for the difficult fields. Roughly thirty experiments over four weeks.

Result. We plateaued at 71% exact-match across all eleven fields. Individual fields varied a lot: policy number was near-perfect, incident date sat in the fifties. The last two weeks moved the number by under a point. I also ran the honest cost comparison, which is the part most people skip. At 71%, an adjuster still has to check every document, because they cannot know which 29% is wrong. Verifying takes about four minutes against six to type it fresh. That saves two minutes of every six, a third of the effort, so against a baseline of two roles it is two thirds of one role, against a build and run cost that was higher than the saving. The technology was not the problem. The economics were.

So I took it to the COO at week six. That meeting lasted about fifteen minutes, and it was short precisely because we had agreed the bar in advance. There was nothing to argue about. I did not present it as a failure; I presented the 71%, the field-level breakdown, and one specific thing we had learned that changed the shape of the product: policy number and claim type were above 95%, and those two fields were the ones that determined routing. So I proposed a different thing. Not extraction, but triage. Use the model for the two fields it was excellent at, route the document to the right adjuster queue automatically, and leave the eleven-field typing alone.

We shipped that in five weeks. It cut routing time and, more importantly, cut misrouted claims, which had been quietly costing us a day of turnaround each time. The golden set outlived the project entirely: it became the regression suite for the triage system and was still in the pipeline two years later. Martin ended up as the person other teams went to when they needed labels.

What I would do differently: the checkpoint should have been week three. Weeks four to six produced under a point of movement and cost us a month.

Why this works

  • The decision rule existed before the work. 90% and six weeks, agreed and signed by a named sponsor. Everything else in the story is downstream of that one move.
  • There is a number, and it has texture. Not just 71%, but the field-level breakdown, which is what made the reframe possible. Knowing the aggregate is table stakes; knowing the distribution is what proves you were close to the work.
  • Named collaborators doing the actual work. Priya built the golden set. Martin knew the domain. The manager's contribution is explicitly the thing only a manager could do: securing Martin's time across an org boundary and protecting Priya's focus. It is not a hero story.
  • The honest cost comparison. Six minutes to type against four minutes to verify, and the recognition that partial accuracy does not save the review step. This is the move that separates commercial judgement from technical enthusiasm.
  • A reframe, not a retreat. The story does not end at the kill. It ends at a shipped product that came out of what the failed experiment taught. The failure produced information, and the information had a use.
  • An artefact that outlived the project. The golden set became the regression suite. This is the tell that the work was real engineering rather than a demo.
  • An unprompted "would do differently" with a number attached. Week three, not week six. Offered, not extracted.

How to break it

Remove the number and it becomes "we tried it, it did not work well enough, we did something else". Ungradeable.

Remove the pre-agreed rule and the kill becomes an opinion. The follow-up is "how did you know 71% was not good enough?" and without a bar set in advance, the honest answer is that you decided afterwards, which is exactly what the round is checking for.

Make yourself the sole hero (I built the golden set, I found the plateau, I made the call) and it is a competent senior engineer's story. The manager grade requires other people to have done things.

Blame the COO for the demo, or the vendor for overselling, and you lose regardless of how right you are. The version where the executive is naive and you are the adult is the single most common way this story fails.


Building your six

You do not need twelve stories. You need about six, because a well-chosen story answers three prompts depending on which part you emphasise. The worked example above covers prompt 1 (killing), prompt 5 (the COO and the demo), prompt 9 (the cost comparison), and prompt 6 (the golden set) with a shift in emphasis and about forty seconds of re-pointing.

Your story Primary Also covers
A project you killed with a pre-agreed rule 1 Decision under uncertainty 5 Executive expectations, 9 Cost
An incident in a probabilistic system 2 Incident leadership 11 Risk, 6 Evaluation
A belief the evidence overturned 4 Intellectual honesty 3 Influence, 6 Evaluation
Someone you grew into the work 7 Talent development 12 Leading through ambiguity, 8 Hiring
A no you defended with arithmetic 10 Strategic judgement 9 Cost, 3 Influence
A hire that went wrong and the loop change 8 Hiring judgement 7 Development, 4 Honesty

Six stories, twelve prompts, full coverage. Note that each story appears in three cells and each prompt is hit at least once. Build your own version of this grid before you build the stories; the gaps in the grid tell you which stories you are missing.

Then the technique that matters most: write the number first. Open a document, write one line per story, and that line is the measurement. "71% against a 90% bar." "Eleven catches in month one, two of them the wrong customer." "Eleven weeks to independent ownership." If you cannot write the line, you do not have the story, and no amount of narrative craft will rescue it.

Only once the number is on the page do you build the story around it. This ordering matters more than it sounds. A story constructed for its narrative and then retrofitted with a number is the story that collapses under follow-up, because the number was chosen to fit the ending rather than the ending being driven by the number. Interviewers cannot always tell a good story from a manufactured one, but they can always tell when the second probe produces hesitation.


Follow-ups that break weak stories

Five probes. They are not clever, they are just load-bearing. Every one of them tests whether the story is remembered or reconstructed.

"What was the number?" Works because invented stories have adjectives where numbers should be. Survive it by knowing the figure, the denominator, and who produced it. "71% exact-match across eleven fields on 50 documents, measured by the harness Priya built."

"Who else was involved?" Works because hero stories have no other characters, and a candidate who has not thought about their team cannot produce names under time pressure. Survive it by having the names ready and knowing precisely what each person contributed that you did not.

"What would you do differently?" Works because the polished story has been told so often that the rough edges are sanded off, and a candidate with no regrets has stopped learning. Survive it by volunteering one before you are asked, and make it specific and slightly costly to admit.

"What did the person who disagreed say?" The sharpest one. Works because it requires you to have actually listened to opposition well enough to reconstruct its logic. Candidates who steamrollered dissent cannot answer; they only remember that someone objected. Survive it by being able to state the opposing case fairly and say which part of it was right.

"What happened after you left?" Works because it separates results from theatre. Things that were real persist. Things that were a performance for a promotion cycle quietly get switched off. Survive it by knowing, and by being honest if the answer is that it was rolled back. "The triage system is still running. The weekly quality review I set up stopped within two months of my leaving, which tells me I built it around myself rather than into the process."

Practise these against your own stories out loud before the loop. Written stories always sound better than spoken ones. See Mock interview scripts for a structure to rehearse against, and note that formats vary between companies and change over time.


Next: Case and strategy questions

Related: Common failure modes | Top 20 questions | Mock interview scripts