A normal status update reports progress against a plan. An AI update reports evidence against a question, because the plan is a set of bets and the honest unit of progress is what you now know that you did not know last week. This template forces four things onto the page every week, and they are the four an executive actually needs: eval movement, cost trend, top risks, and decisions needed. It is deliberately short, because a long update does not get read, and an update that is not read is worse than no update at all (it costs you the writing time and buys you the false belief that you communicated).
Last reviewed: July 2026
Copy the block below. Keep the headings. Delete the italic guidance once you have internalised it, not before.
Initiative: [name] | Week ending: [date] | Owner: [you] | Status: On track for the decision date / At risk / Decision needed
Guidance: the status is about the decision date, not the outcome. "On track" means you will know the answer when you said you would know it. It does not mean the answer will be yes. The outcome is what you are finding out, so claiming to be on track for it is either a guess or a lie. This distinction is the single most useful thing in this template and it will take your stakeholders about three weeks to absorb.
| Decision | Who decides | By when | What happens if it slips |
|---|---|---|---|
| [the actual choice, phrased so a yes or no is possible] | [one named person, not a group] | [date] | [the concrete cost of a week of delay] |
Guidance: this is why the update exists, so it goes near the top, not the bottom. An update with no asks is a broadcast, and broadcasts get filed. If you have no decision to request this week, write "None this week" and mean it, but check first: most weeks in AI work contain a fork that someone above you should own and that you are quietly deciding by default.
Metric: [name] | This week: [number] | Last week: [number] | Baseline: [number] | Target: [number]
What moved it: [one or two lines]
Guidance: one metric, not five. Pick the one that decides the initiative and hold it steady for the life of the project. If it did not move, say so, and say what that eliminated. A week that produces knowledge is a real week. Hiding a flat week teaches everyone on the team that failed experiments are shameful, and a team that believes that will start running only the experiments they expect to win, which is the end of learning anything. See Evaluation-driven development for how the metric gets built in the first place.
Cost per [unit of value]: [number] ([up/down] from [number]) | Spend to date: [number] against [budget] | Direction: [improving / flat / worsening]
Guidance: this lives next to the eval number, always, never on a separate page and never in a different meeting. Separate them and someone optimises one while destroying the other, usually by discovering that a much larger model fixes the quality problem for eight times the money, or that an aggressive cache fixes the cost problem by serving stale answers. Both numbers in one glance is the whole point. See Cost management.
[2-3 lines. Plain sentences.]
Guidance: the most valuable section on the page and the first one people delete when they are busy. In AI work this is the actual output of many weeks. Six months from now, nobody will reread your task list, but somebody will absolutely ask "did we ever try X?" and this section is the only place that answer lives. Write it for that person.
| Risk | Impact | What we are doing | By when |
|---|---|---|---|
| [specific, not "model quality"] | [what it costs if it lands] | [an action with an owner] | [date] |
Guidance: three maximum. Real ones. If a risk has been unchanged for a month, it is not a risk, it is a condition of the work, so either promote it into a decision or drop it off the page. A risk table that never changes is furniture, and people stop reading furniture.
[2-3 lines on the question you are answering next, not the tasks you are doing.]
Guidance: "Run the retrieval sweep" is a task. "Whether retrieval quality or prompt structure is the binding constraint on accuracy" is a question. The question tells the reader what they will learn on Friday. The task tells them nothing they can act on.
[One line and a link.]
Guidance: the demo shows what is possible. The eval number shows what is probable. Both belong in the same email, because a demo on its own has taken more AI programmes off the rails than any technical failure. Someone senior sees a good demo, forms a belief about reliability, and starts making commitments against it.
Under 300 words. If you cannot say it in 300, you do not yet know what happened.
The same metrics every week, so the trend is legible. No new metric introduced quietly in a good week, which is the most common small dishonesty in this genre and the easiest to spot from above once someone is looking. If the metric genuinely needs to change, that is a decision, so put it in the decisions table and let someone else agree to it.
No adjective where a number exists. "Significantly better" is a number you have and chose not to type. "Roughly stable" is a number you have and are hiding behind.
Send it when the news is bad. Especially then. The credibility you spend on a hard week is exactly what makes the good weeks believed. Skip two bad weeks and your next green update reads as noise, because your readers have correctly learned that your updates only appear when things are going well.
Illustrative only. The initiative, numbers, and names are invented to show the shape.
Initiative: Support triage assistant | Week ending: 10 July 2026 | Owner: R. Okafor | Status: On track for the decision date
Decisions needed
| Decision | Who decides | By when | What happens if it slips |
|---|---|---|---|
| Approve annotation budget for 400 additional labelled tickets from the refunds queue | D. Ellis (sponsor) | Fri 17 July | Annotation vendor slot is reallocated and we lose two weeks. The go/no-go on 14 August moves to 28 August. |
Eval movement
Metric: routing accuracy on the held-out set of 600 tickets. This week: 78%. Last week: 78%. Baseline: 71%. Target: 85%.
What moved it: nothing, and that is the finding. We tested the hypothesis that accuracy was limited by prompt structure. Three restructured prompts (few-shot with hard negatives, explicit category definitions, chain-of-thought before the label) all landed within one point of the current prompt. Prompt structure is not the binding constraint. That eliminates the cheapest branch and leaves retrieval quality and label noise, which is where next week goes.
Cost trend
Cost per resolved ticket: down from 4.1p to 2.6p after moving classification to the smaller model, which the flat eval result made safe to do. Spend to date: 11,400 against a 40,000 pilot budget. Direction: improving.
What we learned
Prompt engineering is exhausted on this task at this quality level. The remaining accuracy gap is in the data, not the instructions. Separately, the smaller model matches the larger one on this classification step, which is why cost improved in a week with no quality gain.
Top risks
| Risk | Impact | What we are doing | By when |
|---|---|---|---|
| Refunds category is 31% of volume and our worst performer (54% accuracy). Suspected label noise in training examples. | Blocks the 85% target regardless of model work. | Re-annotating 200 refunds tickets with two annotators and measuring agreement. Budget ask above covers the rest. | 24 July |
| Held-out set was drawn before the April policy change and may no longer represent live traffic. | We could hit target on a stale benchmark and miss in production. | Sampling 150 live tickets this week to compare distribution. | 18 July |
Next week
Whether the refunds gap is label noise or genuine task difficulty. Inter-annotator agreement on the re-annotated sample answers this. If agreement is low, the problem is our labels and is fixable. If agreement is high, the model is failing a task humans find easy, and that is a harder and more interesting problem.
Demo: Staging link. It routes the happy path cleanly. It is wrong about one ticket in five, and refunds tickets are where it is wrong.
Your team's version carries the eval detail in full: every prompt variant, the exact deltas, the failed sweep, the annotation disagreements. That detail is the working material of the next experiment and stripping it out wastes it. Your executive's version carries the decision, the trend, and the ask, because that is the resolution at which they can act, and burying the annotation budget request under four paragraphs of retrieval methodology is how the request goes unanswered. Same facts, different resolution, and never a different story. The moment the team's version and the executive's version disagree about what is true, you have stopped writing status updates and started running two narratives, and the reconciliation always arrives at the worst possible moment. For the wider craft of pitching this well, see Stakeholder communication.
Next: Repository home
Related: Stakeholder communication | Metrics and KPIs | AI project management