New

The AI Operations Lead's First 90 Days

You have taken the job of getting AI into production inside an organisation. Here is what to negotiate before you start, and what the first three months should look like.

What's inside this template

Who it's for

The person joining as an AI Operations Lead, internal Forward Deployed Engineer, AI engineer or AI enablement lead, and anyone moving into AI delivery from a business or IT role internally

When to use

From the offer stage through to the end of your first quarter. The negotiation section only works before you sign, so read that part first.

Key benefit

A plan that gets you to a real, measured result by day 60 without skipping the mapping work that makes months four to twelve possible, and the reasoning to defend where you chose to concentrate

Sections included

  • What to negotiate before you sign, and why you only get one chance
  • Questions that tell you whether the job is set up to succeed
  • Days 1 to 30: map, and resist building
  • Days 31 to 60: ship one thing that someone asked for
  • Days 61 to 90: prove it and set up to scale
  • The three artefacts you have to build, and what good looks like
  • Getting a working operating model in Claude in your first week, free
  • Choosing domains on business impact before you choose use cases
  • Defining an evidenceable metric before any build starts
  • How to score a use case backlog so the ranking survives challenge
  • Directing demand when everyone wants to be first
  • Avoiding the auditor problem
  • What to do if you inherited a stalled programme
  • An installable Claude skill that works through all of it with you

Complete template content

Claude skill ai-operations-lead-first-90-days

This is an installable skill, not a document to fill in. Install it, tell Claude what you are working on, and it works through it with you and produces the output.

Claude desktop or Claude.ai

Download the .zip, then go to Settings, Capabilities, Skills, and upload it.

Claude Code

Unzip it into ~/.claude/skills/ for every project, or .claude/skills/ for one. Then run /ai-operations-lead-first-90-days.

NOTE: To use this guide, copy the content using the “Copy page” button above, then adjust the dates and functions to your organisation.

Your first 90 days as an AI Operations Lead

You have taken a job nobody at this organisation has done before. There is no predecessor to learn from, no established playbook inside the company, and a set of expectations nobody has written down.

The good news is that the bar for a first quarter is clear even when nobody states it: one thing in production that people use, and enough understanding of the organisation that the second and third things come faster than the first.

This guide covers what to agree before you start, and what the three months should look like.

There is also an installable Claude skill that works through it with you, above. Three modes: what to negotiate while you still have an offer, a 90-day plan built for your organisation, and help building the operating model and backlog once you are in the work.


Before you sign

You have leverage exactly once, and it is now. After you start, the same conversations become requests for favours.

Four things to agree, in writing, in the offer discussion.

A named sponsor. Someone senior enough to open a door you cannot open yourself. Get the name, and get a standing meeting in the diary before your first day. A role like this fails quietly when the sponsor is notional.

A budget you control. Platform costs and tooling, without a business case per agent. The amount matters less than the control. Needing approval for every fifty-pound tool turns a delivery job into an administrative one.

System access, agreed with security before day one. Ask what AI agents are permitted to reach today and who decides. If the answer is that nobody has decided, ask for that decision to be made before you start. This one item is the difference between a first quarter and a first half.

A written scope. Which functions are yours this year, and which are not. Roles like this are offered as “own AI across the organisation”, which is three jobs. Narrow it now and expand later from a position of having delivered.

Questions that tell you whether this is set up to succeed

Ask these in the final interview. You are assessing them.

  1. Who owns the AI outcome above me, and what have they committed to?
  2. What use cases do you already have in mind, and who asked for them?
  3. What is security’s current position on what AI agents may reach?
  4. What has been tried before, and what happened?
  5. What does success look like at my first board review, in your words?
  6. Who will be difficult, and why?

Question four matters most. An organisation with a failed Copilot rollout behind it is a different job from a greenfield one, and they rarely volunteer it.

Question six tells you whether they understand their own organisation. An interviewer who says “everyone is very supportive” has either not looked or is managing you.


Days 1 to 30: map, and resist building

The pressure to ship something in week one is real and you should resist it. Anything built in week one rests on assumptions you have not tested, and when it lands badly it costs you the credibility you need for the work that matters.

Week 1. Meet the sponsor and agree the plan below, so it is theirs as well as yours. Meet security and convert whatever was agreed before you started into actual access. Meet the people who will be difficult, first, before they hear about you from someone else.

Set up where the map is going to live before you start filling it. The Kowalah Plugin gives you a working operating model in Claude in about ten minutes and costs nothing to start, which is the shortest route from nothing to somewhere to put what you learn. There is a walkthrough of that first session, from install to one function mapped. A spreadsheet works too. What does not work is three weeks of discovery in a notebook.

Weeks 2 and 3. Discovery in your first function. Sit with people and map what they actually do, not what the process document says. Ask what they put off, what is tedious, what they do outside normal hours because there is no time in the day. Do not ask them how AI could help, that is your job and they will anchor on whatever they have read.

Week 4. Write up the operating model for that function and take it back to them. Getting it wrong in front of the team and being corrected is worth more than getting it approximately right in private, and it is the fastest way to establish that you are interested in how they work.

By day 30 you should have: access sorted, the first function mapped, your two or three domains chosen and agreed with the sponsor, a scored backlog, and a plan the sponsor has signed off.


Days 31 to 60: ship one thing

Pick the use case a team asked for. Not the most impressive one, and not the one you find most interesting. Your first delivery sets how every later approach is received. A team that asked for help will use what you build and tell other people, which is the only marketing that works internally.

Beyond that, the first one should have a named business owner, a process that runs often enough that a small improvement compounds, and a baseline you can measure before you change anything.

Where this sits against your domains. Choosing your first use case for social reasons and your programme for business reasons pulls in two directions. Be deliberate about which one you are serving. Ideally the team that asked sits inside one of the domains you picked, and you get both. If it does not, do it anyway: the first delivery is buying you permission, and permission is worth more in month two than an extra point of impact. What you cannot do is let that become the pattern. From the second use case on, the domains decide.

Capture the baseline first. Once the agent is live you cannot go back and measure what it was like before, and a result you cannot quantify is a result the board discounts. Time per cycle, error rate, volume handled, whatever the team already cares about. An afternoon spent on this in week five is worth more than anything else you do that week.

Ship it to real users and sit with them. Watch the first few runs in person. The gap between how you expected it to be used and how it is actually used is where the next three improvements are, and you will not find it from usage data alone.

By day 60 you should have: one agent in production with a named owner, a measured before and after, and a small number of people who will say publicly that it helped.


Days 61 to 90: prove it and set up to scale

Add the second and third. These should come faster than the first. If they do not, that is information: either the mapping was too shallow or the first one solved something unusual.

Start the champion network. Find the person in each team who was most engaged during discovery and give them a role. They do not need to build, they need to be the person their colleagues ask. This is what makes the fourth, fifth and sixth use cases arrive without you going to look for them.

Establish the agent register properly. Every deployment: owner, purpose, the systems it can reach, the data it handles, what it costs per day. Doing this at three agents feels unnecessary. Doing it at thirty is a project.

Write the quarter-end report. What is live, who uses it, what changed against the baseline, what is next, and what you need. Four agents in production used by 120 people weekly, with the month-end close down from nine days to five, is a board update. “Completed AI discovery across six functions” is activity, and a board told that twice stops asking.

By day 90 you should have: two or three agents in production across two functions, adoption measured, a champion in each, a register that is current, and a roadmap for the next two quarters that your sponsor has agreed.


The three artefacts

These outlast any individual agent, and they are what you are actually judged on at the end of the quarter. They are also what makes the programme survive you, which is the thing that gets you the budget for year two.

1. The operating model

A map of how the organisation actually works: the units, the processes each one owns, who is accountable, and where AI is landing today.

Build it function by function as you do discovery rather than attempting it all at once. Each unit needs its processes named, the volume and frequency of each, who owns it, and what the current pain is. The value is in the coverage, not the detail, so a shallow map of eight functions beats a deep map of one.

Keep it current. An operating model that is accurate in month one and stale by month six is worse than none, because people will rely on it.

2. The scored backlog

Pick the domains before you pick the use cases

The sequence that matters is domain first, use case second. A domain is a top-level business process: order to cash, claims handling, procurement, month-end close, bid production. Choose the two or three where AI moves a number the business already cares about, then find your use cases inside them.

This is the argument in Rewired (Lamarre, Smaje & Zemmel, 2nd ed., 2026), and it holds up in practice. Value concentrates in a small number of domains. Spreading thinly across every function produces a long list of small wins that never add to anything a CFO recognises, and it is the default outcome if you let the backlog be built by whoever asks.

To choose a domain, ask three questions of each candidate:

  1. What number does this move, and who owns that number? If you cannot name both, it is not a domain, it is a collection of tasks.
  2. How big is the gap between how it runs today and how it could run? Look for volume, manual handoffs and rework.
  3. Will the executive who owns it engage? A high-value domain with a disengaged owner delivers less than a mid-value domain with an active one.

Write down the domains you chose and, just as importantly, the ones you did not. The second list is what you will be asked about.

Define the metric before you build

For each domain, agree the metric and its baseline before any build starts. Not afterwards, and not “we will work out how to measure it once it is live”.

The measurement has to be evidenceable: a number someone already produces, or one you can start producing now, owned by the person who owns the domain. Cycle time on the close. Cost per claim. Days from bid request to submission. Error rate on invoice matching.

A programme that cannot evidence its result gets funded once. This single discipline is what separates a second year of budget from a polite conversation about value.

Score the use cases inside the domains

Every use case you find, scored so the ranking survives challenge. Use Impact, Confidence and Ease, scored 1 to 10 each:

  • Impact: how much changes if this works, against the domain metric you agreed
  • Confidence: how sure you are that it will work, based on evidence rather than enthusiasm
  • Ease: how hard it is to build and deploy, including the integration and the approvals

Multiply them, for a score out of 1,000. The point of a visible score is not precision, it is that when a director asks why their idea is fourth, you have an answer that is about the work rather than about who asked.

Record who raised each one. It costs nothing and it means you can go back to them when it reaches the top.

3. The agent register

Every deployed agent with its owner, purpose, the systems it can reach, the data it handles, and what it costs to run per day.

The cost column is the one people skip and the one your CFO will ask about first. Start it at agent one.

The register is also how you retire things. An agent nobody has used for two months should be switched off and recorded as retired, and being the person who does that makes every other number you report more believable.

Where a platform helps

You can build all three in a spreadsheet, and people do. The reason they decay is that they live apart from the work: the map is updated when someone remembers, and the backlog ages between reviews.

This is what the Kowalah Platform holds as the system of record, and what the Kowalah Plugin puts inside Claude, so that mapping a process in conversation writes back to the model rather than into notes. The Agent Hub is the register, including the per-agent daily cost.

You can start this in your first week without an engagement and without talking to us. Install the plugin, sign in, and if nobody from your company is there yet you get your own organisation with you as its admin and an empty model to fill. Map as you do discovery, one function at a time, in the conversation you are already having. If your company already has a Kowalah organisation you join that one instead, as a member, and an admin there can raise your access.

Where it stops being a solo exercise is the point where you want the people who own the processes in there with you, confirming their own and keeping them current. That is also the point where the model stops being one person’s view of the business, which is the real reason to do it.


When everyone wants to be first

Once word gets round that you exist, the demand arrives all at once. Department heads, people who saw a demo, someone’s chief of staff, all with a request and most of them reasonable. This is a better problem than silence, and it is also the point where the role is won or lost.

The job stops being technical here. Nobody hired you to take orders in the sequence they arrive.

Answer with the reasoning, not the ranking. “That is fourth on the list” ends a conversation badly. “We picked claims and procurement because they move numbers the board is already watching, and yours sits outside both, so here is what would have to be true for it to come forward” is a different conversation. The person leaves understanding the logic, and the next thing they bring you is better argued.

Make the scoring visible. A backlog people can see, with the scores and the domains on it, moves the argument off you and onto the criteria. It also means people start pre-qualifying their own ideas, which is the point at which the backlog starts improving without you.

Say no to the right things out loud. Write down the domains you did not pick and why. A no with a reason and a review date is a manageable outcome. A no by silence, where someone’s request simply never moves, is how you acquire quiet opponents.

Teach the criteria rather than applying them alone. Give the champions and the engaged department heads the scoring method and let them rank their own requests before they reach you. You are trying to build an organisation that can make these calls, not a queue with you at the front of it. This is the difference between a role that scales and one that becomes a bottleneck with a person in it.

Route the genuine small stuff elsewhere. A good portion of what arrives is someone who needs twenty minutes of help rather than a project. Have somewhere to send them: office hours, a channel, a champion in their function. Turning those away with the same “it is on the backlog” answer is how you become the department that says no.


Implementation notes

The auditor problem. If the organisation reads you as the person sent to find jobs that can be cut, no amount of skill recovers it. Have your sponsor introduce you in writing as a resource teams can call on. In discovery, ask about the work people put off and find tedious rather than the work that consumes the most headcount. The first three teams set how every later team receives you.

If you inherited a stalled programme. Find out what was built, who owns it and whether anyone uses it before adding anything. Retiring something dead and saying so publicly buys more credibility than shipping something new, because it establishes that you will tell people the truth. Then do the mapping anyway. Inherited artefacts are rarely accurate enough to build on unchecked.

Ask for help early. Work out in the first fortnight whether your gap is the engineering or the business conversation, and ask for support in it while a new joiner asking for help still reads as good judgement rather than as struggling.

Adjust for your sector. In a regulated environment the security and compliance line takes longer and the 90 days stretch, so agree that with your sponsor in week one rather than missing a date in month three. In manufacturing, logistics and construction the highest-value processes are not in the office, so budget time to be where the work happens.

Get started

How to Use This Guide

01

Read the negotiation section first

It only works before you sign. After that you are asking for favours rather than agreeing terms

02

Copy the 90-day plan

Use the 'Copy page' button, then adjust the dates and functions to your organisation

03

Agree it with your sponsor in week one

A plan your sponsor has seen and signed off is what protects you when priorities shift in month two

04

Build the three artefacts as you go

The operating model, the scored backlog and the agent register. They are what you are judged on at the end of the quarter

Questions

Frequently Asked Questions

Common questions from people starting in this role

What should I negotiate before I accept the offer?
Four things, and you have leverage for them exactly once. A named executive sponsor senior enough to open a door you cannot open yourself. A budget you control for platform and tooling costs, without a business case per agent. System access agreed with security before your first day rather than negotiated per project. And a written scope: which functions are yours this year and which are not. Each of these is a reasonable ask that a well-run organisation will already have thought about. If they cannot answer, that tells you something important while you can still act on it.
Should I build something in my first week to show value?
No, and the pressure to is the most common way this role goes wrong. Something built in week one is built on assumptions, and when it lands badly it costs you the credibility you need for the work that matters. Spend the first month mapping and the second shipping. The plan in this guide gets you to a measured result by day 60, which is fast enough for any reasonable sponsor and slow enough to be right.
How do I pick my first use case?
Take something a team asked for, not something you chose. Your first delivery sets how every later approach is received, and a team that asked for help is a team that will use what you build and say so. Beyond that: it should have a named business owner, a measurable baseline you can capture before you change anything, and a process that runs often enough that a small improvement compounds. Resist the most impressive use case in favour of the one most likely to be in daily use at day 90.
What are the three things I have to build?
An operating model, a scored use case backlog, and an agent register. The operating model is a map of how the organisation actually works: the units, the processes each owns, and where AI is landing today. The backlog is your chosen domains, the use cases inside them, and the scoring visible to anyone who asks. The register is every deployed agent with its owner, purpose, the systems it reaches, the data it handles and what it costs to run. These three are what you are judged on at the end of the quarter, and what makes you replaceable in the good sense: the programme survives you.
How do I decide where to concentrate?
Pick domains before use cases. A domain is a top-level business process: order to cash, claims handling, procurement, month-end close, bid production. Choose the two or three where AI moves a number the business already cares about, then find your use cases inside them. This is the argument in Rewired (Lamarre, Smaje and Zemmel, 2nd ed., 2026) and it holds up: value concentrates in a small number of domains, and spreading thinly across every function produces a long list of small wins that never add to anything a CFO recognises. Test each candidate domain on three things: what number it moves and who owns that number, how big the gap is between how it runs today and how it could run, and whether the executive who owns it will engage.
How do I measure whether any of this worked?
Agree the metric and its baseline for each domain before any build starts, not afterwards. It has to be evidenceable: a number someone already produces, or one you can start producing now, owned by the person who owns the domain. Cycle time on the close, cost per claim, days from bid request to submission, error rate on invoice matching. Once an agent is live you cannot go back and measure what came before, and a programme that cannot evidence its result gets funded once.
Everyone is asking me for help at once. How do I handle it?
This is where the role stops being technical. Answer with the reasoning rather than the ranking: 'that is fourth on the list' ends a conversation badly, where 'we picked claims and procurement because they move numbers the board is already watching, so here is what would have to be true for yours to come forward' leaves the person understanding the logic. Make the backlog and its scores visible so the argument is about the criteria rather than about you. Write down the domains you did not pick and why, because a no with a reason and a review date is manageable where a no by silence creates quiet opponents. And teach the criteria to your champions and department heads so they rank their own requests before they reach you. You are building an organisation that can make these calls, not a queue with you at the front of it.
How do I stop people seeing me as the person sent to cut jobs?
Get your sponsor to introduce you in writing as a resource teams can call on, before you meet anyone. Make your first delivery something a team asked for. And in discovery, ask about the work people put off and find tedious rather than the work that takes the most headcount. The first three teams you work with set how every later team receives you, so spend disproportionate effort there.
What if I have inherited a stalled programme rather than a blank page?
Find out what was built, who owns it, and whether anyone uses it, before you add anything. A stalled programme leaves agents in production with no owner and a reputation to repair. Retiring something dead and saying so publicly buys more credibility than shipping something new, because it signals that you will tell people the truth. Then run the mapping work anyway. Inherited artefacts are rarely accurate enough to build on without checking.
Do I need the engineering skills or the change skills more?
It depends on what the organisation already has. If there is a capable integration team you can lean on, your scarcity is the business conversation and the adoption work. If there is not, you will spend more time on connectors, authentication and error paths than you expect. Work out which you are in during the first fortnight and ask for support in the gap early, while a new joiner asking for help still reads as good judgement.
How do I report progress to a board that does not understand the work?
Report the same shapes you would for any programme: what is live, who uses it, what changed against a baseline, and what is next. Avoid capability claims without evidence and avoid counting activity. 'Four agents in production, used by 120 people weekly, month-end close down from nine days to five' is a board update. 'Completed AI discovery across six functions' is not, and a board that has been told the second one twice stops asking.
Where does the Kowalah platform fit into this?
You can start on it in your first week, on your own, without an engagement and without talking to us. Install the Kowalah Plugin in Claude and sign in: if nobody from your company is there yet you get your own organisation with you as its admin and an empty model to fill, and you build the operating model in conversation as you do discovery. The three artefacts in this guide are what the Kowalah Platform holds as the system of record, so what you build is the real thing rather than a demo. Where an engagement starts is when you want the people who own the processes in there with you, confirming and maintaining their own, which is also the point at which the model stops being one person's view of the business.

Ready when you are

Your CIO will ask what the programme around you needs

Two things here are written for them rather than you, and both are worth forwarding. The AI Programme Calculator sizes what a full implementation costs and how long it takes. The AI Operations Lead hiring guide sets out what this role needs from an organisation to succeed, which is the same list you negotiated on.

Size the programme