1. Home
  2. Guides
  3. How to run an AI pilot: a manager's plan from two workflows to a decision
Career

How to run an AI pilot: a manager's plan from two workflows to a decision

A manager's step-by-step AI pilot plan: pick two workflows, set a baseline, choose tools and data rules, train five people, measure two weeks, decide.

Key takeaways

  • Pilot two workflows, not a tool: frequent, text-heavy, measurable, owned by someone on the team, and clear of regulated data.
  • Measure the baseline for a week before anyone touches AI, or you will have nothing to compare against.
  • Pick five people on purpose: two enthusiasts, two skeptics, and the person who is best at the task today.
  • Write the decision rules before the pilot starts, so the readout is a comparison against a bar rather than a debate.
  • Scale in cohorts with a playbook and a named owner; a pilot that ends with a slide and no next cohort was a demo.

Why most pilots produce nothing

The usual AI pilot goes like this: buy thirty licenses, send a launch email, wait a quarter, survey people, learn that "some people love it," and renew or cancel based on vibes. Nothing was measured because nothing was defined, and the decision at the end is the same guess it was at the start.

A pilot that produces a decision has five parts: specific workflows, a baseline, a small group with real training, a fixed measurement window, and decision rules written before the data comes in. The rest of this guide walks through them in order and ends with a plan template and a readout template.

Step 1: pick two workflows

Pilot workflows, not tools. "Try Copilot" is not a pilot. "Use Copilot to produce the first draft of customer QBR decks" is.

Two workflows is the right number. One gives you no comparison and makes a single bad fit look like a verdict on AI. Three or more spreads five people too thin. Choose each against this list:

  • Frequent. Happens at least a few times a week per person, so two weeks yields dozens of data points.
  • Text-heavy or data-heavy. Drafting, summarizing, classifying, extracting, reformatting, first-pass analysis. These are where current tools are strongest.
  • Measurable. You can count units and time them, and someone can judge the quality of the output.
  • Owned. One person on your team is accountable for the workflow today and will be during the pilot.
  • Clear of regulated data, or already covered by an approved tool and plan. Do not make compliance a pilot dependency.
  • Painful. People complain about it. Adoption follows relief.

Examples that fit: first drafts of support replies, meeting notes into action items, weekly status reports, screening job applications against a rubric, turning call transcripts into CRM notes, summarizing contracts for internal review, first-pass RFP responses, cleaning and pivoting a monthly export. Examples that do not: anything done once a month, anything that needs a judgment call nobody can describe, and anything where the output goes straight to a customer with no review.

Step 2: measure the baseline for one week

Before anyone opens an AI tool, spend one week measuring how the workflow runs now. Without this, the pilot produces the sentence "people feel faster," and feelings do not survive a budget review.

Keep the log tiny or nobody will fill it in. One row per unit of work with five fields: date, who, what (one phrase), minutes spent, and a quality note (a 1 to 3 score from the owner, or "sent as-is" versus "needed rework"). A shared sheet is fine. If your ticketing or project tool already stamps start and end times, use those and skip the minutes field.

Also capture cycle time where it matters: how long the customer or colleague waited from request to result, which is often the number leadership actually cares about.

Write the baseline in one line per workflow: "Support first replies: about 14 minutes each, roughly 60 per person per week, 1 in 5 needs a second pass." You will compare against this in three weeks.

Step 3: choose the tool and the data rules

Tool. Use the business tier of something you already have where you can. Microsoft Copilot if you are a Microsoft 365 shop, Gemini if you run on Google Workspace, or a team plan of ChatGPT or Claude. Business tiers matter because their standard terms do not use your inputs for training and they give you admin control over who has access. Pick one tool per workflow, or you cannot tell what caused the result.

Data rules. Write three lists and hand them to the five people on day one:

  • Green: fine to paste. Internal drafts, public information, your own notes, templates, anonymized examples.
  • Yellow: paste only in the approved tool and only what the task needs. Customer names on a support ticket inside your approved Copilot tenant, for instance.
  • Red: never, in any tool, during the pilot. Health records, payroll and financial account data, legal matters under privilege, anything with a contractual confidentiality clause, a client's source code, and personal data of minors.

If you do not already have a company policy, the AI policy generator produces a starting draft and how to write a team AI policy explains the choices. A pilot is a good excuse to write the policy you needed anyway, and it shuts down the shadow AI that is already happening on personal accounts.

Step 4: train five people

Five per workflow. Choose them on purpose: two people who volunteered, two who are skeptical but fair, and the person who is currently best at the task. The enthusiasts find the ceiling, the skeptics find the failure modes, and the expert shows you where the tool falls short of real expertise.

Training is ninety minutes, live, hands-on, with real examples from the workflow. Cover four things:

  1. The prompt pattern for this workflow. Give them three prompts you have already tested, with placeholders, in a shared document they can edit. Not a general prompting course; the three prompts.
  2. What to check. For each prompt, the two or three things that go wrong and how to spot them. Wrong numbers, invented details, tone that does not match the customer.
  3. The data rules. The three lists, read aloud, with an example of each.
  4. The log. Same log as the baseline week, plus one field: "AI used: yes or no." And a free-text note field for anything surprising.

Then open a channel where the five post prompts that worked and outputs that failed, and hold fifteen-minute office hours twice a week. That channel becomes the playbook in Step 7.

Step 5: measure for two weeks

Two weeks is long enough for the novelty to fade and short enough that people keep logging. During the window:

  • Check the log every other day; a participant who stops logging is data you will not have.
  • Hold one twenty-minute check-in per week with all five. Ask three questions: what worked, what failed, what would you change about the prompts.
  • Collect artifacts. Save five outputs that were good and five that were bad, with the prompts that produced them. The readout needs examples, not adjectives.
  • Watch review time. If a reply now takes six minutes to draft and eight to fix, the workflow did not get faster, and only the log will show it.
  • Fix prompts freely, but do not swap the tool or reshape the workflow mid-pilot; that resets the clock.

Step 6: decide

Write the decision rules before the pilot starts and put them in the plan. Something like: "Scale if time per unit falls by at least a quarter with no drop in the quality score. Extend two weeks if the result is mixed or the log is thin. Redesign if the workflow changed shape. Stop if time or quality got worse."

At the end of the two weeks, compare the log against the baseline and the rules. There are exactly four outcomes: scale, extend, redesign, stop. "Stop" is a legitimate result and a cheap one; you spent five people's partial attention for three weeks and learned that this workflow is not the place. Write down why and move to the next candidate.

The full method for turning the log into numbers without fooling yourself is in how to measure AI ROI.

Step 7: scale

Scaling is a second, larger pilot, not a launch email.

  • Turn the channel's best prompts and the failure examples into a two-page playbook per workflow.
  • Train the next cohort of ten with the same ninety-minute session, taught by one of the original five.
  • Name an owner for each workflow who maintains the prompts and reads the log monthly.
  • Keep measuring, at lower intensity: a weekly count of units and a monthly quality sample.
  • Update the data rules and the policy with anything the pilot surfaced.

After two cohorts, you have a repeatable method and enough data to fund the next two workflows.

Worked example: a support team pilot

Suppose you run a twelve-person customer support team. You pick two workflows: first drafts of email replies, and turning resolved tickets into knowledge base article updates. The baseline week shows replies take about 14 minutes each, and an article update takes about 50 minutes and happens twice a week per person, which is why the knowledge base is out of date.

You choose Copilot, since the team lives in Outlook and SharePoint, with the rule that ticket content stays inside the tenant and no payment details ever go in. Five people are trained on three prompts: draft a reply from the ticket and the relevant help article, rewrite a reply in the team's tone, and draft an article update from a resolved ticket.

Two weeks later, the log shows replies averaging around 9 minutes including the fix-up pass, with the quality score unchanged, and article updates at about 20 minutes, with four times as many updates completed. Two of the five said the reply draft was sometimes worse than starting from scratch on complex tickets; the expert flagged that the draft occasionally cited an outdated article. Your decision rules said scale at a quarter reduction with no quality drop. You scale reply drafting with a rule that complex tickets skip the draft, scale article updates as-is, and add "check the article date" to the playbook.

These numbers are an illustration of what a readout looks like, not a benchmark. Your baseline will be your own.

Template: the one-page pilot plan

Fill this in before Step 2 and share it with your manager and the five participants.

Field Your answer
Workflows 1 and 2 One sentence each, with the unit of work named
Owner per workflow Name
Baseline (from week 1) Minutes per unit, units per week, quality measure, cycle time
Tool and plan Name, tier, who administers it
Data rules Green, yellow, and red lists
Participants Five names and why each was chosen
Prompts Link to the shared prompt doc, three per workflow
Measurement Log location, fields, who checks it and when
Window Baseline dates, training date, two-week pilot dates, readout date
Decision rules Scale, extend, redesign, and stop thresholds
Risks Top three, with the mitigation for each

If you would rather draft it in a chat assistant, this prompt gets you a usable first version.

You are an operations consultant helping a manager design a two-week AI pilot. Ask me up to eight questions, one at a time, to fill in a pilot plan with these fields: two workflows and their unit of work, an owner per workflow, how we will measure a one-week baseline, the tool and plan we will use, green/yellow/red data rules, five participants and why each, three prompts per workflow, the log fields, dates, decision rules for scale/extend/redesign/stop, and the top three risks.

Context: I manage [TEAM AND SIZE] at [COMPANY TYPE]. The candidate workflows are [LIST TWO TO FOUR WORKFLOWS]. We already pay for [TOOLS THE COMPANY HAS]. Regulated or confidential data we handle: [DESCRIBE OR SAY NONE].

Once you have my answers, produce the plan as a one-page table, then a short list of what could invalidate the results. Do not invent numbers; where I have not given you a baseline, leave it blank and mark it "measure in week 1."

Template: the readout

One page, in this order, sent before the meeting.

  1. What we tested. Two workflows, five people, dates, tool.
  2. What we measured. Baseline versus pilot for each workflow: minutes per unit, units, quality score, cycle time, with the sample size next to each number.
  3. What happened. Three good examples, three failures, and what the participants said, in their words.
  4. What it costs to scale. Licenses, training time, an owner's time, review overhead.
  5. Risks and rules. Data rules, the human-in-the-loop points, anything compliance needs to see.
  6. Recommendation. Scale, extend, redesign, or stop, tied to the decision rules from the plan.
  7. The ask. Budget, headcount for the next cohort, or a decision by a date.

Lead with number six. Leaders read the recommendation first, and a readout that makes them hunt for it loses them.

Next steps

Frequently asked questions

How long should an AI pilot last?
One week of baseline measurement, a short training session, then two weeks of measured use is enough to make a scale-or-stop decision for a frequent workflow. Longer pilots mostly add drift: people forget to log, the novelty wears off unevenly, and the numbers get harder to trust. If a workflow only happens a few times a month, pick a different workflow.
How many people should be in an AI pilot?
Five per workflow is a good number. It is enough to see whether results hold across different people and small enough that you can talk to each of them every week. Choose deliberately: mix enthusiasts and skeptics, and include the person who is currently best at the task so you learn where AI falls short of an expert.
Which AI tool should we pilot?
Usually the business or enterprise tier of a tool you already pay for, such as Copilot if you are on Microsoft 365 or Gemini if you are on Google Workspace, or a team plan of ChatGPT or Claude. The tool matters less than the data rules and the training. Pilot one tool per workflow so you can tell what caused the result.
What if the pilot shows no time savings?
That is a successful pilot: you learned something cheaply. Check whether the workflow was a poor fit, the training was thin, or the review time ate the savings. Sometimes the win is quality or cycle time rather than minutes. If none of those hold up, stop, write down why, and pick the next workflow.

Keep going

Career

How to measure AI ROI without fooling yourself

How to measure AI ROI honestly: time saved, quality, cycle time, cost, and error rates, with a simple worksheet and a one-page readout for leadership.

Safety

How to write a team AI policy (with a fill-in template)

A manager's step-by-step for a team AI policy: scope, approved tools, data rules, disclosure, verification duties, IP, training, and a fill-in template.

Fundamentals

How to use AI at work: a practical operating manual

An operating manual for using AI at work: pick a tool, learn five daily use cases, prompt well, verify output, protect data, and follow a 30-day plan.

Safety

AI privacy at work: what happens to what you paste

Where your pasted text goes, how consumer and enterprise AI plans differ, a red/yellow/green data test, and how to ask IT for an approved tool.

Tools

Best AI tools for work: a curated shortlist by category

A curated list of AI tools for work by category: chat, writing, meetings, research, slides, automation, data, and coding, plus how to evaluate any tool.

Career

How to talk to your boss about AI: making the case for a tool, budget, or pilot

Pitch an AI tool, budget, or pilot to your boss: frame the ask around one task, answer security, cost, quality, and jobs objections, and use the template.

Job playbook

AI for Operations Managers

Operations managers live in SOPs, incident reports, staffing plans and spreadsheets, which is the material AI handles best. Here is how to use it to get your week back, and where the line is for safety and people decisions.

Job playbook

AI for Project Managers

AI cannot run your project, but it can draft the status report, turn a messy meeting into an action list, and pressure-test your risk register in minutes. Here is how project managers use it without losing the plot.

Job playbook

AI for HR Professionals

First drafts of policies, announcements, and survey analysis in minutes, with employee data and employment decisions kept where the law and your judgment say they belong.

Job playbook

AI for Small Business Owners

AI answers the one-star review calmly, drafts the month of posts, turns your voice memo into an SOP, and preps the questions for your CPA. You still make the calls, sign the checks, and own what goes out under your name.

Job playbook

AI for Customer Success Managers

A CSM's calendar is calls, and the work between them is writing: follow-ups, success plans, QBR decks, renewal notes. AI takes the writing and the first pass at the usage data. Here is how to use it, and which customer data must stay inside your CRM.

Terms in this guide