Key takeaways
- Pilot two workflows, not a tool: frequent, text-heavy, measurable, owned by someone on the team, and clear of regulated data.
- Measure the baseline for a week before anyone touches AI, or you will have nothing to compare against.
- Pick five people on purpose: two enthusiasts, two skeptics, and the person who is best at the task today.
- Write the decision rules before the pilot starts, so the readout is a comparison against a bar rather than a debate.
- Scale in cohorts with a playbook and a named owner; a pilot that ends with a slide and no next cohort was a demo.
Why most pilots produce nothing
The usual AI pilot goes like this: buy thirty licenses, send a launch email, wait a quarter, survey people, learn that "some people love it," and renew or cancel based on vibes. Nothing was measured because nothing was defined, and the decision at the end is the same guess it was at the start.
A pilot that produces a decision has five parts: specific workflows, a baseline, a small group with real training, a fixed measurement window, and decision rules written before the data comes in. The rest of this guide walks through them in order and ends with a plan template and a readout template.
Step 1: pick two workflows
Pilot workflows, not tools. "Try Copilot" is not a pilot. "Use Copilot to produce the first draft of customer QBR decks" is.
Two workflows is the right number. One gives you no comparison and makes a single bad fit look like a verdict on AI. Three or more spreads five people too thin. Choose each against this list:
- Frequent. Happens at least a few times a week per person, so two weeks yields dozens of data points.
- Text-heavy or data-heavy. Drafting, summarizing, classifying, extracting, reformatting, first-pass analysis. These are where current tools are strongest.
- Measurable. You can count units and time them, and someone can judge the quality of the output.
- Owned. One person on your team is accountable for the workflow today and will be during the pilot.
- Clear of regulated data, or already covered by an approved tool and plan. Do not make compliance a pilot dependency.
- Painful. People complain about it. Adoption follows relief.
Examples that fit: first drafts of support replies, meeting notes into action items, weekly status reports, screening job applications against a rubric, turning call transcripts into CRM notes, summarizing contracts for internal review, first-pass RFP responses, cleaning and pivoting a monthly export. Examples that do not: anything done once a month, anything that needs a judgment call nobody can describe, and anything where the output goes straight to a customer with no review.
Step 2: measure the baseline for one week
Before anyone opens an AI tool, spend one week measuring how the workflow runs now. Without this, the pilot produces the sentence "people feel faster," and feelings do not survive a budget review.
Keep the log tiny or nobody will fill it in. One row per unit of work with five fields: date, who, what (one phrase), minutes spent, and a quality note (a 1 to 3 score from the owner, or "sent as-is" versus "needed rework"). A shared sheet is fine. If your ticketing or project tool already stamps start and end times, use those and skip the minutes field.
Also capture cycle time where it matters: how long the customer or colleague waited from request to result, which is often the number leadership actually cares about.
Write the baseline in one line per workflow: "Support first replies: about 14 minutes each, roughly 60 per person per week, 1 in 5 needs a second pass." You will compare against this in three weeks.
Step 3: choose the tool and the data rules
Tool. Use the business tier of something you already have where you can. Microsoft Copilot if you are a Microsoft 365 shop, Gemini if you run on Google Workspace, or a team plan of ChatGPT or Claude. Business tiers matter because their standard terms do not use your inputs for training and they give you admin control over who has access. Pick one tool per workflow, or you cannot tell what caused the result.
Data rules. Write three lists and hand them to the five people on day one:
- Green: fine to paste. Internal drafts, public information, your own notes, templates, anonymized examples.
- Yellow: paste only in the approved tool and only what the task needs. Customer names on a support ticket inside your approved Copilot tenant, for instance.
- Red: never, in any tool, during the pilot. Health records, payroll and financial account data, legal matters under privilege, anything with a contractual confidentiality clause, a client's source code, and personal data of minors.
If you do not already have a company policy, the AI policy generator produces a starting draft and how to write a team AI policy explains the choices. A pilot is a good excuse to write the policy you needed anyway, and it shuts down the shadow AI that is already happening on personal accounts.
Step 4: train five people
Five per workflow. Choose them on purpose: two people who volunteered, two who are skeptical but fair, and the person who is currently best at the task. The enthusiasts find the ceiling, the skeptics find the failure modes, and the expert shows you where the tool falls short of real expertise.
Training is ninety minutes, live, hands-on, with real examples from the workflow. Cover four things:
- The prompt pattern for this workflow. Give them three prompts you have already tested, with placeholders, in a shared document they can edit. Not a general prompting course; the three prompts.
- What to check. For each prompt, the two or three things that go wrong and how to spot them. Wrong numbers, invented details, tone that does not match the customer.
- The data rules. The three lists, read aloud, with an example of each.
- The log. Same log as the baseline week, plus one field: "AI used: yes or no." And a free-text note field for anything surprising.
Then open a channel where the five post prompts that worked and outputs that failed, and hold fifteen-minute office hours twice a week. That channel becomes the playbook in Step 7.
Step 5: measure for two weeks
Two weeks is long enough for the novelty to fade and short enough that people keep logging. During the window:
- Check the log every other day; a participant who stops logging is data you will not have.
- Hold one twenty-minute check-in per week with all five. Ask three questions: what worked, what failed, what would you change about the prompts.
- Collect artifacts. Save five outputs that were good and five that were bad, with the prompts that produced them. The readout needs examples, not adjectives.
- Watch review time. If a reply now takes six minutes to draft and eight to fix, the workflow did not get faster, and only the log will show it.
- Fix prompts freely, but do not swap the tool or reshape the workflow mid-pilot; that resets the clock.
Step 6: decide
Write the decision rules before the pilot starts and put them in the plan. Something like: "Scale if time per unit falls by at least a quarter with no drop in the quality score. Extend two weeks if the result is mixed or the log is thin. Redesign if the workflow changed shape. Stop if time or quality got worse."
At the end of the two weeks, compare the log against the baseline and the rules. There are exactly four outcomes: scale, extend, redesign, stop. "Stop" is a legitimate result and a cheap one; you spent five people's partial attention for three weeks and learned that this workflow is not the place. Write down why and move to the next candidate.
The full method for turning the log into numbers without fooling yourself is in how to measure AI ROI.
Step 7: scale
Scaling is a second, larger pilot, not a launch email.
- Turn the channel's best prompts and the failure examples into a two-page playbook per workflow.
- Train the next cohort of ten with the same ninety-minute session, taught by one of the original five.
- Name an owner for each workflow who maintains the prompts and reads the log monthly.
- Keep measuring, at lower intensity: a weekly count of units and a monthly quality sample.
- Update the data rules and the policy with anything the pilot surfaced.
After two cohorts, you have a repeatable method and enough data to fund the next two workflows.
Worked example: a support team pilot
Suppose you run a twelve-person customer support team. You pick two workflows: first drafts of email replies, and turning resolved tickets into knowledge base article updates. The baseline week shows replies take about 14 minutes each, and an article update takes about 50 minutes and happens twice a week per person, which is why the knowledge base is out of date.
You choose Copilot, since the team lives in Outlook and SharePoint, with the rule that ticket content stays inside the tenant and no payment details ever go in. Five people are trained on three prompts: draft a reply from the ticket and the relevant help article, rewrite a reply in the team's tone, and draft an article update from a resolved ticket.
Two weeks later, the log shows replies averaging around 9 minutes including the fix-up pass, with the quality score unchanged, and article updates at about 20 minutes, with four times as many updates completed. Two of the five said the reply draft was sometimes worse than starting from scratch on complex tickets; the expert flagged that the draft occasionally cited an outdated article. Your decision rules said scale at a quarter reduction with no quality drop. You scale reply drafting with a rule that complex tickets skip the draft, scale article updates as-is, and add "check the article date" to the playbook.
These numbers are an illustration of what a readout looks like, not a benchmark. Your baseline will be your own.
Template: the one-page pilot plan
Fill this in before Step 2 and share it with your manager and the five participants.
| Field | Your answer |
|---|---|
| Workflows 1 and 2 | One sentence each, with the unit of work named |
| Owner per workflow | Name |
| Baseline (from week 1) | Minutes per unit, units per week, quality measure, cycle time |
| Tool and plan | Name, tier, who administers it |
| Data rules | Green, yellow, and red lists |
| Participants | Five names and why each was chosen |
| Prompts | Link to the shared prompt doc, three per workflow |
| Measurement | Log location, fields, who checks it and when |
| Window | Baseline dates, training date, two-week pilot dates, readout date |
| Decision rules | Scale, extend, redesign, and stop thresholds |
| Risks | Top three, with the mitigation for each |
If you would rather draft it in a chat assistant, this prompt gets you a usable first version.
You are an operations consultant helping a manager design a two-week AI pilot. Ask me up to eight questions, one at a time, to fill in a pilot plan with these fields: two workflows and their unit of work, an owner per workflow, how we will measure a one-week baseline, the tool and plan we will use, green/yellow/red data rules, five participants and why each, three prompts per workflow, the log fields, dates, decision rules for scale/extend/redesign/stop, and the top three risks.
Context: I manage [TEAM AND SIZE] at [COMPANY TYPE]. The candidate workflows are [LIST TWO TO FOUR WORKFLOWS]. We already pay for [TOOLS THE COMPANY HAS]. Regulated or confidential data we handle: [DESCRIBE OR SAY NONE].
Once you have my answers, produce the plan as a one-page table, then a short list of what could invalidate the results. Do not invent numbers; where I have not given you a baseline, leave it blank and mark it "measure in week 1."
Template: the readout
One page, in this order, sent before the meeting.
- What we tested. Two workflows, five people, dates, tool.
- What we measured. Baseline versus pilot for each workflow: minutes per unit, units, quality score, cycle time, with the sample size next to each number.
- What happened. Three good examples, three failures, and what the participants said, in their words.
- What it costs to scale. Licenses, training time, an owner's time, review overhead.
- Risks and rules. Data rules, the human-in-the-loop points, anything compliance needs to see.
- Recommendation. Scale, extend, redesign, or stop, tied to the decision rules from the plan.
- The ask. Budget, headcount for the next cohort, or a decision by a date.
Lead with number six. Leaders read the recommendation first, and a readout that makes them hunt for it loses them.
Next steps
- Turn the log into numbers leadership will trust: how to measure AI ROI and the time savings calculator.
- Set the rules before day one: how to write a team AI policy and the AI policy generator.
- Prepare the pitch: how to talk to your boss about AI.
- Check where your team stands with the AI readiness quiz.
- Prompts for the management side: leadership and communication and hiring and management.
- Role playbooks for common pilot teams: AI for operations managers and AI for customer service reps.
Frequently asked questions
How long should an AI pilot last?
How many people should be in an AI pilot?
Which AI tool should we pilot?
What if the pilot shows no time savings?
Keep going
How to measure AI ROI without fooling yourself
How to measure AI ROI honestly: time saved, quality, cycle time, cost, and error rates, with a simple worksheet and a one-page readout for leadership.
SafetyHow to write a team AI policy (with a fill-in template)
A manager's step-by-step for a team AI policy: scope, approved tools, data rules, disclosure, verification duties, IP, training, and a fill-in template.
FundamentalsHow to use AI at work: a practical operating manual
An operating manual for using AI at work: pick a tool, learn five daily use cases, prompt well, verify output, protect data, and follow a 30-day plan.
SafetyAI privacy at work: what happens to what you paste
Where your pasted text goes, how consumer and enterprise AI plans differ, a red/yellow/green data test, and how to ask IT for an approved tool.
ToolsBest AI tools for work: a curated shortlist by category
A curated list of AI tools for work by category: chat, writing, meetings, research, slides, automation, data, and coding, plus how to evaluate any tool.
CareerHow to talk to your boss about AI: making the case for a tool, budget, or pilot
Pitch an AI tool, budget, or pilot to your boss: frame the ask around one task, answer security, cost, quality, and jobs objections, and use the template.
Job playbookAI for Operations Managers
Operations managers live in SOPs, incident reports, staffing plans and spreadsheets, which is the material AI handles best. Here is how to use it to get your week back, and where the line is for safety and people decisions.
Job playbookAI for Project Managers
AI cannot run your project, but it can draft the status report, turn a messy meeting into an action list, and pressure-test your risk register in minutes. Here is how project managers use it without losing the plot.
Job playbookAI for HR Professionals
First drafts of policies, announcements, and survey analysis in minutes, with employee data and employment decisions kept where the law and your judgment say they belong.
Job playbookAI for Small Business Owners
AI answers the one-star review calmly, drafts the month of posts, turns your voice memo into an SOP, and preps the questions for your CPA. You still make the calls, sign the checks, and own what goes out under your name.
Job playbookAI for Customer Success Managers
A CSM's calendar is calls, and the work between them is writing: follow-ups, success plans, QBR decks, renewal notes. AI takes the writing and the first pass at the usage data. Here is how to use it, and which customer data must stay inside your CRM.