Baseline Before You Build: Metrics That Make AI Pilots Fundable

A woman holding a glowing stopwatch timing a friendly robot assistant as it carries a stack of papers along a glowing measuring tape laid on the floor

The pilot review meeting goes like this. The pilot team presents. Users say they love the tool. There are a couple of great quotes and a slide of screenshots. Then someone from finance asks a simple question: “How much time does it actually save?” Nobody knows, because nobody captured baseline metrics on how long the work took before the pilot started. The pilot gets extended “another quarter to gather data,” and the quarter after that, too.

That’s pilot purgatory, and it’s where most AI projects at mid-size organizations end up. It’s rarely because the AI didn’t work. It’s because nobody captured the AI pilot metrics that would prove it did. The fix is cheap, but it has to happen before you build, because once the tool is in people’s hands the “before” is gone for good.

Why the gap between pilot and production is a measurement gap

Getting a pilot started is relatively easy. Enthusiasm and a few licenses will do it. Getting a use case into production takes a funding decision, and funding decisions need evidence. A pilot that can show “this task took a median of 40 minutes, now it takes 15, and the error rate didn’t rise” gets funded. A pilot that shows “people like it” gets extended.

Here’s how the assessment asks the question, and the 0 to 4 ladder I score it against:

U2. How far has AI gone beyond experimentation?

  1. No sanctioned AI use
  2. Individuals use chat tools on their own
  3. One or two formal pilots underway
  4. At least one use case in production with measured results
  5. Multiple use cases in production and scaling

Read level 3 carefully: “in production with measured results.” Both halves matter. A use case in production without measured results can’t justify the next one, which is why so many teams reach level 3 once and then stall before level 4.

Why you get one chance at a baseline

Once people start using AI for a task, three things happen. They stop remembering how long it used to take. The process itself changes around the new tool. And any estimate you try to reconstruct afterward is colored by how people feel about the tool. If they like it, they’ll overestimate the savings; if they don’t, they’ll underestimate. Either way, it’s the weakest evidence you can bring to a budget review.

A baseline captured before the pilot starts is the only number that will hold up when someone skeptical asks about it.

What to measure

Pick from four categories, depending on the use case:

  • Time: minutes of effort per task, and end-to-end cycle time from request to done. They’re different: a task might take ten minutes of effort but three days of waiting.
  • Volume: tasks completed per person per week, or the size of the backlog.
  • Quality: error rate, rework rate, escalations, or a quality score from a reviewer.
  • Cost: labor hours at a loaded rate, plus any outside spend the task drives, such as overtime or outsourcing.

Then choose one primary metric, the number the pilot is trying to move; two supporting metrics that explain it; and one guardrail metric that must not get worse. The guardrail is the one teams forget. If AI makes a task faster but doubles the error rate, that’s not a win, and without a guardrail metric you won’t find out until a customer does.

Adoption numbers, such as active users or tasks completed with AI, are worth tracking during the pilot. But adoption isn’t value. It tells you people are using the tool, not whether the business is better off. That distinction is the subject of measuring AI value.

The asset: a two-week baseline plan and a pilot scorecard

Two-week baseline capture

  1. Day 1: define the task boundaries. Where does the task start and where does it end? “Ticket assigned” to “ticket resolved” is measurable. “Handling a customer issue” isn’t.
  2. Days 1 to 10: capture time and volume. Use system data wherever it exists: ticket timestamps, CRM activity logs, document timestamps. Where it doesn’t, ask each person doing the task to log ten occurrences with start and end times. Ten per person is enough.
  3. Day 10: sample quality. Review 20 to 30 recent outputs of the task for errors or rework. Use the same reviewer and criteria you’ll use after the pilot.
  4. Day 14: compute and sign off. Use medians, not averages; a few unusual cases can distort an average badly. The business owner signs off on the baseline numbers.

The pilot scorecard

One table, agreed and signed before the pilot starts:

  • Metric: primary, supporting, and guardrail.
  • Baseline: the number from the two-week capture.
  • Target: what result would justify production.
  • Measurement method: exactly how you’ll measure it at the end, using the same method as the baseline.
  • Pilot result: filled in at the end.
  • Decision: go to production, extend with a specific change, or stop.

A decision rule for ending pilots

Agree before the pilot starts what result means go, what means extend, and what means stop, and fix the pilot’s length in advance. Six to eight weeks is enough for most assistant and automation pilots. Write the thresholds into the scorecard and have the business owner sign it.

This does two things. It protects good pilots from moving goalposts (“now show us it works for another department”). And it protects the organization from pilots that should end but keep running on goodwill. “Stop” is a legitimate, useful outcome, and it’s much easier to reach when the criteria were set in advance. It’s also one of the biggest differences between pilots that succeed and the patterns in why AI pilots fail at mid-size companies.

Add a comparison group if you can

The most credible pilot design is also a simple one: during the pilot, half the team uses the AI tool and half doesn’t, doing the same kind of work over the same period. You compare the two groups instead of comparing “before” with “after,” which removes seasonal effects and other changes that happen during the pilot. It isn’t always practical, but when the team is large enough, it turns a good result into a result nobody argues with.

Mistakes I see at this stage

Measuring after the fact. Reconstructed baselines are guesses. Capture them first.

Using averages. One task that took four hours can make the average meaningless. Medians are more honest.

Measuring only adoption. High usage with no change in outcomes is a warning sign, not a success.

No guardrail metric. Faster and worse isn’t better. Always measure quality alongside speed.

No business owner signature. If the business didn’t agree to the baseline and targets, it won’t accept the result. The business owner’s role is covered in why IT-only pilots stall.

No end date. A pilot without a fixed length becomes a permanent experiment that nobody funds properly and nobody shuts down.

Where does your team actually stand?

Moving AI beyond experimentation is one of 24 questions in the AI Readiness assessment, which covers six dimensions: data, security, infrastructure, skills, use cases, and governance. The free version is 10 questions and gives you a score in a few minutes.

Get your free AI Readiness Score →

Want to see what the full assessment covers first? Flip through a complete 38-page sample report.

Related guides

Frequently asked questions

Why do AI pilots need a baseline?

Because once people start using AI for a task, the before is gone: memories fade, the process changes, and estimates get colored by how people feel about the tool. A baseline captured before the pilot is the only number that holds up in a budget review.

What metrics should an AI pilot track?

One primary metric the pilot aims to move, two supporting metrics that explain it, and one guardrail metric that must not get worse, usually quality or error rate. Choose from time, volume, quality, and cost, depending on the use case.

How long does it take to capture a baseline?

About two weeks: define where the task starts and ends, capture time and volume from system data or short logs of ten occurrences per person, sample 20 to 30 outputs for quality, then compute medians and have the business owner sign off.

How long should an AI pilot run?

Six to eight weeks is enough for most assistant and automation pilots. Fix the length and the go, extend, or stop criteria before the pilot starts, so good pilots aren't held to moving goalposts and weak ones don't run forever on goodwill.

Scroll to Top