A two-week AI agent pilot: what happens day by day
A pilot answers one question: on your data, in your systems, can this process be handed to an agent, and how much does that actually save. It is not there to prove that AI works — that much is settled. Below is the schedule we use for single-process pilots, split into what we do and what has to happen on your side. The dates are approximate; the order of the steps is not. Each one closes a risk that would otherwise surface in production.
- Two weeks are enough because the scope is narrow: one process, one type of case, production data.
- The biggest schedule risk is not the technology — it is how long it takes to get access to your systems.
- A pilot ends with a number, not a presentation — and with a decision you are free to answer with a no.
Before the clock starts
The pilot starts after the process audit, so on day one we already have the process map, the arithmetic on your data and an agreed scope. That is not a formality — a pilot without that stage turns into two weeks of arguing about what we are building. Before the start we need three things: access to your systems within the boundaries of one process, one person on your side who knows that process from the inside, and permission to work on production data rather than samples.
Week one: from data to the first answers
Days 1-2: access and a sample. Connecting to the mailbox, the ERP and the CRM within the scope of one process. We take a sample of several hundred historical cases and check what they really look like — not what the process description says they look like.
Days 3-4: decision rules. The most important part of the whole pilot, and the only one that needs your team. We agree what the agent does on its own, what it prepares for approval, and what it never touches. That list is written at a table, not in code.
Day 5: first dry runs. The agent handles historical cases without sending anything. We compare its decisions with what the team actually did. The first discrepancies are normal, and they teach us the most.

Week two: from answers to a number
Days 6-8: running in parallel. The agent handles live cases, but nothing leaves the company without a human approving it. The team approves or corrects, and every correction feeds back into the rules. This is where the number of discrepancies drops.
Days 9-10: limited go-live. One narrow, selected case type goes out to customers without approval, with full logging and a single decision that switches it off. Everything else still goes through approval.
Days 11-12: measurement. We count the same things we counted before the start: cases handled without a human, time to response, number of escalations and corrections. On the same definitions as in the audit, so the comparison holds.
Days 13-14: the decision. A report with the numbers, a list of what did not work, and the scope of a full implementation with a price. The decision is yours, and a no is a legitimate pilot outcome.
The difference between a demo and a pilot. A demo shows that the technology works; a pilot shows what percentage of your cases it handles without a human — including the odd ones, which never turn up in a demo.
What usually stretches those two weeks
Almost never the technology. The most common cause is waiting for access — if granting ERP permissions requires a ticket and a week of lead time, the schedule slips by that week before anything begins. The second is the absence of one decision-maker for the process: when escalation rules are set by a committee, days 3-4 stretch into a week and a half.
The third is rarer but the most serious: data that turns out to be worse than the audit suggested. If the sample shows that ERP statuses are sometimes out of date, we stop the pilot and come back with a recommendation to clean up the data first. Carrying on would give you an agent that repeats an existing error faster and across more cases.
Frequently asked questions
Does the pilot run on real customer data?
Yes — on production data, within the scope of one process, with every action logged. A pilot on test data does not answer the question it exists to answer: the unusual cases, the ones that decide the outcome, never appear in test sets.
What do we keep if we say no after the pilot?
The measurement report, the list of data and process problems we identified, and the documentation of the decision rules — all of it stays with you and has value regardless of whether we work together further. The agent itself is switched off and the access is revoked.
Why two weeks specifically?
Because the scope is deliberately narrow: one process and one case type. With a wider scope the same schedule does not close, and it is better to say so at the start than to discover it on day ten. Processes spanning several systems or carrying a lot of exceptions we quote at three weeks.