Most AI pilots do not collapse because the model was not good enough. They collapse quietly, a fortnight in, when the person who was enthusiastic in the kick-off goes back to doing the job the old way because the old way is what fits their Tuesday. That is an organisational failure, not a technical one, and it is avoidable.
Pilot a task, not a tool
The most common design mistake is scoping the pilot around software: three months to evaluate a platform, a committee, a report at the end. Nothing about that produces a changed way of working, and everyone can feel it.
Scope it around one task instead. Not “explore AI in operations” but “the weekly supplier summary that takes someone half of every Wednesday”. A task has an owner, a frequency, a current cost and an obvious before-and-after. A tool has a vendor.
Give it to someone who does the work
The pilot should be built and owned by the person whose task it is, not by an innovation team building on their behalf. Two reasons. They know the exceptions and the judgement calls that never made it into any process document, and those are exactly what a workflow gets wrong. And a workflow handed to someone as a finished object gets abandoned the first time it misbehaves, because they have no idea what is inside it.
If that person cannot build it, that is worth knowing early. It is a stronger argument for teaching them than for building it for them.
Define “better” before you start
Write down, in advance, what the task costs today and what would count as an improvement. Hours, error rate, turnaround, or simply “it stops being the thing they dread on a Wednesday”. Vague success criteria always resolve to “it was interesting” and then to nothing.
Be honest that the comparison is against how the job is done now, not against perfection. A workflow that gets a summary eighty per cent right in four minutes may beat a person doing it perfectly in three hours, or may not - but you cannot tell without having written down which one matters here.
Build the check into the workflow
The reason pilots lose trust is a bad output that nobody caught. A pipeline that drafts and then explicitly reviews its own draft against a stated standard catches far more than one that drafts and stops, and it leaves you a record of what was checked. Where the output is wrong in a way that would be expensive and hard to spot, that is a signal to keep the human step, not to try harder.
Plan for week two
Week one has attention. Week two has the rest of the job. This is where most pilots quietly die, so decide up front what happens there: who notices if it stops being used, when you look at it again, and what small fix is allowed without another approval round. A pilot with a review date survives. One that ends with “let us see how it goes” does not.
And keep the pilot small enough that it can succeed. One task, done properly, changes more minds in an organisation than a broad programme that produces a slide deck.
Where this gets practical
All of this is easier to agree with than to do, and the hardest part is the first working version. That is what a day at the Oxford Agentic Bootcamp is for: bring the real task, build the workflow yourself with someone on hand when you get stuck, and leave with it running rather than with a plan to run it. No coding required.