01 / what to build

Figuring out what to actually build

Most AI shortlists are assembled from what a vendor demonstrated rather than from work anyone actually does. Here is where the better answer is already sitting, and how to tell a real candidate from an expensive one.

A hand sorting through a table covered in identical grey cards, lifting one card outlined in red clear of the pile and pushing the rest aside.

The big picture

The question stopped being whether it can be done

For a few years the interesting question was capability. Could a machine draft this, read that, answer the other. That question is largely settled, and settling it removed the constraint that used to do the choosing for you.

Now almost anything is technically possible, which means the shortlist is no longer produced by what the technology can reach. It gets produced by whoever was in the room, what a vendor demonstrated last month, and which executive saw something impressive at a conference.

So the failure mode changed shape. Organisations are not failing because they picked something too hard. They are failing because they picked something nobody needed, measured it in a way that could never show anything, and only discovered both facts after spending a year.

Where the answer already is

Your company is already running the experiment

Two bars compared: forty per cent of companies have bought an official AI subscription, while over ninety per cent have staff using personal AI tools for work, drawn in red.

There is a finding in the MIT NANDA research that deserves more attention than the headline it came wrapped in. Only 40% of the companies surveyed had bought an official large language model subscription. Yet workers at over 90% of those same companies reported regularly using personal AI tools for work. The report puts it plainly: almost every single person used one in some form for their job.

Sit with what that means. Inside most organisations there is a live, unmanaged, self-funded experiment running right now, in which hundreds of people have independently worked out which parts of their job AI genuinely helps with. They chose the tasks themselves. They kept using it where it worked and quietly stopped where it did not. Nobody had to run a workshop.

That is the most honest signal about what to build that any organisation possesses, and almost nobody reads it. Meanwhile the formal process assembles a shortlist from a vendor slide, and the two lists rarely resemble each other.

The first thing I want to know is what people are already doing without permission. Not to shut it down. To read it.

The four filters

What separates a real candidate from an expensive one

Four filters stacked in sequence with candidate ideas passing through: evidence, edge, measurement and ownership. Most ideas fall out along the way and one continues in red.

Every idea that reaches a shortlist sounds reasonable in the room. That is what makes the room a bad instrument. These four questions are the ones that reliably separate candidates worth money from candidates worth a conversation, and an idea that fails any single one should be killed rather than parked.

  • Evidence. Is somebody already doing this work by hand, at volume, and can you name them? An idea with a named person doing it badly today is worth ten that begin "imagine if we could".
  • Edge. Does doing this well depend on something only you have — your data, your process, your customer relationships? If it does not, you are building what you will be able to buy as a feature within a year, and buying it will be cheaper.
  • Measurement. Is there a number that already exists, that somebody already looks at, that would move if this worked? Inventing a new metric alongside the project is how a project becomes unfalsifiable.
  • Ownership. Who runs this in six months, when it breaks and the enthusiasm has moved on? A candidate with no named owner is not a project, it is an intention.

The problem

Money follows what is easy to measure

A weighing scale tipped heavily toward a pan of coins and notes labelled sales and marketing at seventy per cent, easy to attribute. The raised pan, drawn in red, is labelled back office, where the return actually is. Underneath: visibility is not value.

The same research asked executives to allocate a hypothetical hundred pounds across functions. Sales and marketing took around seventy per cent of it. The report is direct about why that is a problem: back-office automation often yields better returns, and the concentration reflects easier attribution rather than actual value.

This is worth understanding as a mechanism rather than a mistake, because it repeats everywhere and nobody involved is being stupid. A marketing use case produces a number that moves within a quarter and a story somebody can tell upward. A back-office use case produces time returned to people whose time was never costed, in a process nobody outside the team can see.

One of the executives quoted in the research says it almost exactly: how do you sell an idea to a chief executive when it will not directly move revenue or reduce a measurable cost. That is not a failure of nerve. It is what happens when the appraisal system only recognises one kind of evidence.

Which is why the measurement question sits inside the filters rather than after them. If a genuinely valuable candidate has no number attached, the work is to find the number before the funding conversation, not to abandon the candidate.

The solution

The bottleneck is killing things, not thinking of them

A narrowing funnel showing sixty per cent of organisations evaluating a tool, twenty per cent reaching a pilot, and five per cent reaching production, with the final stage in red.

The same study traced what happens to enterprise-grade AI tools once organisations start looking. Sixty per cent evaluated something. Twenty per cent got as far as a pilot. Five per cent reached production.

Read that funnel carefully, because it says something counterintuitive. The expensive failure is not the idea that gets rejected. It is the idea that gets piloted, absorbs a year of attention from your best people, and then quietly does not go anywhere. Pilots are cheap to start and enormously expensive to finish, and almost nothing in a normal organisation is designed to stop one.

So the deliverable that matters from this stage is not a list of possibilities. Anyone can produce those, and increasingly anyone can produce forty of them in an afternoon. The deliverable is a shortlist with the kills attached, and the reasoning written down where people can argue with it.

Writing the kills down matters more than it sounds. An idea that was rejected without a recorded reason comes back in nine months with a new sponsor, and the organisation pays to have the same argument twice.

Being straight about it

The statistic somebody will quote at you

You will be told that 95% of AI pilots fail. It usually arrives without a source, immediately before a recommendation to buy something. It comes from the MIT NANDA report, and it is worth knowing what it actually says, because the version in circulation is not quite it.

The finding is that 95% of organisations were getting zero return, defined as no measurable profit and loss impact, while 5% were extracting millions. That is a claim about measured financial return, which is a different and narrower claim than "failed". The report also attributes the divide to approach rather than to model quality or regulation.

The caveats are in the document and rarely survive the retelling. It describes itself as preliminary findings. It is not peer reviewed. It rests on 52 organisations interviewed and 153 senior leaders surveyed across four industry conferences, which is a self-selected sample. Its own research note says the figures are directionally accurate based on individual interviews rather than official company reporting, that sample sizes vary by category, and that the definition of success differs between organisations.

There is also a straightforward internal inconsistency. The summary line for the investment section says half of budgets go to sales and marketing. The body of the same section says approximately seventy per cent. Both numbers are in the report, a few lines apart. I use the seventy because it is the one attached to the described method, and I mention the discrepancy because anyone quoting either figure at you has probably not opened the document.

None of this makes the research worthless. The shadow usage finding and the pilot funnel are the most useful things published on this subject in a while. It makes it evidence to reason with rather than a number to wave.

My approach

How I work through it

I start with the people doing the work, not the leadership team that commissioned the exercise. The bottlenecks are known with great precision by the people standing in them, and that conversation costs nothing beyond the willingness to hear that a favoured idea solves nothing.

I ask what people are already using unofficially, and I ask it in a way that does not get anyone in trouble, because the answer is worth more than the policy breach is worth punishing.

Then each candidate goes through the four filters with the reasoning visible, so you can disagree with a judgement rather than with a verdict. You end up able to run this yourself next time, which is the only version that survives my leaving.

And I will tell you when the answer is to buy something ordinary, or to fix a process, or to do nothing this year. That happens more often than the market admits, and saying it is most of what you are actually paying for.

Next

Where to start if this sounds familiar

If you have been told to do something with AI and the list keeps growing without anything shipping, the useful next move is not another round of gathering ideas. It is putting the existing list through something that kills most of it.

Send me the list, however rough, and a sentence about who asked for it. I will tell you which ones I would take further and which I would stop, and if there is nothing on it worth building I will say that too.