Workflow diagnosis

How to Choose the First AI Use Case

By: The Nova Labs Team8 min read

Reviewed by Nova Labs AI Operations

The short answer

Choose the first AI use case by scoring candidates on four things: whether you have a measured baseline, whether the output can be verified, how sensitive the data is, and whether a named person will own it. Pick the highest scorer, not the most impressive idea. The first project's job is to produce a result your team can judge and learn from, not to be a showcase.

Key takeaways
The criteriaMeasurable baseline, verifiable output, data sensitivity, and named ownership. Score each candidate on all four.
How many to compareFive to eight candidates. Fewer is not a comparison, more and the scoring drifts.
Best first projectUsually internal, document or drafting shaped, reviewed by a human before anyone outside sees it.
DisqualifiersNo baseline, no owner, unverifiable output, or regulated data with no controls in place.
Set before you buildThe measurement, the review date, and the number that would count as failure.

The first project is a learning instrument

Microsoft's 2026 Work Trend Index put 19 percent of AI users in its high-capability, high-readiness group. The gap that figure describes is not about tool access, which is close to universal. It is about whether an organisation is set up to turn access into working practice.

That shapes what the first project is for. It is not the project with the biggest theoretical payoff. It is the one that will teach your team the most about how these systems fail, while the consequences of failure are still small and reversible. A first project that succeeds but that nobody can explain teaches nothing. A first project that fails visibly and cheaply, with a clear reason, is worth more.

The four criteria

Measurable baseline

You need to know what the process costs today before you change it, derived from timings and volumes rather than impressions. Without a baseline the project has no way to be evaluated, and it will be judged on how people feel about it three months later. The arithmetic is in How to Calculate the Cost of Manual Work.

Verifiable output

Can somebody tell, quickly and reliably, whether a given output is correct? Extracting an invoice total is verifiable: the number matches the document or it does not. Summarising the mood of a client relationship is not. Unverifiable outputs are not automatically bad ideas, but they make poor first projects because you cannot build a calibration loop around them.

Data sensitivity

Rank the data the process touches: public, internal, confidential, regulated. Early projects belong at the low end. This is not only a compliance point. Sensitive data forces controls, approvals, and legal review that will slow the project to the point where the team learns nothing for months.

Named ownership

One person, by name, who is responsible for the system after launch: watching outputs, handling exceptions, deciding when it needs changing. Not a committee and not a job title with nobody currently in it. An implementation without an owner degrades quietly until someone notices it has been wrong for a while.

The scoring method

Score each candidate 0 to 3 against each criterion and add them up. The maximum is 12. The scale is coarse on purpose: fine-grained scoring invites arguing about single points instead of comparing candidates.

Scoring rubric
ScoreBaselineVerifiabilityData sensitivityOwnership
3Measured, with volumesMechanically checkablePublic or internalNamed, has authority
2Partly measuredCheckable by a person quicklyInternal, some confidentialNamed, needs sign-off
1Estimated onlyRequires expert reviewConfidentialTeam owns it, no individual
0NoneCannot be checkedRegulated, no controlsNobody

A worked comparison

The same 40-person services firm from the audit guide finishes its audit with five candidates. Scored:

Five candidates scored
CandidateBaselineVerifyDataOwnerTotal
Draft first-pass client update322310
Extract data from supplier invoices33219
Answer staff policy questions12328
Summarise sales calls into the CRM21115
Customer-facing support chatbot11013

These scores are an illustration of the method, not a benchmark. The same five candidates would score differently at a firm with a finance lead who wanted to own invoice extraction, which is exactly the point: the rubric surfaces local conditions rather than producing a universal ranking.

The chatbot scored lowest, and it was the idea the leadership team arrived with. It is customer-facing, so errors are public. It touches customer data with no controls in place. Nobody had volunteered to own it. And there was no baseline for what support currently costs, so success could not have been measured.

The client update drafting won on ownership and baseline. Output is reviewed by a manager before it goes out, so mistakes stay internal while the team learns what the system gets wrong.

Setting success before the build
  • Baseline = 120 min/week across 3 people (measured)
  • Target = under 45 min/week within 8 weeks
  • Quality gate = manager edits fewer than 1 in 4 drafts heavily
  • Failure condition = above 70 min/week at week 8, or heavy edits above 50%
  • Review date = 8 weeks from launch, booked now
  • Owner = named account manager

Written down before anything is built, so the result can be judged.

The failure condition matters as much as the target. A project that cannot fail cannot succeed either, and without a stated threshold the conversation at week eight becomes a negotiation about whether it feels better.

Candidates to reject outright

Some ideas should not be scored at all. Rejecting them quickly saves the session for real comparison.

  • Anything with regulated data and no controls yet. Build the controls first or pick something else.
  • Processes about to be replaced. Automating a workflow that a migration will delete next quarter wastes the build and the learning.
  • Unstable processes. If the steps change week to week because nobody agreed on them, redesign first. Automating instability produces faster instability.
  • Anything with no human in the loop on a first project. You have not yet calibrated what the system gets wrong, so you cannot judge how much supervision it needs.
  • Ideas nobody will own. Enthusiasm in a workshop is not ownership. Ask who watches the output in month three.

When this does not apply

There is a burning operational problem. If one broken process is actively losing customers, fix that. Scoring is for choosing between reasonable options, not for delaying an obvious emergency.

You have already run several projects. Once the team knows how these systems fail, the criteria shift toward value and strategic fit. This rubric is calibrated for the first one or two.

No candidate scores above about 6. That is a signal about readiness rather than about the candidates. The work to do first is measurement, ownership, and data hygiene, not a build.

The real constraint is not capacity. If the bottleneck is demand, or a decision nobody will make, recovering hours in the back office will not move anything that matters.

Running the session

Ninety minutes with the people who do the work and one person who can commit resources. List candidates from the audit. Score each one together, out loud, with someone recording. Where a score is disputed, record both numbers and the reason rather than averaging them, because the disagreement is usually more informative than the score.

Leave with one chosen candidate, its baseline, its target, its failure condition, its owner, and its review date. Anything less and the session produced a shortlist, not a decision.

Frequently asked questions

What is the best first AI project for a small business?
The one with a measurable baseline, a named owner, low data sensitivity, and a verifiable output. In practice that is usually an internal document or drafting task rather than anything customer-facing, because mistakes stay inside the building while the team learns what the system gets wrong.
Should the first AI project be customer-facing?
Usually not. Customer-facing work carries reputational risk before you have calibrated the system, and it is harder to correct quietly. Start where a human reviews the output before anyone outside the company sees it.
How many use cases should we evaluate at once?
Score five to eight, then pick one. Fewer than five and you have not really compared anything. More than eight and the scoring drifts because nobody can hold the criteria consistently across that many candidates in one session.
What if the highest-scoring use case is not the one leadership wants?
Show the scores and the criteria, then let leadership override with the trade-off in writing. The point of scoring is not to remove judgement but to make it explicit. An override with a stated reason is fine. An override that quietly changes the criteria is not.
How do we know the first project worked?
Define success before the build using the same measurements you used for the baseline, and set the review date at the same time. If nobody can say in advance what number would count as failure, the project has no way to fail and therefore no way to succeed.

Sources and method

Microsoft 2026 Work Trend Index
Cited for the gap between AI access and organisational readiness (19 percent of AI users sat in the high-capability, high-readiness group).
McKinsey, The State of AI in 2025
Cited for adoption and scaling patterns, and for workflow redesign as the condition for measurable impact.

Figures in this article are either cited to a named source above, derived from arithmetic shown in full, or labelled as an estimate with the assumptions stated. Last reviewed August 3, 2026.

Keep reading