The selection rule: your first AI project should solve one narrow, frequent, measurable operational problem. It should improve a real baseline, use data you can responsibly access, and preserve human control where errors could create material harm.
Many first projects begin with a tool demo and then search for a business problem. Reverse that order. The OECD's 2026 SME survey reports rapidly increasing use of AI tools, while strategic and secure integration remains uneven. A narrow use case creates the evidence needed to move from experimentation to an operating capability.
Begin with work, not ideas
Ask team leads where work waits, repeats, gets copied, gets checked again, or depends on one person's memory. Then look at the actual records: queues, inboxes, spreadsheets, tickets, invoices, CRM stages, meeting actions, and approval delays.
Write each candidate as a workflow statement: “When this event happens, the team must use these inputs to produce this outcome, with this person accountable.” If that sentence cannot be completed, the process needs discovery before automation.
The VALUE framework
VALUE is a practical scoring model for comparing first-use-case candidates. Score each dimension from 1 to 5, document the evidence, and treat the result as a decision aid rather than a mathematical truth.
V — Volume
How often does the process occur, and how much work does it create? High volume provides more potential benefit and more examples for testing. Record monthly cases, time per case, waiting time, and seasonal variation.
A — Accessibility
Can the system reliably and lawfully access the required inputs? Identify the source of truth, data owner, permissions, quality problems, retention rules, and any sensitive fields that should be excluded.
L — Loss
What avoidable loss does the current process create? Consider delay, rework, missed follow-up, error, backlog, inconsistent service, or expert time spent on preparation. Use an observed baseline, not an invented ROI claim.
U — User ownership
Is one person accountable for the workflow and able to make decisions about it? A good owner can explain exceptions, approve test cases, define acceptable performance, and decide when the pilot should pause.
E — Evaluation
Can you determine whether an output is correct and whether the business outcome improved? Define acceptance criteria before building: required fields extracted, routing accuracy, cycle-time reduction, draft acceptance, exception rate, or response time.
Add a risk gate before ranking
A high VALUE score does not make a use case safe. Apply a separate gate for consequence, uncertainty, reversibility, sensitivity, and legal or regulatory exposure.
- Consequence: what harm can an incorrect output or action create?
- Uncertainty: how often is context missing, ambiguous, or adversarial?
- Reversibility: can an action be undone before it affects a customer, payment, record, or right?
- Sensitivity: does the workflow expose personal, financial, confidential, or regulated information?
- Exposure: which policies, contracts, laws, and sector rules apply?
NIST's AI RMF Playbook emphasizes mapping the intended purpose, users, operating context, impacts, limitations, and alternatives. It also notes that narrow scopes are generally easier to map, measure, and manage than open-ended systems.
Compare candidates with evidence
| Candidate | Volume | Access | Loss | Owner | Evaluation | Risk |
|---|---|---|---|---|---|---|
| Support-ticket triage | 5 | 4 | 4 | 5 | 5 | Medium |
| Invoice-field extraction | 4 | 5 | 4 | 5 | 5 | Medium |
| Proposal first draft | 3 | 4 | 3 | 4 | 3 | Medium |
| Autonomous pricing decision | 3 | 4 | 5 | 4 | 3 | High |
The scores above are illustrative. In many SMEs, ticket triage or invoice extraction makes a stronger first pilot than autonomous pricing because the correct output is easier to evaluate and the human approval boundary is clearer.
Define the smallest useful pilot
- One trigger: for example, a new email in a specific inbox.
- One outcome: classify it, extract agreed fields, and prepare a routed task.
- One source set: use only approved documents and systems.
- One owner: name who approves cases and reviews failures.
- One evaluation set: test normal cases, edge cases, and known failures.
- One authority level: begin read-only or draft-only.
- One decision date: decide to stop, revise, or expand based on evidence.
OpenAI's current agent guidance recommends agents for workflows involving complex decisions, difficult-to-maintain rules, or heavy reliance on unstructured data—and explicitly notes that deterministic solutions may be sufficient otherwise. Your first AI use case may therefore be a conventional workflow with one AI classification or extraction step. That is often a strength, not a compromise.
Metrics to record before the pilot
- cases per week or month;
- active handling time and total cycle time;
- manual handoffs and repeated data entry;
- backlog and response delay;
- error, exception, and rework rate;
- percentage of outputs accepted without correction;
- human review time after automation.
A first pilot succeeds when it produces decision-quality evidence—not when it demonstrates the maximum possible autonomy.
Common selection mistakes
- choosing the most visible process instead of the most measurable one;
- automating a broken process before clarifying policy and ownership;
- using broad production data before permissions and minimization are defined;
- measuring model output without measuring the end-to-end workflow;
- expanding authority before reviewing real exceptions.
Sources and further reading
- OECD — Empowering SMEs in the age of AI: 2026 D4SME Survey
- NIST AI RMF Playbook — Map
- NIST AI RMF Core
- OpenAI — A practical guide to building agents
Review the 15-process opportunity library, compare chatbots, workflows, and agents, then document your candidate using the system-mapping workbook.
Select with evidence
Score one real workflow before buying another AI tool.
I help SMEs turn operational bottlenecks into narrow, governed, measurable AI pilots.
Map your first use case →