How can a small business launch an AI pilot in 30 days? Select one frequent, bounded workflow with an accountable owner and accessible data. Measure its current performance, define acceptance and stop criteria, build the smallest controlled workflow in week two, test normal and adversarial cases in week three, and launch to limited users or volume in week four. At day 30, expand only if business measures and residual risk support it.

The goal is evidence for a decision. A pilot should tell you whether the workflow creates enough value, whether the organization can operate it, which exceptions require people and what must change before wider use. “The model produced a good answer” is one observation, not a pilot result.

Write the pilot charter before day one

Charter fieldDecision to record
WorkflowOne trigger-to-outcome process, including its system of record.
ProblemThe observable delay, correction work, missed service or capacity constraint.
OwnerOne person accountable for policy, acceptance, rollout and post-pilot operation.
Users and volumeThe limited group, case types and maximum production volume.
BaselineCurrent cycle time, throughput, correction, service, cost and risk measures.
BoundariesApproved data, tools, write actions, limits, review points and prohibited actions.
Decision criteriaNumeric and qualitative thresholds for expand, revise or stop.

Choose the workflow with the VALUE framework. Avoid a high-consequence process merely because it is visible. A document-classification, internal triage, preparation or recommendation workflow usually gives a small team more room to learn than autonomous payments, employment decisions or external commitments.

The four-week pilot roadmap

01Discover and baselineChoose one workflow; map inputs, decisions, tools, exceptions and controls; record current performance; name the owner; approve the pilot charter.
02Build the controlled prototypeConnect only the minimum approved data and tools; separate deterministic rules from AI judgment; add validation, approval and complete operation logging.
03Test and challengeRun a documented evaluation set with normal, edge, hostile and failure cases; compare results with acceptance thresholds; repair the workflow, not only the prompt.
04Launch under controlRelease to a small user group or limited case volume; monitor every exception; compare with baseline; decide to expand, revise or stop.
Four controlled stages: baseline, build, evaluate and limited launch.

Week 1: discover and baseline

Observe actual work, not only the written SOP. Capture the trigger, required inputs, business rules, judgment calls, applications, hand-offs, exception categories, approval points and outcome. Sample recent cases to quantify volume, handling time, waiting time, rework, error categories and service performance. Identify data ownership and access before assuming the prototype can use it.

Finish the week with a signed charter, current-state map, prioritized risk list, test-case plan and go/no-go review. If the workflow has no owner, no baseline or no lawful data path, pause. Discovery has already created value by preventing a blind build.

Week 2: build the controlled prototype

Implement the smallest end-to-end path. Use deterministic code for fixed validation and policy rules; use the model for bounded interpretation, extraction, classification, comparison or drafting. Connect only the necessary sources and narrow tools. Require previews and human approval before consequential actions. Record the input version, policy, model, output, tool call, user decision and result.

Do not add a second workflow to fill spare capacity. Use the time to make the first one observable and recoverable. The SOP-to-AI guide provides a practical decomposition method, while the security checklist covers identities, permissions and incident controls.

Week 3: test and challenge

Create a versioned evaluation set from representative historical cases and designed failures. Include incomplete, ambiguous, duplicate, stale and conflicting inputs; unauthorized requests; prompt injection inside retrieved content; tool errors; values just above thresholds; and cases that must escalate. Record expected outcomes before running the system.

Measure task success by category, exception recall, false actions, false blocks, evidence quality, human correction and recovery. NIST's AI RMF calls for testing before deployment and regularly in operation, under conditions similar to the deployment setting. NIST's AI Resource Center groups this as test, evaluation, verification and validation (TEVV), emphasizing that trustworthy operation cannot be inferred from a handful of demonstrations.

Week 4: launch under control

Limit the release by user, case type, volume, action or time window. Brief users on what the system can do, what it cannot do, how to verify output and how to report a problem. Monitor exceptions daily. Keep rollback, credential revocation and a manual fallback ready. Do not silently widen tool permissions to solve a blocked case.

At day 30, compare results with the charter. Expand only the specific action supported by evidence. Revise when value is promising but controls or performance are insufficient. Stop when economics are weak, risk exceeds tolerance, ownership is absent or integration and maintenance costs overwhelm the benefit.

A practical day-30 scorecard

  • Outcome: Did the workflow improve the agreed business measure against baseline?
  • Quality: What were success, correction and exception results by case type?
  • Safety: Were permissions, approvals, privacy and stop conditions effective?
  • Adoption: Could users understand, verify and recover from the system's work?
  • Economics: Do confidence-adjusted benefits exceed build and ongoing operating costs?
  • Ownership: Is there a funded owner for monitoring, evaluation, incidents and change?

Use the ROI calculator for the economic decision and the failure analysis when the pilot stalls.

Primary sources checked for this guide

Checked 26 August 2026. NIST's TEVV-Athlon document is an initial public draft, not final guidance; it is cited for its current evaluation framing.

One workflow. Real evidence.

Start a guided 30-day AI pilot

I help SME teams choose the workflow, establish the baseline, build controlled access, evaluate real cases and make an evidence-based day-30 decision.

Plan a 30-day pilot