Problem
Many AI pilots begin with a tool and end with a demo. The team can show that something works, but cannot say whether it improves a decision, reduces a complete cost or deserves a place in the operation.
Thesis
A useful AI experiment does not ask, “Can it do this?” It asks, “Which decision changes, with what evidence and at what supervision cost?” Fourteen days are enough to find signal when the scope is narrow.
Framework
Design the test around four elements:
- Decision: the specific moment you want to improve.
- Baseline: how it is handled today, including human time.
- Threshold: the result that would justify continuing.
- Boundary: the harm, error or cost that stops the test.
| Phase | Days | Output |
|---|---|---|
| Map | 1–3 | Decision, cases and baseline |
| Test | 4–10 | Result and exception log |
| Decide | 11–14 | Continue, correct or stop |
Why it matters now
NIST’s AI RMF organises risk work around govern, map, measure and manage. A startup does not need a corporate programme to apply that logic: every experiment needs context, measurement, an owner and a response to observed risk.
Anti-example
The team compares two models with twenty prompts, chooses the one that “sounds better” and integrates it. Nobody records false positives, review minutes or affected decisions. The demo wins; the operation inherits uncertainty.
Protocol (3 steps)
- Write a one-page brief: user, decision, permitted data, metric and stop condition.
- Test real cases and boundaries: include normal, ambiguous and adversarial examples.
- Hold a decision review: continue only when the improvement exceeds total cost and observed risk.
Related
Sources consulted
Next step
Turn the next “we should use AI for…” into an experiment brief. If it does not fit on one page, the problem is not defined tightly enough yet.