BVG Insights
AI ROI Without the Guesswork: Baselines, Small Proofs, and Stop Rules
The most important number in an AI proposal is usually the one that existed before the tool arrived.
Without a baseline, a faster draft can be mistaken for a better workflow. A new subscription can be mistaken for adoption. A few promising examples can be mistaken for a business case.
The missing comparison is not another market statistic. It is the current operating path: volume, elapsed time, active work, review and correction, and the result the workflow is supposed to produce.
That is why “What is the ROI of AI?” is too broad to answer responsibly. A decision-grade question names the work: Can this bounded change improve one observable workflow result after setup, review, and correction are counted?
Adoption is context, not a return calculation
Published adoption figures answer different questions about different populations. The Federal Reserve’s 2026 synthesis explains that estimates vary with sampling, unit of analysis, weighting, question framing, and whether “use” covers any business function or only production of goods and services. Those measures should not be averaged into one market rate.
Two dated examples show why the distinctions matter. In the NFIB’s 2025 member survey, 24 percent of sampled small employers reported AI use. In a separate 2026 survey of selected Goldman Sachs 10,000 Small Businesses participants, 76 percent reported AI use while 14 percent reported full integration into core operations. The definitions and samples differ. The figures belong inside their own boundaries; they are not inputs to a blended adoption claim, and neither establishes ROI.
JPMorganChase Institute transaction research found that paid AI-service adoption accelerated and varied by firm characteristics through 2025. It measures paid transactions—not total use, integration quality, workflow value, or return. A purchase is observable behavior. Whether it improved the business is a separate question.
Use external adoption evidence to understand the market conversation. Build the decision from the company’s own workflow baseline.
Build the baseline before choosing the success story
A baseline does not need to be a perfect accounting model. It needs to be stable enough to compare the current path with a bounded test.
Start with one workflow that has a visible beginning and end. “Use AI in marketing” is too loose. “Turn an approved interview transcript into a first-pass article for human review” is testable. So is “classify incoming service requests into existing routing categories before a dispatcher confirms them.” The workflow should happen often enough to observe and matter enough to justify attention.
Record a compact set of measures:
- Volume: How many items enter during the chosen period?
- Elapsed time: How long does an item take from start to accepted finish?
- Active work: How much human effort does the current path consume?
- Review and correction: What must be checked, fixed, or redone before the output is usable?
- Quality threshold: What makes the result acceptable for this workflow?
- Operating result: What happens next when the work is accepted?
Keep output measures separate from business outcomes. Producing a first draft sooner is an output change. Publishing more often is another output. Qualified demand, retained revenue, or lower service cost would be business outcomes requiring their own evidence and time window. A proof should not leap from the first category to the second.
Put all of the work inside the test boundary
Nominal task time is only one part of the cost.
The test also consumes setup, prompt or rule design, access control, training, exception handling, fact checking, privacy review, correction, and maintenance. Some of that work is front-loaded. Some returns every time the workflow runs. If the owner or a senior employee becomes the permanent verifier, the tool may have moved effort rather than reduced it.
Human review is not a footnote added after the demonstration. It is a named workflow step with an owner, a standard, and a measured burden. State which errors are tolerable, which require correction, and which make the workflow unsuitable for the tool.
When testing an AI workflow, count possible time savings alongside accuracy, setup, complexity, and review burden. These considerations inform wording and test design, not prevalence or error-rate claims. Their useful contribution is the question they force: what correction job does the demonstration leave out?
Write the stop rules while enthusiasm is cheap
A proof becomes easier to interpret when the decision rules exist before the best example appears.
Give the test three possible endings:
Stop when a protected condition is crossed: an unacceptable privacy exposure, a critical error class, review effort above the agreed ceiling, or no meaningful movement in the leading indicator. Stopping is a valid result. The proof prevented a larger commitment.
Change when the workflow still matters but the first design is wrong. Narrow the input set. Move the human review point. Improve the source material. Use an existing tool differently. Test a non-AI process change. The next run should isolate what the first run taught.
Scale only when the bounded workflow meets its quality threshold, the full attention cost is acceptable, the result persists across enough ordinary cases, and the next increment has a named owner. Scaling one workflow is not permission to generalize across the company.
These criteria preserve reversibility, but reversibility is a ranking dimension—not an entry condition. Some material workflows are difficult to reverse and still deserve analysis. That risk should raise the evidence threshold and tighten the controls, not hide behind a small-pilot label.
An illustrative proof card
Consider a company that prepares a weekly internal operations summary. This is a test design, not a reported result.
Workflow: Convert approved source notes into a first-pass internal summary.
Baseline window: Four ordinary weekly cycles using the current process.
Leading indicator: Time from complete source notes to an accepted first pass.
Quality threshold: Every factual statement traces to an approved source; required sections are present; a named reviewer accepts the summary.
Full effort counted: Setup, source preparation, generation, review, correction, and exception handling.
Stop: A protected source is exposed, a critical claim cannot be traced, or reviewer burden crosses the pre-agreed ceiling.
Change: The first pass is useful only for one section or source type, so the workflow narrows.
Scale: The bounded process meets the threshold across the agreed sample, and the next volume increase has an owner and review capacity.
The card does not forecast savings. It defines the evidence required before making a larger claim.
Do not ask one metric to carry the whole decision
Even a clean test can create conflicting evidence. Cycle time may improve while correction work rises. Quality may become more consistent while setup remains too specialized. The tool may help a skilled reviewer and frustrate a novice. Keep those tradeoffs visible.
Use a four-column ledger: observation, interpretation, confidence, and next proof. “Reviewer corrected six unsupported statements” is an observation in a hypothetical test. “The model is unreliable for our work” is an interpretation. It still needs context: sample size, source quality, error severity, and whether the test design can change.
One partner-fielded 2025 survey found that its Explorer segment wanted clearer ROI evidence, easier tools, and practical training. The sample covered businesses with reported revenue from $25,000 to $5 million and involved Reimagine Main Street and PayPal, so it is directional adoption-support evidence—not a universal buyer finding. Its relevance here is the question it sharpens: what proof would make this workflow decision clearer?
A bounded next step
Use the proof-card exercise in this article to define one workflow baseline, its human-review burden, and its stop, change, or scale rules. For a broader inventory, get the Owner’s Field Guide.
The honest path to AI ROI is not a more confident forecast. It is a baseline another person can inspect, a small proof that counts the hidden work, and decision rules strong enough to stop a tool that does not earn its place.
Source notes
- Federal Reserve Board, “Monitoring AI Adoption in the U.S. Economy.” Read the source. Source date: 2026-04-03. Retrieval date: 2026-08-27. Caveat: adoption estimates differ by sample, unit, weighting, definitions, and framing; they must not be averaged.
- NFIB, “2025 Small Business and Technology Survey.” Read the source. Source date: 2025-06-25. Retrieval date: 2026-08-27. Caveat: NFIB member sample and a definition that differs from other studies; the 24 percent figure is one bounded estimate only.
- Reimagine Main Street/PayPal, “Beyond Efficiency: Small Businesses Look to AI for Competitive Edge.” Read the source. Source date: 2025-06-10. Retrieval date: 2026-08-27. Caveat: partner/vendor involvement and a $25,000–$5 million revenue sample; directional adoption-support evidence only.
- Goldman Sachs, “Small Businesses Embrace AI but Need Training and Support to Fully Harness It.” Read the source. Source date: 2026; no exact publication day was recorded in the source review. Retrieval date: 2026-08-27. Caveat: selected 10,000 Small Businesses participants, not a representative small-business sample; use only for the reported use/integration gap within that sample, never as ROI evidence.
- JPMorganChase Institute, “Understanding AI Use by Small Businesses.” Read the source. Source date: 2026-04-14. Retrieval date: 2026-08-27. Caveat: paid transactions measure paid adoption behavior, not total use, integration quality, or ROI.
Use the illustrative proof card
Illustrative example, not a reported result.
Open live proof card