AI Automation

Scope an AI Automation Project So You Do Not Buy a Demo

Scope an AI Automation Project So You Do Not Buy a Demo

Opening answer (BLUF)

If you cannot name the process, the success metric, the human checkpoints, and the data the system may touch, you are buying a demo. To scope an AI automation project, write those four items down, then run a 90-day pilot against one workflow with a kill date. A vendor walkthrough that classifies sample emails is a product tour, not a specification. The scope is the contract you give operations, security, and finance before anyone connects a model to a live system of record.

Why a demo is not a scope

U.S. Census Bureau Business Trends and Outlook Survey (BTOS) data from December 14, 2025 through May 3, 2026 show overall business AI usage between 17 percent and 20 percent (19.8 percent in the period ending May 3, 2026), with 37 percent of firms that have at least 250 employees reporting use in operations.[1] A Census working paper on the 2026 BTOS AI supplement (November 2025 through January 2026) found that 18 percent of firms used AI in a business function (32 percent employment-weighted), and that 57 percent of those users integrate AI in three or fewer functions.[2] Most use is still narrow.

Eurostat reported that in 2025, 20.0 percent of EU enterprises with 10 or more employees used AI technologies, up from 13.5 percent in 2024.[3] The OECD, using 2024 figures, put AI implementation at 13.5 percent of enterprises in EU27 countries and 13.9 percent in the OECD area, with 39 percent of large firms using AI compared with 12 percent of small firms.[4] Adoption is rising. Depth is not.

RAND researchers, reporting in August 2024 from interviews with 65 data scientists and engineers with at least five years of AI and machine learning experience, identified five leading root causes of failure: misunderstanding the problem, lacking the data, chasing new technology instead of a user problem, lacking infrastructure, and applying AI to problems it cannot solve. The same report notes that, by some estimates, more than 80 percent of AI projects fail, twice the rate of IT projects that do not involve AI.[5]

In a July 2025 review, GAO found that across 11 selected agencies with AI inventories, reported AI use cases nearly doubled from 571 in 2023 to 1,110 in 2024, while generative AI use cases rose from 32 to 282. Officials at 10 of 12 selected agencies said existing federal policy, such as data privacy policy, could present obstacles to adoption.[6] That is the conversation a plant in Charlotte, NC, or a shared-services team in Philadelphia, PA, has when a vendor asks for production data on day one.

Start with a process map

NIST's Artificial Intelligence Risk Management Framework 1.0 (January 2023) organizes work into four functions: Govern, Map, Measure, and Manage. The Map function establishes the context that frames risks related to an AI system, including intended purpose, capabilities, stakeholders, and impacts. After governance is in place, NIST expects most users to start with Map.[7] That is the process map, not a slide of model names.

Walk the live path, not the policy path: who starts the work, what arrives, which systems are read and written, where work waits, who handles exceptions, and what "done" looks like in the system of record. Time a sample of real cases. Count the exception types. If the map is "invoices go in and payments go out," you do not yet have a map. A usable map names objects, people, systems, and the exception branch the demo never shows. That branch is where most project risk lives.

NIST notes that AI actors in one part of the lifecycle often lack full visibility over other parts, and that early decisions about purposes can alter later behavior.[7] If operations, finance, and IT each describe a different process, the model will be optimized for the wrong one. RAND found that trained models are often deployed that have been optimized for the wrong metrics or do not fit the overall business workflow.[5] For a manufacturer with operations in Raleigh, NC, and finance in another city, the map must name plant codes, company codes, and approval limits.

One success metric, measured before go-live

The NIST Measure function uses quantitative, qualitative, or mixed-method tools to assess and monitor AI risk. NIST states that AI systems should be tested before their deployment and regularly while in operation, and that measurements include documenting functionality and trustworthiness.[7]

GAO's AI Accountability Framework (June 2021) organizes accountability around four principles: governance, data, performance, and monitoring. Under governance, it calls for clear goals and objectives so intended outcomes are achieved.[8] Those 2021 practices still apply to a 2026 pilot.

Pick one primary metric operations already understands: cycle time from receipt to posting, first-pass match rate, exception rate, or hours of clerk time per 100 cases. The pilot succeeds or fails on that number. Baseline it on the current process before the model is connected. If you cannot produce a baseline, you cannot claim improvement. Do not use "accuracy" as the only number unless you define the unit (field, document, or case) and the gold-standard labeler. Write the threshold in the scope. "Make AP faster with AI" is not a metric.

Human checkpoints as design, not afterthought

NIST Map subcategory 3.5 requires that processes for human oversight are defined, assessed, and documented. Govern 3.2 requires policies that define and differentiate roles for human-AI configurations and oversight. Appendix C states that human roles in decision making and overseeing AI systems need to be clearly defined, and that configurations can span fully autonomous to fully manual.[7]

NIST's Generative AI Profile (NIST AI 600-1, July 2024) names confabulation (confidently stated but erroneous content) and human-AI configuration as risks unique to or exacerbated by generative AI. It warns that automation bias, excessive deference to automated systems, can worsen confabulation and bias, especially in consequential decisions.[9] A scoped project decides, in writing, where a person must see the output before it becomes an action.

CISA and international partners, in April 2024 joint guidance on deploying AI systems securely, recommend human-in-the-loop as a failsafe for advanced deployments and advise organizations not to run models immediately in the enterprise environment.[10] In May 2026, CISA and partners published Careful Adoption of Agentic AI Services: avoid granting broad or unrestricted access (especially to sensitive data or critical systems), begin with low-risk and non-sensitive use cases, and account for agentic AI in the existing security model.[11]

Put that into the scope as named gates: the model may classify and draft, a clerk must accept before the ERP write; amounts above a dollar threshold route to the controller; a new vendor never auto-posts; a kill switch stops writes if exception rate exceeds the baseline. Name the role, the SLA, and the queue. If no one owns the queue, the checkpoint does not exist. The scoping question is which actions the pilot may take without a person.

Data access is part of the brief

RAND's second root cause of failure is that the organization lacks the data needed to train or run an effective model.[5] GAO's 2021 framework treats data as its own principle: use data that are appropriate for the intended use of each AI system.[8] In July 2025, officials at 10 of 12 selected agencies told GAO that existing federal policy, such as data privacy policy, could present obstacles to generative AI adoption.[6] Commercial teams hit the same wall when a vendor asks for production files on day one.

CISA's 2024 deployment guidance tells organizations to identify and protect proprietary data sources, maintain a catalog of trusted sources, apply access control, encrypt sensitive AI information at rest, and secure exposed APIs with authentication, authorization, and input validation.[10] The 2026 agentic guidance is sharper: do not grant broad or unrestricted access, especially to sensitive data or critical systems.[11]

A scoped brief answers, before the demo, what the system may read, what it may write, where copies live, and who can approve an expansion. List the objects (for example, closed invoices from the last 24 months, vendor master without bank details). Use a dedicated integration account, not a shared clerk login. If personal data is in scope, name the legal basis and the minimization rule.

GAO's April 2026 review of federal AI acquisitions found that selected agencies were not yet systematically collecting lessons learned, including contract terms related to data rights or testing requirements. Industry reportedly invested over $250 billion in AI in 2024 alone.[12] Put data rights and a labeled test set in the 90-day scope: a non-production extract, a rule that vendor models do not train on your data unless you say so, and labels from the people who do the work today. An operations lead in Asheville, NC, who cannot get a clean extract of last year's exceptions does not have a model problem. They have a scoping stop.

A 90-day pilot with a kill date

RAND recommended that leaders be prepared to commit each product team to a specific problem for at least a year, and that if a project is not worth that commitment it is probably not worth starting.[5] A 90-day pilot does not contradict that. It is the decision gate before the year-long commitment: one process, one metric, one environment, and a keep-or-stop date.

Days 1 through 30: freeze the process map, success metric, human checkpoints, and data-access list. Pull the baseline. Stand up a non-production environment. Confirm identity, least-privilege access, and logging. Do not connect writes to production.

Days 31 through 60: run on historical and then live-shadow cases. Compare outputs to gold-standard labels. Measure the primary metric and the exception mix. Exercise the human gates until the queue has an owner and an SLA.

Days 61 through 90: limited production writes on a bounded set (one vendor group, one ticket queue, one plant). Keep the kill switch staffed. At day 90, compare to the threshold. If it is not met, stop or recast. If it is met, write the year-one plan, including monitoring. NIST Manage 4.1 calls for post-deployment monitoring plans, including mechanisms for capturing input from users.[7]

Kill criteria belong in the scope. Examples: no labeled baseline by day 21; cannot restrict the integration account to the agreed objects; exception rate worse than baseline for two consecutive weeks; no named owner for the human queue. CISA's advice to start with low-risk, non-sensitive use cases is the right first pilot, not a permanent ceiling.[11] The day-90 packet is the map, the metric with baseline and result, the checkpoint log, the data inventory, and a keep-or-stop recommendation. That is how you scope an AI automation project so the next dollar buys production work instead of another tour.

Practical takeaways

  1. Refuse to start from a demo. Require a process map that names objects, systems, people, and exception branches before a vendor connects to anything live.[7][5]
  2. Write one primary success metric with a baseline and a 90-day threshold. Test before deployment and on a schedule after.[7][8]
  3. Put human checkpoints in the scope: role, queue, SLA, dollar or risk thresholds, and a kill switch.[7][9][10]
  4. Treat data access as a first-class requirement: what is read, what is written, which identity is used, where copies live, and whether the vendor may train on your data.[5][10][11][12]
  5. Run 90 days as a gated pilot on one workflow, then decide on a year-long production commitment. Start with a low-risk slice, do not grant broad access, and capture lessons so the next process does not repeat the first mistakes.[5][11][12]

How we can help

Our team at Idea Forge Studios scopes AI automation the same way we would audit it: process first, metric second, human gates and data access before any production write. We work with operations and technology leaders in Charlotte, NC, Raleigh, NC, Asheville, NC, and Philadelphia, PA, on 90-day pilots with a kill date. Our AI business tools work sits behind that same discipline.

Have more questions or want to get in touch?

https://ideaforgestudios.com/contact-us-idea-forge-studios/ · (980) 322-4500 · [email protected]

Citations

  1. U.S. Census Bureau, "Large Firms With at Least 20 Employees Biggest AI Users" (2026)
  2. U.S. Census Bureau, Center for Economic Studies, "The Microstructure of AI Diffusion: Evidence from Firms, Business Functions, and Worker Tasks" (2026)
  3. Eurostat, "20% of EU enterprises use AI technologies" (2025)
  4. OECD, "Emerging divides in the transition to artificial intelligence" (2025)
  5. RAND Corporation, "The Root Causes of Failure for Artificial Intelligence Projects and How They Can Succeed: Avoiding the Anti-Patterns of AI" (2024)
  6. U.S. Government Accountability Office, "Artificial Intelligence: Generative AI Use and Management at Federal Agencies" (2025)
  7. National Institute of Standards and Technology, "Artificial Intelligence Risk Management Framework (AI RMF 1.0)" (2023)
  8. U.S. Government Accountability Office, "Artificial Intelligence: An Accountability Framework for Federal Agencies and Other Entities" (2021)
  9. National Institute of Standards and Technology, "Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile" (2024)
  10. CISA / NSA AI Security Center, "Joint Guidance on Deploying AI Systems Securely" (2024)
  11. CISA, "CISA, US and International Partners Release Guide to Secure Adoption of Agentic AI" (2026)
  12. U.S. Government Accountability Office, "Artificial Intelligence Acquisitions: Agencies Should Collect and Apply Lessons Learned to Improve Future Procurements" (2026)
Our Strongest Offering

Forge Your Next Website

Forged Sites are custom-built, static-first websites with a full AI content engine on board — no CMS to log into, no plugins to break, no builder to fight.

  • Near-perfect PageSpeed scores, static-first architecture
  • ADA + WCAG 2.2 AA accessibility, built in and re-checked on every deploy
  • MOG, an AI Site Director, lives inside your site and deploys changes in minutes
  • DraftDash auto-drafted blogs keep your content engine running
  • Ethel AI-powered forms filter spam and capture genuine leads