A small group inside a CPA firm tests an AI use case and gets promising results. The tool saves time, the users see potential and leadership agrees it should expand. A few months later, the same handful of people are still the only ones using it.
Recent Thomson Reuters research found that 81% of tax and audit professionals use AI at least several times a week, yet 34% of senior leaders have not made significant changes to their operating models in response. Usage is moving faster than the structure needed to support it.
AI projects usually stall after the pilot because the firm has proved capability, not operability. Scaling requires the surrounding workflows, data, permissions, ownership, support and measurement to hold up under normal client conditions.
Pilots create conditions that are hard to reproduce
Pilots are supposed to be controlled. They usually start with a narrow use case, selected information and a small group of motivated users. When the tool produces a weak result, the project team can step in quickly.
Those conditions make it easier to test an idea, but they can also hide the work required for broader use. An AI-assisted tax research pilot may perform well when an experienced manager frames the questions, reviews every source and edits each response. Add more users, client files and deadlines, and the experience changes. Data quality varies, exceptions increase and informal support becomes harder to sustain.
McKinsey’s 2025 State of AI survey found that 88% of organizations use AI in at least one business function, but only about one-third have begun scaling their AI programs to match that scale.
Scaling AI requires the full CPA firm workflow
Most pilots test whether AI can complete a task. They do not always account for everything that happens before and after it.
An AI-assisted tax research tool may produce a strong first draft, but broader use requires the firm to define which sources are authoritative, what client information the tool can access, how citations are verified, where the output is stored and who approves the final position.
The same principle applies across tax, audit and advisory work. AI has to fit into the complete process, including intake, access, review, documentation and exception handling. Otherwise, each user develops a different method, making the results difficult to trust or repeat. CPA.com’s 2025 AI in Accounting Report identifies workflow automation and human verification as central considerations for accounting firms adopting AI.
AI ownership becomes more complicated in production
During a pilot, ownership often feels obvious. A partner champions the use case, IT manages the tool and early users provide feedback.
Production use introduces split responsibility. The service line owns the workflow and professional judgment. IT owns identity, access, integration and technical support. Risk or quality leaders set acceptable-use and review expectations. The technology vendor controls the product, including model updates, feature changes and service availability.
A practical ownership model should clarify three areas:
- Business ownership: Who is accountable for the workflow, intended result and professional standards?
- Technology ownership: Who manages access, integrations, administration and support?
- Risk and quality ownership: Who defines review requirements, monitors exceptions and has authority to pause use?
Those responsibilities need to be explicit. Otherwise, IT becomes the default owner of questions involving professional judgment, while service-line leaders assume the technology team is monitoring risks it may not be able to see.
This is one reason the harder part of AI adoption is often operational rather than technical. The firm needs clear decision rights around approving changes, resolving exceptions and determining whether the workflow still serves its intended purpose.
Production exposes data, permission and integration gaps
A pilot may rely on users manually uploading a few documents. Production use may require integration with a document management system, tax platform, audit application or Microsoft 365 environment.
That introduces questions the initial test may never answer:
- Does the tool inherit the firm’s existing identity and access controls?
- Can it distinguish final work from drafts, duplicate documents or outdated guidance?
- Are prompts, outputs and user activity logged appropriately?
- Can access be limited by client, engagement, role or service line?
- Who supports the workflow when an integration fails during a deadline?
- Can the firm suspend or reverse the rollout without disrupting client work?
Permissions deserve particular attention. An AI tool that can search across internal information may surface content a user would never know to look for manually. Production readiness depends on whether the underlying access model is accurate, not simply whether the AI product includes security features.
The same is true for data quality. Firms often find that fragmented data, inconsistent permissions and unclear ownership slow AI adoption long before the technology itself becomes the constraint. The National Institute of Standards and Technology’s voluntary AI Risk Management Framework treats AI risk management as an ongoing process across design, deployment, monitoring and use.
Adoption depends on how the work feels to users
Training matters, but training alone will not fix a workflow that creates more effort than it removes. Staff may avoid an approved tool because it requires extra steps, produces inconsistent results or adds review time. Partners may disagree about acceptable use. Reviewers may quietly redo AI-assisted work because they do not trust the output. Employees may continue using an unapproved tool because it performs better or is easier to access.
For AI to become part of everyday work, users need a process that is easier to follow, expectations they understand and enough experience to recognize where the tool performs well and where professional judgment must take over.
Measure the full workflow, not just output speed
Positive feedback from pilot users is useful, but it is not enough to justify broader investment. A firm needs to know whether AI improves the complete workflow, not simply whether it generates an output quickly.
The most useful measures may include total cycle time, reviewer effort, the percentage of outputs requiring substantial correction, exception volume, support requests, use within the approved process and cost per completed workflow.
A tool can save 15 minutes on a first draft and add 30 minutes of review. In that case, the pilot worked technically, but did not improve the work.
The 2026 Thomson Reuters AI in Professional Services Report found that only 18% of respondents said their organizations track AI return on investment. Another 40% did not know whether it was measured.
Use a production-readiness gate
Before an AI pilot moves into broader use, the firm should be able to confirm that:
- The complete workflow and intended boundaries are documented.
- Authoritative data sources are defined.
- Identity, permissions and retention requirements have been validated.
- Human review and quality thresholds are clear.
- Business, technology and risk owners understand their responsibilities.
- Support, monitoring and escalation processes are in place.
- The rollout can be paused or reversed without affecting client delivery.
- The pilot produced measurable improvement under realistic conditions.
The outcome may still be to revise the use case, run another test or stop altogether. Ending a weak pilot is not failure. It is considerably better than maintaining a permanent experiment that consumes licenses, support time and meeting agendas without improving client work.
From AI capability to operability
The firms that make progress will not necessarily be the ones running the most pilots. They will be the ones that can identify which use cases deserve to become standard work, then build the operating structure needed to support them.
For firms with a promising AI project that has not moved forward, the next step may not be another tool or another pilot. It may be a closer examination of the workflow, data, access and support requirements surrounding the use case. Netgain works with CPA firms to evaluate those dependencies and determine what is genuinely ready for broader use. Discuss your firm’s AI readiness with our team.
