Insights

REI Systems AI FAQ: Cost Control, Governance and Scaling AI in Government
September 30, 2026
Reading Time: 7 minutes

AI costs and value

What should we include in the full cost of an AI initiative?

We budget for the entire life of the capability: data preparation and governance, model or service fees, computing, integration, security, testing, specialist and program staff, training, and monitoring after launch. We estimate expected volume and cost per task as well as the initial build. A pilot budget is rarely a reliable estimate for agency-wide operations. Explore our AI cost guide.

Why can AI costs rise after a successful pilot?

A pilot usually serves a small group with limited data and predefined use cases. Production adds users, system connections, support, monitoring, retention, and more frequent model calls. Basically, adding both scale and scope (in some cases) between pilot to production.  To counter that, we measure actual usage in the pilot, model likely production workload variations compared to pilots thorugh in-depth stakeholder consultations, and revisit costs before expanding. Explore our AI cost guide.

How can we control AI spending without sacrificing performance?

We test the least costly design that meets the mission’s quality, security, and response-time needs. Model routing is core to our AI solution design. That means not using the best frontier models for every task, and switching to smaller models for routine work, fewer repeated calls, appropriate storage tiers, or batch processing when an immediate answer is unnecessary. We track both cost and performance so a cheaper approach does not quietly reduce the value of the result. Read our cost-control perspective.

What role does FinOps or TokenFinOps play in AI?

We use FinOps and TokenFinOps practices to make variable AI costs visible and assign them to the use cases, models, or teams generating them. Mission, technology, and finance leaders can then review usage, forecasts, and cost per useful outcome together. This is especially important when inference and data costs grow with adoption. Spending data should inform design and scaling decisions throughout the capability’s life. Explore our AI cost guide.

Should we build an AI solution or buy an existing service?

We compare both paths against the agency’s data, security, integration, performance, and mission needs. An existing service can shorten the path to a pilot, but licensing, customization, usage fees, and dependence on a vendor remain part of its lifetime cost. A custom approach may offer more control while requiring more engineering and maintenance. We make that tradeoff explicit before committing. Explore our AI cost guide.

How do we know whether AI is worth the investment?

We define the outcome and baseline before selecting technology. Depending on the mission, we may measure time to complete a case, answer quality, avoidable contacts, staff effort, or the cost of serving a user. We compare the AI approach with other ways to improve the work, then use measured results to decide whether to expand, revise, or end the initiative. Read our perspective on AI economics and outcomes.

Governance and public trust

What AI governance should be in place before deployment?

We build on your existing data, cybersecurity, and AI governance. For each use case, we identify an accountable owner, permitted data and actions, testing criteria, human review points, audit records, and a way to handle incidents. Those controls should reflect what an error could do to the mission or the public and continue after launch. Read our governance framework.

Does every AI use case need the same level of oversight?

No. We scale oversight with the system’s autonomy and the consequences of its output or action. An isolated prototype using test data calls for different controls than a production agent acting on live records. We reassess risk as the use case changes, particularly when it begins affecting benefits, safety, rights, or other consequential outcomes. Read our governance framework.

How should governance change when AI agents can take action?

An agent that can update records, trigger a workflow, or contact someone needs limits on its actions as well as checks on its answers. We define permissions, action boundaries, approval gates, activity logs, and intervention procedures. We test those controls before connecting an agent to live systems and monitor behavior after deployment. Read our perspective on agentic AI oversight.

Where should people remain in control of AI-assisted decisions?

We believe in human-in-the-lead by design, not just human-in-the-loop. Human reviews stage-gate incrementally as the risks and consequences increase. AI can organize information, draft a response, or route routine work, while authorized staff handle exceptions and make or approve consequential decisions. We also define how staff can challenge an output, correct information, and escalate a problem. The right checkpoints depend on the program, action, and people affected. Read more about consequence-based oversight.

How do we oversee several AI agents working together?

We give each agent a defined role and permission boundary, then track the handoffs between them. We test what happens when their recommendations conflict, a source changes, or an upstream agent makes an error. Typically, an orchestrator in a multi-agentic system runs a validator or governance agent on the system events. Admist conflicting recommendations, we use LLM or agent-as-a-judge approach to resolve with an independently trained and run agent/model combination. Production design needs a way to see the full workflow, resolve conflicts, and intervene before a mistake spreads across systems. Read our governance framework.

What happens if an AI agent takes unexpected action?

We design intervention into the workflow. That includes alerts, clear ownership, a way to pause or stop the agent or connected workflow, and records that show what happened. We pre-define how staff contain an incident, review affected actions, correct records when needed, and decide when operation can resume. These procedures matter especially for agents that act across several systems. Read our governance framework.

How can we protect sensitive data in AI applications?

We begin with your data classifications, privacy obligations, access rules, and approved environment. We limit what users and AI components can retrieve, send, or change; apply appropriate security controls; and log access and actions. The design also accounts for data passed to models, tools, and connected services. We confirm those boundaries for each use case. Read our governance framework.

How do we test AI for errors, bias, and unfair outcomes?

We test against representative cases, difficult exceptions, and the consequences of a wrong result. For higher-impact uses, we examine whether outcomes vary across relevant populations, run adversarial or failure tests, and compare AI recommendations with staff decisions where appropriate. We monitor data and concept drift when the system is live in production as performance for AI systems can change over time. Read our governance framework.

Data and technology readiness

What makes agency data ready for AI?

We look for data that is accurate enough for the task, understandable in context, accessible to authorized users, and current enough for the decision. We assess quality, ownership, definitions, formats, and gaps between systems before choosing a model. Data readiness is specific to the use case: the records needed to answer a public question differ from those needed to support a case decision. Read more about production data readiness.

Must we move all our data into one system before using AI?

Data integration is not always the best solution. We actively compare intergration against interoperability as a way forward. We determine where the required data lives and how to provide governed access while preserving permissions and context. Connecting to existing sources can be more practical than moving everything first. The approach depends on data quality, interfaces, security, response time, and how the result will be used. Read more about governed data orchestration.

How can users check where an AI answer came from?

We design for traceability when the use case calls for it. That means preserving source references, the relevant version of a document or rule, and records of the data and steps used to produce an answer or action. We test whether staff can verify the output and recognize when evidence is incomplete. Traceability matters most when AI informs a consequential recommendation or decision. Read more about trusted operational data.

Can AI work with legacy systems and disconnected data?

Yes, if the needed information can be accessed reliably under the right controls. We evaluate interfaces, data definitions, permissions, and the work needed to reconcile records across systems. An API, data virtualization, or another integration pattern may help, depending on the environment. We test those connections with real mission questions before treating a pilot as production-ready. Read more about the data barriers behind AI pilots.

From pilot to production

How should we choose our first AI use case?

We start with a specific mission problem, an accountable owner, usable data, a measurable baseline, and a clear way to review the result. We also assess the harm an error could cause and whether a simpler solution would meet the need. A focused use case makes it easier to learn, govern the work, and decide whether broader deployment is justified. Read our guidance on defining an AI problem.

How do we test an AI agent safely before production?

We begin in an isolated setting with limited permissions and appropriate test or de-identified data. We set success and failure criteria, log activity, and have subject matter experts review outputs. A later experiment can use more realistic conditions with stronger oversight and explicit boundaries. We connect an agent to live operations only after testing its behavior, controls, and response to failures. Read our phased governance framework.

Why do government AI pilots struggle to reach production?

A pilot can work with a narrow dataset and controlled demonstration. Production must handle actual records, changing workflows, access rules, legacy systems, staffing, and support. We address those dependencies early by testing cross-system access and data quality, defining ownership and oversight, estimating operating costs, and planning for user adoption. A convincing demo is only one part of readiness. Read more about moving beyond AI pilots.

What should an agency ask for when acquiring an AI capability?

We recommend defining the mission outcome and the evidence needed to evaluate it. The requirement should also address data access, security, permitted actions, human oversight, auditability, integration, performance, operating costs, support, and how the agency can respond when the system fails or changes. We work with program, technology, security, and acquisition teams to turn those needs into testable acceptance criteria for the specific procurement. Read more about outcomes and governance.

How do we measure AI performance after launch?

AI performance postlaunch is not an afterthought. While business-outcomes are already validated in the pre-production stage, in production, we actively monitor for latency, costs, data and concept drifts, etc.  We compare results with the baseline and review failures and exceptions, not just averages. As policy, data, and workloads change, we check whether the capability still performs appropriately and adjust its workflow, controls, or model when needed. Read more about ongoing evaluation.

People and agency capability

How should we prepare our workforce to use AI?

We involve the people who know the mission work in use-case design and testing. We explain what AI will do, where staff retain authority, how to check an output, and how to report a problem. Training should reflect each role and continue as workflows change. Adoption improves when staff can see how the tool helps them serve the mission and can give feedback that shapes it. Read our AI in government framework.

Does an agency need an AI Center of Excellence?

An AI Center of Excellence can help an agency share expertise, set common practices, evaluate tools, and coordinate governance and training across programs. We recommend giving it a clear relationship with mission owners and existing data, security, and technology teams. Its structure should fit the agency’s size and maturity; the goal is consistent decisions and reusable learning. Read our AI in government framework.