Beyond the pilot: A framework for scaling AI safely and efficiently

Presented by Intel Intel's logo

As federal agencies move from AI experimentation to enterprise-wide adoption, turning isolated pilots into production-ready capabilities remains a major challenge. Success requires a unified strategy and infrastructure that connect fragmented workflows and support a consistent path from testing to deployment.

That need for a coordinated approach is becoming more urgent as AI technology advances. Agentic AI offers powerful new capabilities, but it also introduces operational and financial complexity. Token consumption can rise quickly, increasing costs, while disconnected workflows add friction and make solutions harder to scale.

Addressing these challenges requires agencies to look beyond model size and performance and design an architecture that can scale AI efficiently. Rather than defaulting to a cloud-only strategy, agencies can use a hybrid architecture to place each workload in the environment best suited to its cost, performance and security requirements.

Barriers to operationalizing AI

When agencies begin to move AI out of sandbox environments, they often encounter systemic operational and data hurdles. Dr. Darren Pulsipher, chief enterprise architect for public sector at Intel, noted that the barriers are rarely about AI models themselves.

Pulsipher said some of the largest barriers are process-related: Agencies create promising experiments but often lack the repeatable processes needed to move a pilot into production and expand it across the enterprise.

Data practices in the experimental phase can create another barrier. For security reasons, agencies often isolate data during a pilot. When the solution moves toward production, however, those isolated datasets can become silos that do not align with enterprise data structures or standards.

“They grab all the data and put it in a container,” Pulsipher said. “Then when they go, ‘We want to do this in production,’ all of a sudden, the data isn’t the same, there’s no standard schema, it’s all unstructured.”

Without a repeatable operating model, agencies struggle to turn proofs of concept into enterprise value. Industry research underscores the gap: An MIT study found that 95% of AI pilot projects failed to deliver measurable impact. Pulsipher added that agencies often struggle to explain how they measure the success of their pilots.

“It might do really cool stuff, but if you can’t find business value, and you can’t measure that business value, then why are you doing it?” he said.

A framework for success: The AI Augmented Operating System (AAOS)

To create a bridge from isolated experiments to repeatable success, Pulsipher outlined a six-stage framework referred to as the AI Augmented Operating System:

  1. Diagnose: Assess the current state of the organization across five domains of change: strategic, organizational, procedural, digital systems and physical infrastructure. Leaders should determine whether the project supports the agency’s strategy, whether the organization is structured for the change and whether employees have the skills needed to adopt the new capability.
  2. Activate: Translate the assessment into a deployment plan and redesign processes as needed. This stage identifies training requirements, assigns responsibilities and defines how the organization will introduce the new capability.
  3. Control: Establish governance for cybersecurity, regulatory compliance, data privacy and AI system identities. AI systems should never use an employee’s credentials. Each system should have a distinct identity with access that can be time-limited or revoked.
  4. Execute: Move into active production. Once governance, plans and technical controls are in place, the agency deploys the AI solution.
  5. Measure: Evaluate results against the key performance indicators established during planning. These measures should focus on mission and operational outcomes rather than raw usage statistics, such as token consumption.
  6. Scale: Expand proven capabilities across the enterprise. At this stage, agencies apply lessons from the previous five stages to strengthen return on investment and make each subsequent deployment faster and more repeatable.

The Diagnose stage is especially important because it helps agencies identify where AI can create the greatest practical value. Technologists should begin by examining existing workflows, particularly processes that require substantial manual effort, and look for opportunities to use AI to support employees rather than replace them.

Pulsipher framed the goal as using AI to reduce the burden on overwhelmed employees without removing human judgment. “The decision still needs to be with the human, but if the human is gathering data from lots of different places to make a decision, this is a perfect spot for AI,” he said.

GPUs vs. CPUs: Optimize costs through a hybrid approach

Alongside the AAOS framework, agencies need an infrastructure strategy that supports long-term scale without driving unnecessary cost. Earlier cloud-first initiatives encouraged organizations to move as many workloads off-premises as possible. Today, agencies are taking a more selective approach, placing workloads according to their performance, security and cost requirements.

Pulsipher described this evolution: “After we did ‘cloud first,’ then we were, ‘cloud smart.’ But now we’re seeing the costs can be so high that we’re repatriating. I call it ‘cloud wise.’”

The same discipline should guide how agencies use AI models. Some organizations treat token consumption as a sign of adoption and even encourage employees to use more tokens. Pulsipher cautioned that this shortsighted approach rewards activity rather than measurable organizational value.

“Everyone should be leveraging AI, but we have to really think about how we're leveraging it,” he said. “I don't need the latest model to do a lot of work, because it's small stuff that's augmenting what I do. You don’t need to have all this power sitting in this big GenAI that you don’t need but you’re paying for it.”

One important part of that decision is choosing between GPU- and CPU-based infrastructure. Large, cloud-hosted generative AI models may require substantial GPU capacity, but not every task needs that level of compute. For narrower use cases, smaller private models may deliver the required results at lower cost on existing CPU infrastructure or endpoint devices.

Smaller private models can also be tailored to the needs and context of individual users. Pulsipher noted that many can run on a laptop, reducing the cost to the underlying compute. By capturing relevant user and workflow context within a controlled environment, these models can become more productive and reliable for focused tasks.

Grounding AI in cybersecurity fundamentals

As agencies scale AI deployment, security and compliance remain central to long-term resilience. Rather than requiring an entirely new security approach, scaling safely depends on reinforcing foundational cybersecurity principles such as zero-trust architectures, strict identity management and time-bounded access controls.

“AI is a magnifying glass,” Pulsipher said. “If you already have good security, if you’re following zero trust principles already, then AI is going to shine. … If you don’t, or if you’re doing it halfway, it’s going to expose you like there’s no tomorrow.”

Private generative AI models hosted on enterprise laptops, workstations or internal servers can provide additional security and governance benefits. Local processing can help agencies keep sensitive, regulated or classified data within organizational boundaries and maintain greater control over AI assets. When paired with appropriate safeguards, this approach can reduce exposure to public-cloud environments while avoiding some of the costs associated with large public models.

Scaling AI successfully isn’t about replacing human expertise or expanding compute budgets indiscriminately. It’s about augmenting human capabilities and deploying the right tools in the right environment to build a resilient, mission-ready ecosystem.

Learn more about how Intel can help your organization scale AI efficiently and effectively.

This content is made possible by our sponsor Intel. It is not written by and does not necessarily reflect the views of GovExec's editorial staff.

NEXT STORY: Social Security Timing & Medicare Maximization