Insights

The Data Silo Problem Isn’t the Silos
October 6, 2026
Reading Time: 4 minutes

By Anand Trivedi, AI Lead at REI Systems

NIH recently put a spotlight on a challenge facing nearly every agency pursuing AI: the data AI needs is often spread across systems, programs, organizations, and environments.

For NIH, the stakes are particularly high. Its new Bio Genesis Mission aims to double the pace of biomedical innovation over the next five to 10 years by bringing AI, advanced computing, data, and scientific infrastructure together in new ways. But as NIH Assistant Director for Catalytic Data Resources Chris Kinsinger recently emphasized, AI alone won’t make that happen. Researchers first need to be able to find, access, and use data distributed across NIH’s 27 institutes and centers.

That raises a larger question for government: If AI needs data from many places, how much infrastructure do we need to build around that data before we can actually put it to work? And how quickly and at what costs?

NIH is federating data. What comes next?

NIH has already moved beyond the idea that solving the data silo problem requires putting everything into one giant repository. Its Common Fund Data Ecosystem, or CFDE, links data platforms maintained by individual programs and provides researchers with a central way to discover and query across them. NIH is also developing cloud workspaces where researchers can access and compute across datasets. The underlying data can remain distributed.

That is an important evolution: federation instead of wholesale consolidation. But federation and orchestration are not quite the same thing.

A federated ecosystem creates the standards, metadata, portals, infrastructure, and coordination that allow distributed resources to be found and used together. Orchestration adds an active operating layer on top of those distributed resources.

Instead of stopping at “Where is the data, and how can I access it?” orchestration asks: What data do I need for this question? How do these sources relate? What needs to be transformed or reconciled? What rules govern their use? And how do I execute that work across systems?

What if the data could play in place?

Think about an orchestra. The strings don’t need to move into the percussion section. The brass doesn’t need to become woodwinds. Each section stays where it belongs, with its own role, characteristics, and requirements. But simply knowing where each section sits isn’t enough to create music. Someone has to bring them in at the right moment, interpret the score, coordinate their contributions, and keep the whole performance working together.

Government data environments aren’t all that different.

Mission data may live in agency systems, cloud platforms, document repositories, operational databases, analytics environments, or specialized research platforms. Different owners may govern it. Different security rules may apply. Schemas and definitions may not line up neatly. A federated environment can make those resources easier to find and reach. An orchestrated environment would make them work together.

Federation connects the players. Orchestration conducts the performance.

GovOrch™ AI is REI Systems’ governed data orchestration platform, designed to help agencies work across distributed data environments without first centralizing everything into a single repository.

It connects to data where it already lives and creates an AI-enabled governed orchestration layer across existing systems. It can discover data and schemas, reconcile structures and definitions, build cross-system data workflows, enforce policies, and turn a natural-language question into a governed, executable pipeline.

GovOrch connects to data where it already lives and creates a governed orchestration layer across existing systems. It can discover data and schemas, reconcile structures and definitions, build cross-system data workflows, enforce policies, and turn a natural-language question into a governed, executable pipeline.

The goal isn’t simply to give a user a doorway into several data sources. It is to coordinate what happens across those sources so the user can get to an answer.

Consider the difference: A researcher or program leader wants to answer a question that depends on information in six different systems.

  • A federated approach can help identify those resources, establish common ways to describe them, and provide access across them. But data retrieval and cross-tabulations would still need manual work.
  • An orchestration approach can go further: determine which sources are relevant, understand how their structures relate, assemble the necessary workflow, apply governance, execute across the sources, and return an answer with a traceable path back to the underlying data.

Silos themselves aren’t necessarily the enemy. Disconnection is.

There are often good reasons for data to remain distributed. Different systems serve different missions. Security boundaries matter. Data ownership matters. Existing investments matter. And not every dataset needs to be copied, standardized, or relocated simply because an AI application may need to use it.

The more useful question is whether an agency can empower its mission users to securely reach the right information, understand it in context, reconcile it when necessary, apply appropriate controls, and use it alongside information from other authoritative sources. If it can, physical co-location becomes much less important.

That matters even more as agencies move from AI pilots toward AI that participates in real mission work. A model pointed at one repository can only work with the context inside that repository. Mission decisions often require context and data distributed across operational systems, financial platforms, case records, documents, analytics environments, and other sources, carried forward in a state-aware way across workflows.

AI doesn’t just need data access. It needs coordinated, governed access to the right data at the moment the mission requires it.

Don’t automatically make data integration step one

NIH’s Bio Genesis Mission is an ambitious example of a broader shift across government. Agencies increasingly recognize that their AI ambitions are inseparable from their data architecture.

NIH’s CFDE demonstrates one important part of the answer: distributed data can be made interoperable and accessible without putting everything into a single repository. The next evolution is to ask how much of the work of integration can happen dynamically.

Some data will still need to be standardized. Some will need to move. Some systems will need modernization. And some domains will benefit from purpose-built data ecosystems like CFDE. But agencies shouldn’t automatically assume that every new AI use case requires another major data integration effort before useful work can begin.

Sometimes the better question is: Can we orchestrate what we already have?