Infino unifies data retrieval for AI agents
Infino proposes a single retrieval layer over Parquet so that agents can query structured and unstructured data without multiple systems.

The way AI agents query data doesn't look like how humans or traditional applications do it. While a user launches a query, waits, and reads the result, an agent formulates many small questions to solve a big one, often in parallel, and each answer consumes context, adds cost and latency to subsequent model calls. That difference in behavior is what has led Infino, a company founded by OpenSearch veterans, to launch a data retrieval platform designed specifically for agents.
The diagnosis made by its CEO, Ekechi Nwokah, is clear: agents are the largest consumer of data since the web browser, but they are still served by a fragmented stack built for another era. Data warehouses for SQL, search engines for keywords, vector databases for semantic search, and ETL pipelines to keep everything synchronized. Infino proposes collapsing that stack into a single layer that allows agents to search, sort, filter, join, aggregate, and reason directly on the data.

A single reader, a single copy of the data
The technical key to the proposal lies in how data is stored and queried. Infino stores a single copy in Apache Parquet format on object storage, with search indexes embedded alongside the file footer. This way, any tool that reads Parquet can access the data with or without Infino. On that basis, the platform offers an interface that combines structured and unstructured queries: ranked searches, exact counts, joins, filters, and aggregations at low cost and large scale.
The idea is not for the agent to learn SQL, but for it to express a complete question in a single query. According to Infino, that eliminates pipelines, ETL jobs, schemas to keep aligned, and glue code to combine results from multiple systems. Furthermore, having a single copy of the data centralizes governance: which rows and columns an agent can see, who can read what, and what is logged.
The real problem: retrieval loops
During development, the Infino team encountered a challenge that goes beyond query speed. Even when retrieval is fast, agents spend a lot of time in loops: formulating a query, searching, deciding if the results are sufficient, trying something else, validating the answer, and repeating. To address this, the platform integrates specific inference models that perform these simple tasks faster and more economically than frontier models.
The core engine is open source under the Apache-2.0 license and is available on GitHub. Infino Cloud is the hosted service the company offers as a managed alternative.
What it means for an SME or an IT team
For an SME starting to build agents, the promise of a single retrieval layer is attractive because it reduces operational complexity. Instead of maintaining a data warehouse, a search engine, and a vector database with their respective pipelines, it points to a single system that the agent queries. That can translate into lower infrastructure costs, less engineering time on integrations, and simpler governance.
However, caution is advisable. Infino's proposal is recent and not without risks: depending on an emerging provider for a critical piece of the data architecture requires evaluating its maturity, its community, and its roadmap. Additionally, migrating to a single-copy model in Parquet may require rethinking how current data is ingested and transformed.
As we have seen in other blog articles, agent autonomy and the governance of their accesses are hot issues. In AI in the SOC: how much autonomy to cede without losing control? we already analyzed the importance of limiting what an agent can do without supervision. And in Graph RAG: when relationships are the key evidence we explored how data structure conditions the quality of responses. Infino fits into that conversation: it's not just about storing, but about enabling the agent to retrieve the right information at the lowest possible cost.
Recommendations before taking the leap
- Evaluate your current stack: if you already have a data warehouse and a vector database working, calculate the real cost of maintaining them versus a unified layer.
- Test with a limited case: before migrating everything, experiment with a specific agent and measure latency, cost, and response quality.
- Review governance: a single copy of data simplifies permissions, but requires clearly defining what each agent can see and what is logged.
- Don't lose sight of the format: having data in standard Parquet is an advantage to avoid dependencies, but make sure your current tools support it.
Infino's proposal points to a real problem: agents don't query like humans and deserve infrastructure designed for them. For an SME, the decision involves weighing operational simplification against provider maturity and adoption effort. As always, starting small and measuring is the wisest path.
Source: The New Stack. Analysis and adaptation: ForgeNEX.