The AI Agent Infrastructure Stack in 2026, Layer by Layer
The six layers under every production agent, who supplies each, and the three layers still soft enough that a buyer should be cautious.
Toolradar data: we track 341 MCP servers, 83% published by vendors on their own domains and 30% free. The tool layer is one of six, and the one MCP changed most.
Every production agent stands on the same six layers, whoever built it. This is the map, with who supplies what, from the 10,000+ tools we track, and where the stack is still soft enough to be worth caution.
The six layers
1. Models. Frontier APIs, or self-hosted open weights for bounded high-volume jobs (the decision guide). This is the commodity layer; everything above it is where systems actually differ, so do not over-invest attention here.
2. The gateway. One endpoint owning keys, routing, budgets and logs across providers: Portkey, Helicone and peers. What a gateway buys you turns on whether you have two-plus apps or providers; below that it is overhead.
3. Tools, standardised by MCP. The layer that changed fastest. 341 MCP servers in our catalog, 83% vendor-official, meaning agent access to a product is now configuration rather than integration. Start from the servers worth installing; govern them centrally with an MCP gateway once agents multiply.
4. Execution surfaces. Where the agent acts when there is no API: browser infrastructure (Browserbase, Steel, agents like Skyvern), and the newer computer-use surface where agents drive full desktops. This layer is the least mature of the six, which is worth knowing before you bet a product on it working reliably at scale.
5. Memory and state. Still the softest layer, and the one with no standard. Conversation state, work-in-progress state and long-term knowledge get conflated by most stacks, which is the root of agents that forget mid-task. The boring answer, a database the agent reads and writes plus retrieval for knowledge, outperforms most dedicated memory products today.
6. Observability and evals. Traces, cost attribution and scored test sets: LangSmith, Langfuse, Arize, compared in our evaluation guide. Below two agents this feels optional; above two it is the difference between debugging and guessing.
Cross-cutting all six: security, from prompt-injection defence (Lakera, Prompt Security) to agent identity and scoped credentials, covered in our agent security guide.
Where the stack is still soft
Three layers are not settled, and if you are choosing where to be conservative, choose these:
Memory has no winner and no standard. Every vendor conflates the three kinds of state differently, and the database-plus-retrieval baseline still beats most dedicated products. Do not marry a memory product yet.
Computer-use execution works in demos and struggles at scale. The reliability gap between "drove the desktop in a demo" and "drives the desktop unattended for a thousand users" is real and unclosed.
Identity is the gap the security vendors are racing to fill. What an agent is allowed to do, as distinct from what its human's credentials allow, is not yet a solved layer. An agent acting with a user's full permissions is the default and the wrong default.
How to read the map as a buyer
Buy commodity layers, build the job-specific top. The failure mode we see in the catalog data is the inverse: teams building their own gateway (a solved, commodity problem) and buying an off-the-shelf "agent" whose job does not fit their work. The layers below your agent are undifferentiated infrastructure; the job definition is the only part that is yours, so that is the only part worth building.
FAQ
Is this whole stack necessary for one agent?
No. One bounded agent needs a model, its tools and an eval set. The stack earns itself as agents and teams multiply; the build guide covers the minimum viable version.
What changes next?
Watch the tool layer finish consolidating (MCP already won it) and the memory layer produce a standard. Nothing else in the stack looks settled enough to predict with confidence.
Where should a small team start?
Model, tools via MCP, and evals. Add the gateway when you have a second app or provider, observability when you have a second agent, and everything else only when a specific pain demands it.
Related
From the team behind Toolradar
Growth partner for B2B tech
Toolradar also helps B2B tech companies grow, content marketing & distribution through 5 newsletters (720K+ tech professionals), AI Academy, and the Toolradar directory.
See how we work
Written by
Louis Corneloup
Founder & Editor-in-Chief at Toolradar. Founder & CEO of Dupple, the publisher of 5 industry newsletters reaching 720K+ tech professionals. Reviews B2B software using a public methodology, see /how-we-rate and /editorial-policy.