
Enterprise AI teams appear to be settling on platforms for “agentic” systems faster than they are solving the harder problem of making those systems trustworthy in production. A set of four June 2026 VentureBeat Pulse Research surveys, each based on separate samples of enterprise respondents, points to the same conclusion from different angles: the bottleneck is no longer choosing a base platform, but deploying agents with reliable orchestration, governed context, credible evaluation, and basic security controls.
The sharpest signal comes from the orchestration wave. According to VentureBeat’s survey of 101 organizations with more than 100 employees, enterprises are concentrating agent orchestration on model-provider platforms, with Anthropic’s Claude cited as the primary platform by 40% of respondents, followed by Microsoft at 18% and OpenAI at 13%. But that consolidation sits beside an uncomfortable admission: 71% said a quarter or fewer of their deployed “agents” are actually multi-step orchestrated workflows rather than chatbot wrappers.
Taken together with three companion surveys on context infrastructure, evaluation, and security, the message is broader than one vendor’s lead. Enterprises say they want reliable, autonomous, multi-step systems. In practice, many are still shipping brittle assistants, using provider-native defaults, and discovering failures after deployment. For AI builders, enterprise buyers, and founders selling into this market, that shifts the conversation from model choice to operational discipline.
VentureBeat’s orchestration survey argues that enterprises are choosing platforms largely because of “model gravity,” meaning the pull of the underlying foundation model rather than a distinct orchestration control plane. In that sample, the large providers dominated while open frameworks such as LangChain/LangGraph and custom in-house builds remained in single digits for primary usage.
That matters because enterprises also told VentureBeat that orchestration success is judged mainly by task completion reliability and multi-step workflow management. Those priorities suggest buyers are looking for systems that can coordinate tools, maintain state, and complete work over several steps. Yet the same respondents acknowledged that most current deployments do not meet that standard.
This gap between ambition and production reality is central to the story. If most deployed systems are still single-prompt assistants, then enterprise spending on orchestration is partly preemptive infrastructure: teams are building the layer before the portfolio of truly orchestrated agents exists. That is not necessarily irrational. It may reflect a normal enterprise pattern of standardizing procurement and controls early. But it does mean a lot of “agent” adoption figures likely bundle together very different kinds of systems.
The survey also found that 68% plan to adopt a new, additional, or replacement orchestration platform within a year, with many still lacking a shortlist. That suggests this layer is far from settled, even if a few vendors currently dominate primary deployments.
The other three VentureBeat surveys help explain why enterprises are struggling to move from chatbot wrappers to dependable agents.
On context, a separate 101-respondent survey found that 57% of enterprises had traced a confident but wrong AI-agent answer to missing or inconsistent business context in the past six months. Retrieval-augmented generation is already the main context source for many teams, and VentureBeat reported that provider-native retrieval tools such as OpenAI file search and Google Vertex AI Search now lead dedicated vector databases in production usage within its sample.
That result does not mean retrieval is solved. The same survey suggests the opposite: most enterprises are still building the governed semantic layer needed to make retrieved context consistent and trustworthy. VentureBeat found that 58% either run or are building such a semantic layer, but only 25% have one in production. In other words, AI agents are sounding authoritative before the underlying business definitions are stable.
On evaluation, the numbers are similarly sobering. In VentureBeat’s 157-respondent survey, 50% said they had shipped an AI feature that passed internal evaluations and then failed in a customer-facing setting. Only 5% said they fully trust automated evaluation today. Even so, 66% either already allow fully automated deployment for low-risk agents or are engineering toward it within a year.
That combination may be the clearest sign that enterprise AI operations are outrunning their assurance practices. Teams want faster release cycles and more autonomy, but the evidence in this survey suggests the test layer is not yet strong enough to support that safely.
On security, VentureBeat’s 107-respondent survey found that 54% had already experienced either a confirmed agent security incident or a near-miss. The structural issue, according to the report, is identity and containment. Only 32% said every agent gets its own scoped, managed identity, while many still rely on shared credentials. Just 30% isolate their highest-risk agents in sandboxes.
That is especially relevant as enterprises grant agents real access to internal systems. Without per-agent identity and isolation, a compromised or over-permissioned system can have a much wider blast radius.
A consistent pattern across all four surveys is that provider-native tooling is becoming the default operating choice. In orchestration, model-platform vendors dominate primary usage. In context, OpenAI file search and Google Vertex AI Search lead the retrieval stack in VentureBeat’s sample. In security, OpenAI guardrails and cloud-native controls from major providers were the most common tools. In evaluation, provider-native tools were also at the top, tied with having no dedicated tooling at all.
That pattern matters for the market because it says convenience is beating specialization in the first wave of enterprise deployment. Buyers appear to be choosing what ships with their model platform or cloud stack, especially when teams are still early in production maturity.
But VentureBeat’s own data also suggests those defaults are provisional. In security, 59% planned to adopt or switch tooling within a year. In context, 57% planned to switch or add a retrieval provider. In evaluation, 64% planned to adopt a new or replacement platform. In orchestration, switching intent was also high.
So while the current stack looks concentrated, the buying cycle is still open. Enterprises may start with bundled controls and retrieval because that reduces friction, but many do not appear convinced they will stay there. Vendor lock-in was the top risk cited when orchestration control sits inside a model-provider platform, which helps explain why 51% in the orchestration survey expected a hybrid control plane by the end of 2026.
All four source items are VentureBeat Pulse Research surveys from June 2026, not public filings or audited market data. That gives them value as a directional read on enterprise sentiment and deployment behavior, but also imposes real limits.
The samples were self-selected, mid-market weighted in several cases, and modest in size: 101 respondents for orchestration, 101 for context, 157 for evaluation, and 107 for security. VentureBeat itself says the findings should be treated as directional rather than precise measurement. The surveys also used different respondent groups, so readers should be cautious about treating percentages across waves as directly comparable market-share figures.
That caution is especially important around vendor standings. Anthropic’s Claude leading orchestration in this sample does not establish broader market share. The same applies to OpenAI file search, Google Vertex AI Search, OpenAI guardrails, or provider-native evals. These are survey-based snapshots of primary usage or presence in the stack among AI-active respondents, not spend-weighted rankings.
Still, the repeated structural pattern across separate surveys is harder to dismiss. Different enterprise groups reported similar tensions: bundled tools lead current usage; specialist categories remain fragmented; and most organizations plan to change or add tooling within a year. Even if the exact percentages move, that pattern looks credible.
For AI product teams, the practical lesson is that “agent” claims will face heavier scrutiny. If enterprise customers increasingly recognize that many deployed agents are just wrappers around prompts, vendors will need to show where orchestration actually exists: state management, tool coordination, runtime controls, fallback behavior, and measurable task completion.
For enterprise buyers, these surveys suggest that platform selection is becoming the easy part. The hard part is operationalizing enterprise AI across four layers at once: a trusted context layer, an evaluation layer tied to real-world outcomes, an identity and isolation layer for security, and a control plane that can manage both costs and vendor dependence.
For startups, the opening may be narrower than broad “agent platform” messaging implies. Provider bundles are soaking up the default entry point. The opportunities appear more specific: governed semantic layers, non-human identity, runtime isolation, output-quality monitoring, and orchestration controls that sit above the model vendor.
The cost angle also deserves attention. VentureBeat’s orchestration survey found that 27% had no real-time, programmatic way to stop a runaway agent before the bill arrived. That is not only a FinOps issue. It is a sign that runtime control planes are still immature.
The most important follow-up signal is whether enterprises start reporting growth in genuinely multi-step production workflows, not just more “agent” deployments. If that number remains low while orchestration spending rises, the category may be over-labeled and under-delivered.
Second, watch whether hybrid control architectures become the norm. If teams continue standardizing on Anthropic’s Claude, Microsoft, or OpenAI while adding external orchestration and policy layers, the market may settle into a split stack rather than winner-take-all platforms.
Third, pay attention to whether governed semantic layers move from pilot to production. If OpenAI file search and Google Vertex AI Search keep growing without better context governance, the confident-but-wrong problem is likely to persist.
Fourth, the evaluation layer bears watching. If enterprises continue moving toward automated deployment while provider-native evals and fragmented tools remain the norm, customer-facing failures may become a major procurement trigger.
Finally, security spending and architecture will be the real test of seriousness. If OpenAI guardrails remain the default while scoped identity and sandboxing stay uncommon, enterprises may be normalizing exposure rather than reducing it.
These surveys point to a maturing enterprise AI market in which the language of agents has moved faster than the operating reality. The key takeaway is not that enterprises picked the wrong platforms. It is that too many teams treated platform choice as the main decision when the deployment challenge lives in context quality, runtime control, evaluation fidelity, and security design.
That is why the most interesting competitive battleground may not be the base model at all. It may be the control layers around it: the systems that decide what an agent knows, what it can touch, how it is tested, when it is stopped, and who can prove what happened afterward. Until those layers harden, many enterprise “agents” will remain expensive assistants with better branding than operational autonomy.
VentureBeat surveys suggest enterprises are rushing AI agents into production, but gaps in orchestration, context, evals and security remain.