Across 108 enterprises, trust in automated agent evaluation rose sharply in July — and the failure rate it is supposed to predict did not move at all. The share of organizations that fully trust automated evaluation nearly tripled, from 5% in June to 13%, and the complaint that evaluations don’t match real-world outcomes fell 10 points. Yet the same share as last month — just under half — shipped an agent that passed its evals and then failed a customer. The reason is visible in the cross-tabs: the new trust belongs almost entirely to enterprises that have not yet been burned. Among those that have, 4% trust automated evaluation; among those that haven’t, 24% do. And getting burned does not slow the march to autonomy — it speeds it up.

This is the second wave of the VentureBeat Pulse Research agent reliability tracker, and the first fielded on an instrument identical to the month before it. That makes July the first read on direction rather than position: what moved, what held, and what the movement means.

What moved is confidence. In June, only 5% of enterprises said they fully trusted automated evaluation, and the most-cited limitation was that evaluations align poorly with real-world outcomes (29%). In July, 13% fully trust automated evaluation and the alignment complaint has fallen to 19%, no longer the leading objection. Both shifts are large enough to read as real rather than noise.

What held is the failure. Just under half of organizations (49%) deployed an agent or LLM feature in the past year that passed internal evaluations and then caused a customer-facing failure — statistically indistinguishable from June’s 50% — and a quarter (24%) have seen it happen more than once. Confidence improved; correctness did not. That is the July gap: not between autonomy and trust, as in June, but between trust and the evidence for it.

The cross-tabs explain where the new confidence comes from, and it is not from better evaluations. Trust is concentrated almost entirely among enterprises that have not experienced a false-confidence failure: 24% of them fully trust automated evaluation, against 4% of those that have. The trust curve is being lifted by inexperience. Meanwhile the enterprises that have been burned are not retreating from autonomy — 85% of them already allow zero-human deployment or are engineering toward it, against 61% of those that have not been burned. Overall the autonomy trajectory is flat at 67%, but the population inside it has shifted toward the organizations with the most direct evidence that evaluations miss things.

The vendor market, by contrast, is finally showing signs of settling. The share of enterprises running no dedicated evaluation tooling fell from 17% to 12%; specialist platforms gained, with Braintrust nearly doubling to 15% and DeepEval reaching 17%; and switching intent cooled, with those planning no change rising from 36% to 44%. Selection criteria moved with it: ease of integration overtook cost as the top factor, jumping from 27% to 39%. Enterprises are done shopping on price and have started buying on fit.

Methodology

VentureBeat fielded this survey as part of its ongoing Pulse Research series. This wave — the agentic reliability and evals tracker — examines how technical leaders evaluate agent performance and reliability. Responses are filtered to organizations with 100 or more employees (n=108), drawn from a July 2026 fielding. Because the July instrument is identical to June’s, this report makes month-over-month comparisons where they are warranted; where questions were multiple-select, shares can sum to more than 100%.

Comparisons against June (n=157) are tested for significance, and only a handful of the month’s movements clear a conventional threshold: the rise in full trust in automated evaluation (5% to 13%), the fall in the real-world-alignment complaint (29% to 19%), the jump in ease of integration as a selection factor (27% to 39%), and the gain in Braintrust as a primary platform (8% to 15%). Movements described in this report as flat — the failure rate, the autonomy trajectory, the production monitoring mix, the investment ranking — are statistically indistinguishable between waves, and that stability is itself the finding. Differences of a few points elsewhere should be read as sample variation, not trend.

By role the sample is senior and buyer-credible: 44% are final decision-makers for AI purchases and another 25% recommenders or influencers, a slightly more senior mix than June. Product and program managers (18%), consultants and advisors (12%), CIOs/CTOs/CISOs (11%), and directors of engineering/IT (11%) lead the named titles, alongside a large “Other” function (30%). By organization size the sample is again mid-market-weighted: 100–499 (33%) and 500–2,499 (30%) employees lead, with 2,500–9,999 (23%), 10,000–49,999 (9%), and 50,000+ (5%) above them.

One composition change is worth flagging because it bears on the trust finding. The industry mix shifted between waves: Technology/Software fell from 23% of the June sample to 14% in July, while Retail/Consumer rose from 15% to 19% and now leads. A less technology-weighted sample plausibly carries less hands-on exposure to agent evaluation, and some of the month’s rise in trust may reflect who answered rather than what changed. The burned-versus-unburned split reported in Finding 2 holds within the July sample regardless, but readers should treat the headline trust movement as directional.

At 108 respondents the sample is large enough to support directional conclusions but should not be treated as a precise measurement; it is self-selected and is not a probability sample. Cross-tabs reported here rest on subgroups of 40 to 68 respondents and are correspondingly coarse.

Finding 1: The failure rate did not move

Just under half still ship agents that pass evals and fail customers

We asked whether, in the past 12 months, organizations had deployed an agent or LLM feature that passed their internal evaluations but then caused a customer-facing failure. The answer is the same as last month.

Finding 1 — The Failure Rate Did Not Move

Forty-nine percent of organizations shipped an AI feature that cleared internal evaluations and then failed in front of a customer — an incorrect output, a broken workflow, or a quality incident — against 50% in June. A quarter (24%) have seen it happen more than once, unchanged. Across two waves and 265 enterprises, the rate at which evaluations certify agents that then fail is stable to within a percentage point.

That stability is the anchor for everything that follows. Every other movement this month — rising trust, consolidating tooling, shifting purchase criteria — has to be read against a failure rate that has not responded. Whatever enterprises did between June and July, it did not change how often a passing evaluation turns out to be wrong.

Finding 2: Trust rose — among those who haven’t been burned

Full trust nearly tripled, and the alignment complaint fell ten points

We asked which limitation most reduces trust in automated agent evaluations today. The distribution shifted materially from June.

Finding 2 — Trust Rose — Among Those Who Haven’t Been Burned

Two things moved together: Full trust in automated evaluation nearly tripled, from 5% to 13%, and the objection that most directly describes a false-confidence failure — poor alignment with real-world outcomes — fell from 29% to 19%, surrendering the top spot to evaluation bias and inconsistency (22%), now tied with data-leakage concerns (22%). On the surface this reads as an evaluation layer beginning to earn its keep.

The cross-tab says otherwise. Splitting the sample by whether an organization has actually experienced a false-confidence failure, trust divides almost completely. Among the 53 enterprises that shipped an agent which passed evals and then failed a customer, 4% fully trust automated evaluation. Among the 41 that have had no such failure, 24% do — a six-fold difference, and the sharpest split in the dataset. Direct contact with the failure mode is what removes the trust.

This is the month’s central caution. The improvement in sentiment is not evidence that evaluations got better; the failure rate in Finding 1 rules that out. It is what a trust curve looks like when a cohort of less-burned organizations enters the sample and reports its priors. Enterprises reading their own rising confidence as validation of their evaluation stack are reading a number that measures inexperience.

Finding 3: Being burned accelerates autonomy rather than restraining it

85% of the burned are on the zero-human path, against 61% of the REST

We asked whether organizations would let an autonomous agent deploy a code or system change to production on automated evaluation results alone, with no human-in-the-loop validation. The aggregate held; the composition did not.

Finding 3 — Being Burned Accelerates Autonomy Rather Than Restraining It

At the top line, nothing changed: 67% of organizations either already allow zero-human-in-the-loop deployment for low-risk agents (37%) or are actively engineering their pipelines to permit it within a year (30%), against 67% in June. The share ruling it out for the foreseeable future slipped from 22% to 18%. The autonomy ceiling stopped rising, but it did not come down.

Underneath, the picture inverts the intuitive one. Among enterprises that have shipped an evaluation-passing agent that then failed a customer, 85% are on the autonomy trajectory. Among those that have not, 61% are. Organizations with direct, expensive evidence that their evaluations miss things are substantially more likely to be removing the human check, not less — and only 11% of them rule out full automation, against 24% of those that haven't been burned. The pattern is identical for those burned once and those burned repeatedly.

The most plausible mechanism is not recklessness but maturity: the organizations that ship agents at enough volume to hit a customer-facing failure are the same ones with pipelines sophisticated enough to automate, and they are treating the failure as a cost of operating rather than a reason to stop. That is a defensible read. It is also precisely the dynamic that turns Finding 1’s stable failure rate into a growing absolute number of incidents, since the enterprises most likely to fail are the ones scaling their capacity to deploy without review.

One June finding did not replicate. Last month, larger enterprises appeared slightly further down the autonomy path than smaller ones (70% versus 64%). In July the two converge — 65% for organizations with 2,500+ employees against 68% below that, with near-identical failure rates (48% and 50%) — which suggests the June gap was sample variation rather than a size effect. Company size is not what separates the aggressive adopters; experience of failure is.

Finding 4: The stack begins to consolidate

Specialists gain, and the “Nothing at all” share shrinks

We asked which agent reliability or evaluation platform enterprises primarily use today. The field is still crowded, but it is no longer tied at the top with nothing.

Finding 4 — The Stack Begins to Consolidate

The most consequential number is the one that fell. In June, having no dedicated agent-evaluation tooling was tied for the most common answer at 17%; in July it is 12% and fifth. Enterprises are acquiring evaluation tooling, and the specialists are capturing most of that movement: Braintrust nearly doubled its share of primary usage to 15%, and DeepEval reached 17%. Provider-native tooling held roughly flat — OpenAI at 18%, Anthropic at 12% — meaning the growth came at the expense of running nothing rather than at the expense of the model providers.

Counting any use rather than primary platform, the footprints are wider and the ordering is similar: OpenAI native evals reach 31% of enterprises, DeepEval 27%, Braintrust 22%, Anthropic native evals 20%, custom in-house tooling 14%, and Weave and Langfuse 11% each. Nineteen percent still report using no dedicated tooling anywhere in their stack. The category now has three plausible independent contenders where in June it had none with double-digit primary share — the first evidence in this series of an evaluation layer starting to take shape.

Finding 5: Production monitoring still watches the wrong thing

Half monitor whether the agent runs; under a third monitor whether it’s right

Production monitoring for an AI agent can watch two very different things. It can watch whether the system is functioning — is the agent up and responding, did each request complete, how fast, at what cost, with any errors. Or it can watch whether the agent’s output is correct — automated checks that evaluate the content of each answer as it goes out. A confidently wrong answer is invisible to the first kind: the request completes, the response is fast, no error is thrown, and every functioning-metric reads healthy. We asked which kind live production monitoring is built for today.

Finding 5 — Production Monitoring Still Watches the Wrong Thing

Grouped by what is actually being watched, the split is essentially June’s: 50% of organizations monitor only whether the agent is functioning, while 26% run automated checks on whether its answers are right. Counting ad-hoc reviewers and don’t-knows, nearly three-quarters of organizations have no automated, real-time evaluation of output correctness in production. Inline quality assertions and transaction trace logging are tied as the most common approach at 26% each on a base of 106 — no single monitoring posture leads.

This is the finding that most directly contradicts the month’s rising confidence. Trust in automated evaluation went up eight points while the runtime capacity to detect an evaluation being wrong went nowhere. Among enterprises that already permit zero-human deployment, only 28% run inline quality checks on production traffic — which means the majority of organizations that have removed the human from the deployment decision have also not replaced that human with anything watching output quality afterward. The gate is automated and the alarm is not installed.

Finding 6: Bought on fit now, not on price

Ease of integration overtakes cost as the top selection factor

We asked what most influenced enterprises’ choice of an evaluation vendor, and what they treat as their primary measure of success. One answer moved sharply; the other did not move at all.

Finding 6 — Bought on Fit Now, Not on Price

Ease of integration jumped 12 points to 39% and displaced cost as the leading selection criterion, the clearest purchasing shift in the data. Evaluation accuracy rose modestly to 28%, cost fell to 23%, and breadth of observability (6%) and vendor roadmap (2%) remain marginal. Read alongside Finding 4, the two move together: enterprises adopting their first dedicated evaluation tooling are optimizing for what will slot into an existing pipeline this quarter, not for what is cheapest or most capable in the abstract. That is what a market looks like when it stops evaluating and starts installing.

What did not move is what enterprises want from the tool once installed. Evaluation consistency remains the primary success metric at 38%, essentially identical to June’s 36%, well ahead of reduction in failures (20%), speed of experimentation (18%), production visibility (16%), and compliance (7%). The priority is still repeatability — the same verdict on the same behavior every time — which is notable given that bias and inconsistency is now the top-cited trust limitation in Finding 2. Enterprises are buying for integration and measuring for stability, and are not yet getting the second. Satisfaction with current tooling remains moderate, averaging 3.9 on a five-point scale across overall satisfaction, ease of implementation, and value for money, barely changed from June’s 3.8.

Finding 7: Human review becomes the top line item

And the enterprises that have been burned fund it hardest

We asked which reliability and evaluation investment will grow most over the next year. Human review edged into first place.

Finding 7 — Human Review Becomes the Top Line Item

Human review workflows (31%) and production observability (30%) swapped positions at the top, a change small enough to be noise on its own — but the underlying pattern is the same one June identified and it has strengthened. Enterprises plan to grow spending on human reviewers faster than on the automated evaluation pipelines (19%) that would replace them, at the same moment two-thirds are engineering the human out of the deployment decision. Only 6% report a flat budget, down from 8%.

The cross-tab makes the hedge explicit. Among enterprises that have shipped an evaluation-passing agent that failed a customer, 38% name human review as their fastest-growing investment; among those that have not, 24% do, and they favor observability tooling instead. So the burned cohort is doing both things at once: it is the most aggressive on autonomy (85% on the zero-human path, per Finding 3) and the most committed to funding human reviewers. That is not a contradiction so much as a strategy — automate the deployment decision, and pay people to catch what the automation misses. Whether that scales is the open question, since human review is the one part of the stack that does not get cheaper as agent volume grows.

Finding 8: The switching wave cools

Those planning no change rise from a third to nearly half

We asked whether enterprises plan to adopt a new, additional, or replacement evaluation platform, and which they are considering. Fewer are shopping than last month.

Finding 8 — The Switching Wave Cools

A majority (56%) still intend to adopt a new, additional, or replacement platform within twelve months, but that is down from 64%, and the near-term cohort thinned from 31% to 24%. The share standing pat rose from 36% to 44%. Neither movement clears a significance threshold on its own, but both point the same direction, and they point it consistently with Finding 4: as enterprises actually acquire tooling, the population still looking for it shrinks.

The consideration set has reordered, too. Among the 60 enterprises planning a change, OpenAI’s native evals lead what they are evaluating (20%), followed by Braintrust (18%), Weights & Biases Weave (12%), and DeepEval (10%), with a further 10% actively evaluating but holding no shortlist. DeepEval led June’s consideration set at 20%; it has since converted much of that interest into primary usage, which is what a consideration-to-adoption handoff looks like. Braintrust now occupies the position DeepEval held — high interest ahead of installed base — and is the vendor to watch in the next wave.

The bottom line: Confidence moved, correctness didn’t

June found a gap between the autonomy enterprises were granting their agents and the trust they placed in the evaluations meant to govern it. July finds that gap closing from the wrong side. Trust rose — full confidence in automated evaluation nearly tripled and the complaint that evaluations miss reality fell ten points — while the thing that trust is supposed to track held exactly still. Just under half of enterprises still ship agents that pass their evals and then fail a customer, the same as last month.

The cross-tabs locate the new confidence precisely, and it is not in the evaluations. Twenty-four percent of enterprises that have never had a false-confidence failure fully trust automated evaluation; 4% of those that have do. Trust in this market is a function of exposure, not of evidence. And exposure does not produce caution: the burned cohort is the most autonomous in the sample, with 85% already deploying without human review or building toward it. What it produces instead is a hedge — the same organizations fund human review workflows hardest, at 38%, while removing humans from the deployment gate.

The vendor market is the month’s genuinely encouraging story. Running no dedicated tooling fell from 17% to 12%, specialists gained real share for the first time in this series, buyers shifted from price to integration fit, and switching intent cooled as adoption completed. An evaluation layer is finally forming. But the runtime picture has not followed: half of enterprises still monitor only whether their agents are running, and among those that already deploy without human review, just 28% run real-time checks on output quality.

At 108 respondents in a mid-market-weighted, self-selected sample, and with an industry mix that shifted away from technology between waves, this is a directional read. The direction, though, is legible: enterprises are tooling up, buying for fit, and growing more confident — and none of that has yet changed how often a passing evaluation turns out to be wrong. The question this series carried out of June was whether assurance would catch up to autonomy. July’s answer is that confidence caught up first, which is the harder problem, because an enterprise that trusts a broken gate has less reason to fix it than one that knows the gate is broken.

This report presents the July 2026 wave of an ongoing longitudinal series on enterprise AI agent reliability and evaluation, based on 108 qualified respondents at organizations with 100 or more employees. Comparisons are drawn against the June 2026 wave (n=157), fielded on an identical instrument. At this sample size, results should be read as a directional signal rather than a precise measurement — the sample is self-selected, not a probability sample. Respondents span final decision-makers, technology recommenders/influencers, and end business users, across a mid-market-weighted range of industries and company sizes.