MLPerf, the AI industry's standard speed test, adds its first test of AI agents as a record 30 organisations take part

MLCommons published MLPerf Inference v6.1 with a record 30 submitting organisations and, for the first time, a test of AI agents doing multi-step work such as coding.

· 5 min read · L3 Iron and cloud · The Pulse, 16 September 2026

In this story: Nvidia · Lambda · Oracle · Microsoft Azure · CoreWeave · Nebius

What happened

On 16 September 2026 MLCommons published the results of MLPerf Inference v6.1, its industry-standard, peer-reviewed benchmark of how fast AI systems run. A record 30 organisations submitted results, among them the chipmaker Nvidia and the cloud providers Microsoft Azure, Oracle, CoreWeave, Nebius and Lambda. For the first time the suite includes a test of AI agents, the Edge Agentic Inference test. It measures multi-turn work such as agentic coding, where each step depends on the ones before, on a single device of the kind that serves one user at a time. MLCommons also added an end-to-end test of retrieval-augmented generation, a chain of several AI models that finds documents and then answers from them.

Why it matters

The numbers are useful because they are comparable. Entrants run the same tests under the same rules, and MLCommons calls the results peer-reviewed. On speed, it says the best per-chip result on its visual-language test improved 2.99 times on the round six months earlier, and the best per-chip result on the DeepSeek R1 reasoning test was 5.7 times the best from a year earlier. MLCommons says the new tests reflect the industry's move towards more complex, multi-step and agentic uses of AI.

Lambda's own account shows what an agent test measures. In the open division, where entrants may change the model, Lambda replaced the test's 27-billion-parameter reference model with Kimi K2.6 from Moonshot AI, a model with more than a trillion parameters, and ran it on a data-centre server. Lambda says it completed all 1,007 turns of the recorded agentic-coding task with no failures and scored 86.83 percent on the test's tool-calling accuracy check. Lambda says this was the first model over a trillion parameters in MLPerf and the only agentic submission on data-centre hardware this round.

Nvidia used the round to show a preview of its Vera Rubin NVL72 system. Nvidia says it delivered up to 3.7 times the throughput of its current GB300 NVL72 on the visual-language test and up to 2.5 times on DeepSeek R1. These are Nvidia's own figures for a system MLCommons lists as in preview. Nvidia also cites a 30 times gain on SemiAnalysis AgentX, a separate benchmark outside MLPerf, from preview testing, so it cannot be set against the MLPerf results.

What would change it, and when

MLCommons says its agent test borrows its method from a data-centre version, the MLPerf Agentic benchmark, which it describes as soon to be released. That version would test agents on the kind of hardware cloud providers run. The previous round, v6.0, came six months before this one; if the next round follows the same gap, it would come around March 2027.

What we do not know

These releases do not list each organisation's results, so how the 30 submitters compare with one another is not in the documents here. Nvidia's 3.7 times and 2.5 times figures are its own summary of its submission. The agent test so far measures work on a single edge device, and Lambda's data-centre run was in the open division, so neither is a like-for-like ranking of the services people pay for.

What this changes for you

If you pay for an AI coding tool or any product that works as an agent, there is now a public, reviewed test of how fast such systems run, alongside the companies' own claims. It changes no price or product today. The data-centre version of the agent test is the one to watch, because that is where hosted services like these would be measured.

Sources

Everything above is written from these. Each line says what that document proves.

  1. globenewswire.com: MLCommons' own release, 16 September 2026: a record 30 submitting organisations and their names, the two new tests including Edge Agentic Inference, the 2.99x and 5.7x per-chip gains, and Vera Rubin NVL72 listed as in preview.
  2. lambda.ai: Lambda's own blog, 16 September 2026: its open-division agentic submission with Kimi K2.6 on data-centre hardware, 1,007 of 1,007 turns completed, 86.83 percent tool-calling accuracy, and its claim to the first trillion-parameter model in MLPerf.
  3. blogs.nvidia.com: Nvidia's own blog, 16 September 2026: its preview Vera Rubin NVL72 figures of up to 3.7x and 2.5x over GB300 NVL72, and a separate 30x result on SemiAnalysis AgentX, outside MLPerf.

Keep reading

Strata, the whole AI stack, explained simply. Every number carries a source and a confidence label.

Today · The Stack · Stories · Archive · Glossary · Break the Chain · Focus · About · RSS feed

Privacy · Terms · Corrections · Report an error

Also from Dheeco: Dheeco · AP Fact Check

Published by dheeco.com · © 2026 Dheeco