Independent benchmarks
Structured evaluations of AI products and models on real professional workflows. Published independently.
Independent AI Evaluation
We benchmark AI products and frontier models on real-world professional tasks, and publish what we find. Independent. No vendor involvement in results.
Structured findings, and the common standard the market currently lacks.
View our latest model benchmarksThe problem
Enterprise AI products make capability claims. Accuracy. Reliability. Performance that exceeds a frontier model on domain-specific tasks.
Buyers cannot verify them. Generic benchmarks measure academic reasoning. They say nothing about whether an AI product performs on real professional workflows.
The gap between what exists and what the market needs is the problem we exist to solve.
Hands-on assessments of AI products, run against real professional tasks. Three published so far.
Agent-based commercial diligence run against a real diligence brief, with every sourced claim checked back to its document.
Deck generation from a brief and a source pack, assessed on structure, chart fidelity and how much editing the output needed.
Natural-language model building inside Excel, tested on construction and on whether it catches formula errors it did not introduce.
What we do
Structured evaluations of AI products and models on real professional workflows. Published independently.
Every evaluation produces a capability finding. What the product does well, where it fails, and how the vendor's claims hold up.
The evaluation framework is public. Methodology, scoring criteria, and task design are documented openly. Any party can verify our results with access to the same inputs.