Four systems ran the same diligence task and produced six reports each. The pipeline pulled every number from each report and recorded what it measures, its unit, its time period, its source, and whether it is stated firmly or hedged. The same scoring runs across all four systems.
One adjustment matters. Binocs outputs slides, and slides tend to state firm figures by default. So for the fairness of the comparison, one dimension is scored on written text only, not slide tables.
Six dimensions, four systems. Binocs is clearly ahead on how widely it sources and how many things it measures. On the three dimensions that tell a buyer whether a number can be trusted, citation coverage, self-consistency and forward view, it is matched or beaten by GPT Deep Research. That is an AI configuration built to work the same step-by-step way Binocs does. The two single-shot models sit behind on most of the table.
| Dimension | Binocs | Claude Opus 4.8 | GPT 5.5 | GPT Deep Research |
|---|---|---|---|---|
Citation infrastructure Traceable to live, diverse sources? |
STRONG 833 domains · 84% live · 56% cited |
MODERATE 96 domains · 70% live · 48% cited |
MODERATE 64 domains · 77% live · 52% cited |
MOD–STRONG 81 domains · 79% live · 75% cited |
Claim commitment Firm figures or hedged? |
STRONG 82% point estimates, prose |
LIMITED 34% point · high hedging |
MOD–STRONG 59% point estimates, prose |
STRONG 72% point estimates, prose |
Analytical coverage Breadth of quantification? |
STRONG 1,062 entities · 282 unit types |
LIMITED 131 entities · 30 unit types |
MODERATE 202 entities · 29% unitless |
MODERATE 344 entities · 50 unit types |
Analytical character Historical or forward-looking? |
MODERATE 54% dated · 8% forward |
MODERATE 56% dated · 14% forward |
MODERATE 69% dated · 21% forward |
MODERATE 77% dated · 28% forward |
Internal consistency Does it agree with itself? |
STRONG 66.7% consistency rate |
LIMITED 20.5% consistency rate |
MODERATE 34.1% consistency rate |
STRONG 66.0% consistency rate |
Measurement specificity Precision of unit vocabulary? |
STRONG 282 unit types · 4% unitless |
MODERATE 30 unit types · 6% unitless |
LIMITED 33 unit types · 29% unitless |
MODERATE 50 unit types · 30% unitless |
Binocs cites 7,722 links across 833 websites, about ten times as many sources as any other system here, and 84 percent of them still work. The sources are also well spread, not reliant on a handful of sites.
Wide sourcing is one thing; putting a source on each claim is another. Here Binocs is behind. It attaches a citation to 56 percent of its figures, against 75 percent for GPT Deep Research. In practice, more of the specific numbers a buyer wants to check come without a source attached.
A number is useful to a buyer when it is stated firmly, carries a unit, and comes with enough others to build a picture. Binocs is strongest on all three, even after the slide-format adjustment. The two charts below show how firmly each system states its figures, and how much ground it covers.
Consistency means a report does not disagree with its own earlier numbers. A market size in the summary should match the same figure later on. Binocs is strong here at 66.7 percent, level with GPT Deep Research at 66.0, and well ahead of GPT 5.5 at 34.1 and Claude at 20.5.
Where it is weaker is the outlook. Binocs leans on past data, with only 8 percent of its figures looking forward, against 28 percent for GPT Deep Research. That makes it a strong base for the historical picture, and a lighter one for projections.
A pattern runs through the report. The two systems that work step by step, Binocs and GPT Deep Research, behave alike on trust; the two that answer in a single pass fall behind. What matters is less the underlying model than how the system is built.
Caliper Lab is an independent AI capability measurement institution. Products make capability claims and buyers cannot check them, while generic benchmarks measure academic reasoning and say nothing about real professional work. The Lab builds the evidence base that closes the gap.
Independence is structural, not asserted. The Lab publishes findings regardless of any commercial relationship, grounds every finding in real workflows, and applies the same measurement to every system it tests. This assessment is one entry in that base, and the pipeline behind it keeps running after publication.
Every figure above resolves to a table here.
For each output the pipeline extracts every numeric value: market sizes, financial figures, growth rates, scores, percentages and ratios. Each value is tagged with the metric it measures, the entity it describes, the unit it uses, the time period it covers, the source citation attached to it if any, and the form it is stated in: a firm point estimate, a range, an approximation or a bound. Metrics are calculated across those records per system, pooled across six reports per system.
The text-tier control. Binocs produces structured PPTX output; Claude, GPT 5.5 and GPT Deep Research produce HTML prose. Tables and charts commit to point estimates by structure, so the claim-commitment dimension uses prose-extracted values only, across all four systems.
| System | Citations | Domains | Reachable | Claims cited | Source HHI | Paywalled | Median age |
|---|---|---|---|---|---|---|---|
| Binocs | 7,722 | 833 | 84% | 56% | 0.012 | 11% | 579 d |
| Claude Opus 4.8 | 169 | 96 | 70% | 48% | 0.021 | 25% | 176 d |
| GPT 5.5 | 142 | 64 | 77% | 52% | 0.037 | 17% | 294 d |
| GPT Deep Research | 200 | 81 | 79% | 75% | 0.034 | 17% | 522 d |
Source HHI: 0 is perfectly distributed, 1 is a single dominant source. Binocs scores 0.012, low concentration even at 7,722 URLs.
Claim commitment, prose values only.
| System | Point | Range / bounds | Approximate | n (prose) |
|---|---|---|---|---|
| Binocs | 82% | 12% | 6% | 2,515 |
| Claude Opus 4.8 | 34% | 36% | 30% | 745 |
| GPT 5.5 | 59% | 39% | 2% | 955 |
| GPT Deep Research | 72% | 21% | 7% | 1,399 |
Percentages rounded; rows may not sum to exactly 100.
Topic distribution. Share of quantified values per topic.
| Topic | Binocs | Claude | GPT 5.5 | GPT DR |
|---|---|---|---|---|
| Financials | 27% | 30% | 27% | 26% |
| Scoring / rating | 23% | 3% | 1% | 10% |
| Other | 22% | 17% | 31% | 25% |
| Product / tech | 12% | 6% | 8% | 8% |
| Market sizing | 9% | 33% | 23% | 19% |
| Competition | 2% | 2% | 5% | 6% |
| Other categories | 5% | 8% | 6% | 7% |
Top-entity concentration: Binocs 6.0%; Claude 16.1%, all on one company (ICEYE); GPT 5.5 13.7%; GPT Deep Research 12.4%. Unitless: Binocs 4%, Claude 6%, GPT 5.5 29%, GPT Deep Research 30%.
Internal consistency. Groups where the same metric, entity, unit and period recur in one output.
| System | Repeated groups | Consistent | Contradictions | Rate |
|---|---|---|---|---|
| Binocs | 977 | 652 | 284 | 66.7% |
| Claude Opus 4.8 | 73 | 15 | 58 | 20.5% |
| GPT 5.5 | 82 | 28 | 54 | 34.1% |
| GPT Deep Research | 100 | 66 | 34 | 66.0% |
Binocs generates 10 to 17 times more quantified values than the LLMs, producing proportionally more repeated groups; PPTX repeats figures by design across slides. GPT Deep Research matches Binocs on the rate; its lower contradiction count reflects far fewer repeated groups. Claude's low rate may partly reflect heavier use of ranges and approximations.
Analytical character. Share of quantified values by time orientation.
| System | Time-stamped | Historical | Forward-looking |
|---|---|---|---|
| Binocs | 54% | 32% | 8% |
| Claude Opus 4.8 | 56% | 12% | 14% |
| GPT 5.5 | 69% | 23% | 21% |
| GPT Deep Research | 77% | 22% | 28% |
Numeric precision. All four show sensible rounding discipline and low false-precision risk.
| System | Mean decimal places | Max | 3+ dp share |
|---|---|---|---|
| Binocs | 1.29 | 6 | 0.8% |
| Claude Opus 4.8 | 1.06 | 2 | 0.0% |
| GPT 5.5 | 1.08 | 4 | 0.3% |
| GPT Deep Research | 1.04 | 3 | 0.9% |
Binocs' occasional high decimal count likely reflects derived calculations or precise extraction artefacts; at 0.8% of values it stays low.
Scope. Findings are based on automated analysis of AI-generated outputs during the evaluation period. Results reflect the specific systems, configurations, prompts and datasets used; different configurations may produce materially different results, and findings should not be generalised beyond the sector, task types and system versions evaluated.
Source data. Data has been obtained from sources believed reliable. Caliper Lab makes no representation as to the accuracy or completeness of third-party data and has accepted it without independent verification of the underlying source claims.
Not investment advice. This assessment is not investment advice, a recommendation to buy or sell any product or service, or an opinion on the fairness of any transaction. It informs procurement and product evaluation decisions and should be read in that context.
Temporal validity. Findings are valid as of June 2026. AI capabilities and platform features change frequently, and Caliper Lab assumes no obligation to update this assessment for later changes.
Independence. Caliper Lab conducted this evaluation independently and without commercial bias. No fee arrangement with any evaluated vendor has influenced the findings, methodology, scoring or conclusions, and standards are applied consistently across vendors regardless of commercial relationship.