Vendor Radar Top
Caliper Lab.
Independent AI capability measurement
Platform assessment

Assessment of Due Diligence Platform - Binocs

CitationCommitmentCoverageCharacterConsistencySpecificity
Four AI systems · six dimensions · June 2026
01 · Method

How the assessment works: pull every number, then score it six ways.

Four systems ran the same diligence task and produced six reports each. The pipeline pulled every number from each report and recorded what it measures, its unit, its time period, its source, and whether it is stated firmly or hedged. The same scoring runs across all four systems.

One adjustment matters. Binocs outputs slides, and slides tend to state firm figures by default. So for the fairness of the comparison, one dimension is scored on written text only, not slide tables.

Exhibit 1 · The measurement pipeline
EXTRACT every number metricentityunit periodsourceform TAG six attributes SCORE six dimensions
Six reports per system · pooled per system · satellite-sector task · June 2026
BinocsAgentic platform
Standard interface. Active web sourcing, multi-step reasoning, slide output.
Claude Opus 4.8Frontier LLM · single-shot
One pass, web access enabled, prose output.
GPT 5.5Frontier LLM · single-shot
One pass, web access enabled, prose output.
GPT Deep ResearchAgentic configuration
Multi-step retrieval and reasoning, web enabled, prose output. The closest architectural comparable to Binocs.
02 · Scorecard

Binocs performs strongly on sourcing and breadth, and has potential to improve in-line citation and adding more scenario analysis in diligence.

Six dimensions, four systems. Binocs is clearly ahead on how widely it sources and how many things it measures. On the three dimensions that tell a buyer whether a number can be trusted, citation coverage, self-consistency and forward view, it is matched or beaten by GPT Deep Research. That is an AI configuration built to work the same step-by-step way Binocs does. The two single-shot models sit behind on most of the table.

Exhibit 2 · Six-dimension scorecard
Dimension Binocs Claude Opus 4.8 GPT 5.5 GPT Deep Research
Citation infrastructure
Traceable to live, diverse sources?
STRONG
833 domains · 84% live · 56% cited
MODERATE
96 domains · 70% live · 48% cited
MODERATE
64 domains · 77% live · 52% cited
MOD–STRONG
81 domains · 79% live · 75% cited
Claim commitment
Firm figures or hedged?
STRONG
82% point estimates, prose
LIMITED
34% point · high hedging
MOD–STRONG
59% point estimates, prose
STRONG
72% point estimates, prose
Analytical coverage
Breadth of quantification?
STRONG
1,062 entities · 282 unit types
LIMITED
131 entities · 30 unit types
MODERATE
202 entities · 29% unitless
MODERATE
344 entities · 50 unit types
Analytical character
Historical or forward-looking?
MODERATE
54% dated · 8% forward
MODERATE
56% dated · 14% forward
MODERATE
69% dated · 21% forward
MODERATE
77% dated · 28% forward
Internal consistency
Does it agree with itself?
STRONG
66.7% consistency rate
LIMITED
20.5% consistency rate
MODERATE
34.1% consistency rate
STRONG
66.0% consistency rate
Measurement specificity
Precision of unit vocabulary?
STRONG
282 unit types · 4% unitless
MODERATE
30 unit types · 6% unitless
LIMITED
33 unit types · 29% unitless
MODERATE
50 unit types · 30% unitless
All figures pooled across six reports per system. GPT Deep Research matches Binocs on consistency (66% vs 67%) and leads on citation coverage (75% vs 56%) and forward content (28% vs 8%); Binocs keeps its lead on source breadth and unit vocabulary.
03 · Citation

Binocs sources far more widely, but attaches citations to fewer of its claims.

Binocs cites 7,722 links across 833 websites, about ten times as many sources as any other system here, and 84 percent of them still work. The sources are also well spread, not reliant on a handful of sites.

Wide sourcing is one thing; putting a source on each claim is another. Here Binocs is behind. It attaches a citation to 56 percent of its figures, against 75 percent for GPT Deep Research. In practice, more of the specific numbers a buyer wants to check come without a source attached.

Exhibit 3 · Source breadth
0300600900 833966481 BinocsClaude 4.8GPT 5.5GPT Deep R. UNIQUE CITATION DOMAINS · SIX REPORTS PER SYSTEM
Exhibit 4 · Claim coverage
BinocsClaude 4.8GPT 5.5GPT Deep R. 56%48%52%75% 0%20%40%60%80% SHARE OF EXTRACTED VALUES CARRYING AN INLINE SOURCE
Reachability, source concentration, paywalled share and median source age for all four systems in appendix A3.
04 · Usable figures

Binocs produces the most usable numbers: firm, labelled, and wide-ranging.

A number is useful to a buyer when it is stated firmly, carries a unit, and comes with enough others to build a picture. Binocs is strongest on all three, even after the slide-format adjustment. The two charts below show how firmly each system states its figures, and how much ground it covers.

82%
of Binocs prose figures are firm point estimates, against 34% for Claude
1,062
distinct entities tracked, three times GPT Deep Research, no entity above 6%
4%
of values left unitless, where both GPT systems leave nearly a third
Exhibit 5 · How firmly each system states a figure
BinocsClaude 4.8GPT 5.5GPT Deep R. 82%34% 59%72% 1236 3039 POINT■ navy RANGE■ grey APPROX■ light
Prose values only, controlling for the slide format. n: Binocs 2,515 · Claude 745 · GPT 5.5 955 · GPT Deep R. 1,399. Rounded; rows may not sum to 100. Small segments unlabelled; full table in A4.
Exhibit 6 · Entity breadth and unit vocabulary
1,062131202344 ENTITIES TRACKED 282303350 DISTINCT UNIT TYPES 4%6% 29%30% UNITLESS → BinCla5.5DR BinCla5.5DR
Top-entity concentration: Binocs 6.0% · Claude 16.1% · GPT 5.5 13.7% · GPT Deep R. 12.4%. A figure without a unit cannot be verified or modelled without hunting for context.
05 · Consistency and outlook

Binocs rarely contradicts itself, but it looks backward more than forward.

Consistency means a report does not disagree with its own earlier numbers. A market size in the summary should match the same figure later on. Binocs is strong here at 66.7 percent, level with GPT Deep Research at 66.0, and well ahead of GPT 5.5 at 34.1 and Claude at 20.5.

Where it is weaker is the outlook. Binocs leans on past data, with only 8 percent of its figures looking forward, against 28 percent for GPT Deep Research. That makes it a strong base for the historical picture, and a lighter one for projections.

A pattern runs through the report. The two systems that work step by step, Binocs and GPT Deep Research, behave alike on trust; the two that answer in a single pass fall behind. What matters is less the underlying model than how the system is built.

Exhibit 7 · Agreement rate when a figure recurs
020406080% 66.7%20.5%34.1%66.0% BinocsClaude 4.8GPT 5.5GPT Deep R. SAME METRIC, ENTITY, UNIT AND PERIOD RECURRING IN ONE OUTPUT
Binocs: 977 repeated groups, 652 consistent, 284 contradictions. It repeats far more figures than the LLMs, so the rate is the like-for-like read; full breakdown in A5.
Exhibit 8 · Historical grounding against forward view
BinocsClaude 4.8GPT 5.5GPT Deep R. 32% 8% 12% 14% 23% 21% 22% 28% 0%10%20%30% HISTORICAL SHARE FORWARD-LOOKING SHARE
Share of quantified values by time orientation. Time-stamped overall: Binocs 54% · Claude 56% · GPT 5.5 69% · GPT Deep R. 77%. Binocs is the most backward-anchored system; GPT Deep Research the most forward.
06 · Caliper Lab

The Lab measures what AI products can actually do, and publishes it regardless of who benefits.

Caliper Lab is an independent AI capability measurement institution. Products make capability claims and buyers cannot check them, while generic benchmarks measure academic reasoning and say nothing about real professional work. The Lab builds the evidence base that closes the gap.

Independence is structural, not asserted. The Lab publishes findings regardless of any commercial relationship, grounds every finding in real workflows, and applies the same measurement to every system it tests. This assessment is one entry in that base, and the pipeline behind it keeps running after publication.

For vendors
Third-party evidence that a capability claim is real and measurable. Independent findings carry weight that self-reported benchmarks cannot.
For buyers
A repeatable basis for judging AI tools on real work, with results that compare across tools and tasks rather than vendor decks.
For investors
A way to test whether a capability premium priced into a valuation is real and durable, turning an assertion into evidence.
Appendix

Method, definitions, full tables and limits.

Every figure above resolves to a table here.

A1 · Method in full

For each output the pipeline extracts every numeric value: market sizes, financial figures, growth rates, scores, percentages and ratios. Each value is tagged with the metric it measures, the entity it describes, the unit it uses, the time period it covers, the source citation attached to it if any, and the form it is stated in: a firm point estimate, a range, an approximation or a bound. Metrics are calculated across those records per system, pooled across six reports per system.

The text-tier control. Binocs produces structured PPTX output; Claude, GPT 5.5 and GPT Deep Research produce HTML prose. Tables and charts commit to point estimates by structure, so the claim-commitment dimension uses prose-extracted values only, across all four systems.

A2 · Six-dimension definitions
Citation infrastructure
Are claims traceable to live, diverse sources? Total citations, unique domains, reachability, share of claims with an inline source, concentration and paywalled share.
Claim commitment
Firm or hedged? Prose values only: share stated as point estimates versus ranges, bounds and approximations.
Analytical coverage
How broad is the quantification? Distinct entities, top-entity concentration, topic distribution, unit vocabulary.
Analytical character
Historical or forward-looking? Share of values time-stamped, tied to historical periods, or tied to future periods.
Internal consistency
Does the output agree with itself? Agreement rate where the same metric, entity, unit and period recur in one output.
Measurement specificity
How precise is the unit vocabulary? Distinct unit types, unitless share, numeric precision discipline.
A3 · Citation detail
SystemCitationsDomainsReachableClaims citedSource HHIPaywalledMedian age
Binocs7,72283384%56%0.01211%579 d
Claude Opus 4.81699670%48%0.02125%176 d
GPT 5.51426477%52%0.03717%294 d
GPT Deep Research2008179%75%0.03417%522 d

Source HHI: 0 is perfectly distributed, 1 is a single dominant source. Binocs scores 0.012, low concentration even at 7,722 URLs.

A4 · Commitment, coverage and topic detail

Claim commitment, prose values only.

SystemPointRange / boundsApproximaten (prose)
Binocs82%12%6%2,515
Claude Opus 4.834%36%30%745
GPT 5.559%39%2%955
GPT Deep Research72%21%7%1,399

Percentages rounded; rows may not sum to exactly 100.

Topic distribution. Share of quantified values per topic.

TopicBinocsClaudeGPT 5.5GPT DR
Financials27%30%27%26%
Scoring / rating23%3%1%10%
Other22%17%31%25%
Product / tech12%6%8%8%
Market sizing9%33%23%19%
Competition2%2%5%6%
Other categories5%8%6%7%

Top-entity concentration: Binocs 6.0%; Claude 16.1%, all on one company (ICEYE); GPT 5.5 13.7%; GPT Deep Research 12.4%. Unitless: Binocs 4%, Claude 6%, GPT 5.5 29%, GPT Deep Research 30%.

A5 · Consistency, character and precision detail

Internal consistency. Groups where the same metric, entity, unit and period recur in one output.

SystemRepeated groupsConsistentContradictionsRate
Binocs97765228466.7%
Claude Opus 4.873155820.5%
GPT 5.582285434.1%
GPT Deep Research100663466.0%

Binocs generates 10 to 17 times more quantified values than the LLMs, producing proportionally more repeated groups; PPTX repeats figures by design across slides. GPT Deep Research matches Binocs on the rate; its lower contradiction count reflects far fewer repeated groups. Claude's low rate may partly reflect heavier use of ranges and approximations.

Analytical character. Share of quantified values by time orientation.

SystemTime-stampedHistoricalForward-looking
Binocs54%32%8%
Claude Opus 4.856%12%14%
GPT 5.569%23%21%
GPT Deep Research77%22%28%

Numeric precision. All four show sensible rounding discipline and low false-precision risk.

SystemMean decimal placesMax3+ dp share
Binocs1.2960.8%
Claude Opus 4.81.0620.0%
GPT 5.51.0840.3%
GPT Deep Research1.0430.9%

Binocs' occasional high decimal count likely reflects derived calculations or precise extraction artefacts; at 0.8% of values it stays low.

A6 · Qualifications and limiting conditions

Scope. Findings are based on automated analysis of AI-generated outputs during the evaluation period. Results reflect the specific systems, configurations, prompts and datasets used; different configurations may produce materially different results, and findings should not be generalised beyond the sector, task types and system versions evaluated.

Source data. Data has been obtained from sources believed reliable. Caliper Lab makes no representation as to the accuracy or completeness of third-party data and has accepted it without independent verification of the underlying source claims.

Not investment advice. This assessment is not investment advice, a recommendation to buy or sell any product or service, or an opinion on the fairness of any transaction. It informs procurement and product evaluation decisions and should be read in that context.

Temporal validity. Findings are valid as of June 2026. AI capabilities and platform features change frequently, and Caliper Lab assumes no obligation to update this assessment for later changes.

Independence. Caliper Lab conducted this evaluation independently and without commercial bias. No fee arrangement with any evaluated vendor has influenced the findings, methodology, scoring or conclusions, and standards are applied consistently across vendors regardless of commercial relationship.

Caliper Lab · Independent AI capability measurement Assessment of Due Diligence Platform - Binocs · June 2026 · thecaliperlab.com