Vendor Radar Top
Caliper Lab.
Independent AI capability measurement
Platform assessment

Assessment of AI Financial Modelling Platform - Tracelight

Model constructionHistory & traceabilityForecast & self-checkError detection
Three systems · four measurement blocks · July 2026
01 · Method

Three systems, the same job, judged against the real filings.

Each system built a financial model for Apple and Nike across the 2023 and 2024 filing years. That produced twelve models in total, all created from the same untouched raw filings.

The scoring asked whether the model rebuilt history, traced figures to source, used live forecast formulas, stayed balanced, and responded correctly when assumptions or model inputs changed. Scores run from 0 to 100, and the four blocks remain separate rather than being rolled into one winner.

Exhibit 1 · The evaluation pipeline
COMMON INPUTS Apple + NikeFY2023 / FY2024 12 Excel models3 systems × 4 cases Answer-key scoringexact filing cells retained STRESS TESTS Assumption-change testraise one growth input and follow the modelformula survival · propagation · closure Planted-error testreplace forecast cash with a wrong hardcodeindependent detection · model checks · location Four blocks remain separate · no single combined winner
Two companies · two filing years · three systems · one repeatable scoring harness
TracelightPurpose-built financial modelling platform
The subject system: filing-based model generation with source traceability and model checks.
ClaudeGeneral-purpose AI model
A capable comparison system completing the same modelling task from the same filing workbooks.
GPTGeneral-purpose AI model
A second comparison system, scored on the same answer keys and stress tests.
02 · Scorecard

Tracelight's strength is traceability. It matches the other systems on building the model, and is a step behind on checking its own work.

Tracelight behaves like a purpose-built model checker. It is strongest at tying numbers back to source and at catching a planted error. Claude is slightly ahead on next-year prediction. Both Claude and GPT score higher than Tracelight on self-audit agreement, although GPT trails on forecast accuracy and change propagation.

Exhibit 2 · Four-block scorecard
BlockTracelightClaudeGPT
Building the model
Complete, connected and formula-driven?
MOD–STRONG
connections 90% · formulas 95%
MOD–STRONG
connections 90% · formulas 95%
MODERATE
connections 82% · formulas 97%
History and traceability
History right and every figure sourced?
STRONG
coverage 84% · location 75%
MOD–STRONG
coverage 91% · location 73%
MODERATE
coverage 95% · location 39%
Forecast and self-check
Forecast holds and the model checks itself?
MOD–STRONG
forecast 65% · change pass 75%
STRONG
forecast 68% · change pass 100%
LIMITED
forecast 47% · change pass 25%
Error detection
Does it catch a planted model error?
STRONG
fault pass 100%
STRONG
fault pass 100%
MODERATE
fault pass 67%
A note on the labels. The labels are a plain summary of the numbers in the sections that follow. They are not a grade, and they are not added up. Each block is read on its own terms because each measures a different thing.
03 · Building the model

All three produce the required model structure, with differences in completeness and connectivity.

The first block asks whether each system produced a proper model to start from: the right sheets, no links to files the reader would not have, connected statements, and forecasts built on formulas rather than typed-in numbers.

100%
of all models were self-contained, with no outside file links
89.6%
Tracelight connectivity, almost level with Claude at 89.5%
96.9%
GPT formula-driven forecast rate, the highest of the three
Exhibit 3 · Construction metrics
TracelightClaudeGPTCompletion of the workbookTracelight65.3Claude65.3GPT53.7Self-contained, no outside linksTracelight100.0Claude100.0GPT100.0Statements wired togetherTracelight89.6Claude89.5GPT82.1Forecasts driven by formulasTracelight94.9Claude94.6GPT96.9
Score out of 100 · higher is better

The three systems land within a few points of each other. Tracelight has every required connection in place. GPT is slightly ahead on formula quality. Every system keeps the model self-contained.

What this means: building a sound model is the basic bar for this category. All three clear it, so this is not where the systems separate.

04 · History and traceability

Tracelight links its numbers back to the filing better than the other two.

The history was rebuilt at about the same level by all three systems, around 85 out of 100. The difference is in traceability: whether declared source cells exist, point to the right place, and cover the expected historical lines.

100%
of the filing cells cited by Tracelight actually exist
75.0%
source-location accuracy, narrowly ahead of Claude and far ahead of GPT
84.1%
lineage coverage, the one traceability measure where Tracelight trails

Every source Tracelight cites is real, and it points to the correct cell more often than the other two. The one place it trails is coverage: it sources 84 of every 100 lines, compared with 91 for Claude and 95 for GPT.

What this means: what Tracelight cites is real and correctly located. The improvement opportunity is to attach that traceability to more lines.

Exhibit 4 · Traceability metrics
TracelightClaudeGPTEvery source it cites is realTracelight100.0Claude85.8GPT95.5It points to the right cellTracelight75.0Claude73.1GPT39.2Every line has a sourceTracelight84.1Claude91.4GPT94.9
Source-cell existence · source-location accuracy · lineage coverage
05 · Forecast response

All three move revenue in the right direction; downstream propagation and closure separate them.

The first forecast check compares the next-year forecast with the actual result that was held back during model creation. The second changes a growth assumption and asks whether the forecast moves correctly without breaking formulas or creating new Excel errors.

Exhibit 5 · Forecast and change response
TracelightClaudeGPTPredicting the next yearTracelight65.4Claude67.5GPT47.4Responding when an assumption changesTracelight98.4Claude100.0GPT80.9
Next-year prediction uses the held-back FY2024 actual result

On predicting the next year, the systems are close, with GPT further back. On responding to a change, Tracelight is near the top. When the growth assumption is raised, revenue moves the right way, no new errors appear, and the formulas hold.

Takeaway: the forecast responds well. The next section covers the stricter checks, where the other two are ahead.

06 · Self-check and robustness

Claude leads the strict change test; GPT leads only on baseline self-audit agreement.

The tougher checks ask whether the accounts still balance after a change, whether the whole controlled change passes cleanly on every model, and whether the model's own audit agrees with an independent recalculation.

87.5%
Tracelight closure retention after the assumption change
75.0%
full knob-test pass rate: three of four Tracelight models
59.1%
Tracelight self-audit agreement, the clearest measured gap

Tracelight keeps the accounts balanced better than GPT and just behind Claude. On the whole change passing cleanly, it clears three of four models. The clearest gap is the self-check: the model's own audit is right 59 times out of 100, the lowest of the three.

What this means: Tracelight is very good at catching errors from outside. Its own internal check is the part with the most room to improve.

Exhibit 6 · Strict change and self-audit checks
TracelightClaudeGPTAccounts still balance after a changeTracelight87.5Claude100.0GPT68.8The whole change passes cleanlyTracelight75.0Claude100.0GPT25.0The model self-audit is correctTracelight59.1Claude65.3GPT79.4
Closure retention · full controlled-change pass · baseline self-audit agreement
07 · Error detection

Tracelight catches a planted error every time.

A correct forecast cash formula was replaced with a wrong fixed number. The test asks whether independent checks and the workbook's own checks notice the error, and whether they point to the right location without disturbing the rest of the model.

100%
Tracelight overall fault-test pass rate
100%
full fault detection across every eligible Tracelight model
66.7%
GPT fault-test pass rate, based on three eligible models
Exhibit 7 · Planted-error results
MeasureTracelightClaudeGPT
Fault test pass10010066.7
Full fault detection10010066.7
Independent cash-mismatch detection100100100.0
Balance-sheet check triggered10010033.3
Cash check triggered10010066.7
Error localization100100100.0
GPT could run this test on three of its four models because one already had a broken cash figure before the planted error was added.

Across every check, Tracelight caught the planted error on every model. Claude did the same. On catching an error placed inside a finished model, Tracelight is at the top.

08 · Caliper Lab

The Lab measures what AI products can actually do, and publishes it regardless of who benefits.

Caliper Lab is an independent AI capability measurement institution. Products make capability claims and buyers cannot check them, while generic benchmarks measure academic reasoning and say little about real professional work. The Lab builds the evidence base that closes the gap.

Independence is structural, not asserted. The Lab publishes findings regardless of any commercial relationship, grounds every finding in real workflows, and applies the same measurement to every system it tests.

For vendors
Independent evidence that a capability claim is real and measurable.
For buyers
A repeatable basis for comparing AI tools on real professional work.
For investors
A rigorous way to test whether a capability premium is real and durable.
Appendix

Method, definitions, full tables and limits.

Every figure above resolves to a definition or table here.

A1 · Method in full

Each system built a financial model for Apple and Nike from the untouched FY2023 and FY2024 filing workbooks. Every answer-key value retained its source worksheet and source cell. Deterministic scoring compared the generated models with those answer keys, while controlled perturbation tests changed one forecast assumption and planted one cash error. The assessment groups the evidence into four blocks and does not combine them into a single overall winner.

Eligibility. Most metrics use four models per system. Next-year prediction uses the FY2023-origin models for which FY2024 actual results could be revealed at scoring. GPT's planted-error result uses three eligible models because one GPT base model already had a nonnumeric forecast cash value.

A2 · Four-block definitions
Building the model
Whether the workbook is sufficiently complete, contains the required sheets, is self-contained, preserves the source filing, connects the statements, and drives forecast cells through formulas.
History and traceability
Whether historical values match the filing with the correct signs and units, expected lines are present, and every declared source exists and points to the right filing cell.
Forecast and self-check
Whether forecasts land close to the held-back actual result, respond correctly to a controlled assumption change, preserve formulas and closure, and agree with an independent model audit.
Error detection
Whether a planted cash hardcode is detected independently and by the model, triggers the relevant checks, is correctly located, and leaves the rest of the workbook intact.
A3 · Building the model detail
MeasureTracelightClaudeGPT
Workbook completion65.365.353.7
Self-contained, no outside links100.0100.0100.0
Statements wired together89.689.582.1
Forecasts driven by formulas94.994.696.9
A4 · History and traceability detail
MeasureTracelightClaudeGPT
History rebuilt correctly≈85≈85≈85
Every cited source cell exists100.085.895.5
Source points to the right cell75.073.139.2
Every historical line has a source84.191.494.9

The historical reconstruction result is described in the report as approximately 85 out of 100 for all three systems.

A5 · Forecast and self-check detail
MeasureTracelightClaudeGPT
Predicting the next year65.467.547.4
Responding when an assumption changes98.4100.080.9
Accounts still balance after a change87.5100.068.8
The whole change passes cleanly75.0100.025.0
The model self-audit is correct59.165.379.4
A6 · Error detection detail
MeasureTracelightClaudeGPT
Fault test pass10010066.7
Full fault detection10010066.7
Independent cash-mismatch detection100100100.0
Balance-sheet check triggered10010033.3
Cash check triggered10010066.7
Error localization100100100.0

GPT's error-detection figures are based on three eligible models rather than four.

A7 · Measure definitions
Workbook completion
Whether the model was produced complete rather than partial.
Required sheets
Whether every sheet the model should have is present.
Self-contained
Whether the model avoids linking to files the reader would not have.
Filing preserved
Whether the source filing is kept intact.
Connections in place
Whether the statements are wired together as they should be.
Formula-driven forecasts
Whether forecast cells use formulas rather than typed-in numbers.
History rebuilt correctly
Whether the reported historical numbers match the filing.
Signs and units correct
Whether positives, negatives, and reporting units match the filing.
History complete
Whether all expected historical lines are present.
Source is real
Whether a cited source cell actually exists in the filing.
Source is correct
Whether a cited source points to the right cell.
Every line sourced
Whether every historical line carries a declared source.
Carried across cleanly
Whether values carry from the source into the model unchanged.
Next-year accuracy
How close the held-back forecast lands to the real result.
Lines within 10%
The share of forecast lines within ten percent of the real result.
Change applied correctly
Whether the assumption change was applied at the right size.
Forecast responds
Whether the forecast moves, and in the right direction, after a change.
Change flows through
The share of forecast lines that update after a change.
Formulas survive
Whether forecast formulas hold after a change.
Accounts still balance
Whether the balance-sheet and cash checks close independently.
Change passes cleanly
The overall pass for the controlled assumption change.
Assumptions documented
Whether assumptions are clearly labelled and editable.
Self-audit is correct
Whether the model self-audit agrees with an independent recalculation.
Fault size correct
Whether the planted cash error was the intended size.
Hidden hardcode found
Whether a formula replaced by a fixed value was detected.
Independent check catches it
Whether the independent check caught the mismatch.
Model catches it
Whether the model self-audit caught the planted error.
Cash and balance checks fire
Whether the built-in checks triggered.
Error located
Whether the error was pinned to the right place.
Rest of model intact
Whether formulas away from the error were left untouched.
Full detection
Whether both the independent check and the model caught it.
Test passes overall
The overall pass for the planted-error test.
A8 · Qualifications and limiting conditions

Purpose and intended audience. This assessment has been prepared solely for the party to whom it is addressed. Reproduction, quotation or distribution without Caliper Lab's prior written consent is prohibited.

Scope of findings. Results reflect the specific systems, configurations, prompts and datasets used during the evaluation period. Different configurations may produce materially different outcomes, and findings should not be generalised beyond the task types and versions evaluated.

Source data. Information was obtained from sources believed to be reliable. Caliper Lab makes no representation as to the accuracy or completeness of third-party source claims beyond the filing-based checks conducted in this evaluation.

Not investment advice. This assessment does not constitute investment advice, a recommendation to buy or sell any product or service, or an opinion on the fairness of any transaction.

Temporal validity. AI capabilities change frequently. Findings are valid as of the evaluation date, and Caliper Lab assumes no obligation to update them.

Independence. No fee arrangement with any evaluated vendor influenced the findings, methodology, scoring or conclusions. The same standards were applied to all systems.

No third-party liability. Caliper Lab accepts no liability to third parties for decisions made on the basis of this assessment.