Each system built a financial model for Apple and Nike across the 2023 and 2024 filing years. That produced twelve models in total, all created from the same untouched raw filings.
The scoring asked whether the model rebuilt history, traced figures to source, used live forecast formulas, stayed balanced, and responded correctly when assumptions or model inputs changed. Scores run from 0 to 100, and the four blocks remain separate rather than being rolled into one winner.
Tracelight behaves like a purpose-built model checker. It is strongest at tying numbers back to source and at catching a planted error. Claude is slightly ahead on next-year prediction. Both Claude and GPT score higher than Tracelight on self-audit agreement, although GPT trails on forecast accuracy and change propagation.
| Block | Tracelight | Claude | GPT |
|---|---|---|---|
Building the model Complete, connected and formula-driven? | MOD–STRONG connections 90% · formulas 95% | MOD–STRONG connections 90% · formulas 95% | MODERATE connections 82% · formulas 97% |
History and traceability History right and every figure sourced? | STRONG coverage 84% · location 75% | MOD–STRONG coverage 91% · location 73% | MODERATE coverage 95% · location 39% |
Forecast and self-check Forecast holds and the model checks itself? | MOD–STRONG forecast 65% · change pass 75% | STRONG forecast 68% · change pass 100% | LIMITED forecast 47% · change pass 25% |
Error detection Does it catch a planted model error? | STRONG fault pass 100% | STRONG fault pass 100% | MODERATE fault pass 67% |
The first block asks whether each system produced a proper model to start from: the right sheets, no links to files the reader would not have, connected statements, and forecasts built on formulas rather than typed-in numbers.
The three systems land within a few points of each other. Tracelight has every required connection in place. GPT is slightly ahead on formula quality. Every system keeps the model self-contained.
What this means: building a sound model is the basic bar for this category. All three clear it, so this is not where the systems separate.
The history was rebuilt at about the same level by all three systems, around 85 out of 100. The difference is in traceability: whether declared source cells exist, point to the right place, and cover the expected historical lines.
Every source Tracelight cites is real, and it points to the correct cell more often than the other two. The one place it trails is coverage: it sources 84 of every 100 lines, compared with 91 for Claude and 95 for GPT.
What this means: what Tracelight cites is real and correctly located. The improvement opportunity is to attach that traceability to more lines.
The first forecast check compares the next-year forecast with the actual result that was held back during model creation. The second changes a growth assumption and asks whether the forecast moves correctly without breaking formulas or creating new Excel errors.
On predicting the next year, the systems are close, with GPT further back. On responding to a change, Tracelight is near the top. When the growth assumption is raised, revenue moves the right way, no new errors appear, and the formulas hold.
Takeaway: the forecast responds well. The next section covers the stricter checks, where the other two are ahead.
The tougher checks ask whether the accounts still balance after a change, whether the whole controlled change passes cleanly on every model, and whether the model's own audit agrees with an independent recalculation.
Tracelight keeps the accounts balanced better than GPT and just behind Claude. On the whole change passing cleanly, it clears three of four models. The clearest gap is the self-check: the model's own audit is right 59 times out of 100, the lowest of the three.
What this means: Tracelight is very good at catching errors from outside. Its own internal check is the part with the most room to improve.
A correct forecast cash formula was replaced with a wrong fixed number. The test asks whether independent checks and the workbook's own checks notice the error, and whether they point to the right location without disturbing the rest of the model.
| Measure | Tracelight | Claude | GPT |
|---|---|---|---|
| Fault test pass | 100 | 100 | 66.7 |
| Full fault detection | 100 | 100 | 66.7 |
| Independent cash-mismatch detection | 100 | 100 | 100.0 |
| Balance-sheet check triggered | 100 | 100 | 33.3 |
| Cash check triggered | 100 | 100 | 66.7 |
| Error localization | 100 | 100 | 100.0 |
Across every check, Tracelight caught the planted error on every model. Claude did the same. On catching an error placed inside a finished model, Tracelight is at the top.
Caliper Lab is an independent AI capability measurement institution. Products make capability claims and buyers cannot check them, while generic benchmarks measure academic reasoning and say little about real professional work. The Lab builds the evidence base that closes the gap.
Independence is structural, not asserted. The Lab publishes findings regardless of any commercial relationship, grounds every finding in real workflows, and applies the same measurement to every system it tests.
Every figure above resolves to a definition or table here.
Each system built a financial model for Apple and Nike from the untouched FY2023 and FY2024 filing workbooks. Every answer-key value retained its source worksheet and source cell. Deterministic scoring compared the generated models with those answer keys, while controlled perturbation tests changed one forecast assumption and planted one cash error. The assessment groups the evidence into four blocks and does not combine them into a single overall winner.
Eligibility. Most metrics use four models per system. Next-year prediction uses the FY2023-origin models for which FY2024 actual results could be revealed at scoring. GPT's planted-error result uses three eligible models because one GPT base model already had a nonnumeric forecast cash value.
| Measure | Tracelight | Claude | GPT |
|---|---|---|---|
| Workbook completion | 65.3 | 65.3 | 53.7 |
| Self-contained, no outside links | 100.0 | 100.0 | 100.0 |
| Statements wired together | 89.6 | 89.5 | 82.1 |
| Forecasts driven by formulas | 94.9 | 94.6 | 96.9 |
| Measure | Tracelight | Claude | GPT |
|---|---|---|---|
| History rebuilt correctly | ≈85 | ≈85 | ≈85 |
| Every cited source cell exists | 100.0 | 85.8 | 95.5 |
| Source points to the right cell | 75.0 | 73.1 | 39.2 |
| Every historical line has a source | 84.1 | 91.4 | 94.9 |
The historical reconstruction result is described in the report as approximately 85 out of 100 for all three systems.
| Measure | Tracelight | Claude | GPT |
|---|---|---|---|
| Predicting the next year | 65.4 | 67.5 | 47.4 |
| Responding when an assumption changes | 98.4 | 100.0 | 80.9 |
| Accounts still balance after a change | 87.5 | 100.0 | 68.8 |
| The whole change passes cleanly | 75.0 | 100.0 | 25.0 |
| The model self-audit is correct | 59.1 | 65.3 | 79.4 |
| Measure | Tracelight | Claude | GPT |
|---|---|---|---|
| Fault test pass | 100 | 100 | 66.7 |
| Full fault detection | 100 | 100 | 66.7 |
| Independent cash-mismatch detection | 100 | 100 | 100.0 |
| Balance-sheet check triggered | 100 | 100 | 33.3 |
| Cash check triggered | 100 | 100 | 66.7 |
| Error localization | 100 | 100 | 100.0 |
GPT's error-detection figures are based on three eligible models rather than four.
Purpose and intended audience. This assessment has been prepared solely for the party to whom it is addressed. Reproduction, quotation or distribution without Caliper Lab's prior written consent is prohibited.
Scope of findings. Results reflect the specific systems, configurations, prompts and datasets used during the evaluation period. Different configurations may produce materially different outcomes, and findings should not be generalised beyond the task types and versions evaluated.
Source data. Information was obtained from sources believed to be reliable. Caliper Lab makes no representation as to the accuracy or completeness of third-party source claims beyond the filing-based checks conducted in this evaluation.
Not investment advice. This assessment does not constitute investment advice, a recommendation to buy or sell any product or service, or an opinion on the fairness of any transaction.
Temporal validity. AI capabilities change frequently. Findings are valid as of the evaluation date, and Caliper Lab assumes no obligation to update them.
Independence. No fee arrangement with any evaluated vendor influenced the findings, methodology, scoring or conclusions. The same standards were applied to all systems.
No third-party liability. Caliper Lab accepts no liability to third parties for decisions made on the basis of this assessment.