The evaluation opened with a wider field. Every system was given the same task: take a detailed brief and a pack of source reports, and build a finished deck end to end. Microsoft Copilot and a GPT-based PowerPoint plug-in could not do this reliably, and were set aside at screening.
The two systems that cleared that bar, Perceptis and Claude for PowerPoint, then ran the same three briefs, six decks each, scored across thirteen measurement families. The metrics are drawn from a reference library of professional decks from leading strategy firms, not from opinion.
Four blocks, two systems. Read across them, the two systems have distinct profiles. Perceptis behaves like a deck production system, building visual structure by default. Claude for PowerPoint behaves like a strong writer working inside PowerPoint. Perceptis is clearly ahead on how the deck is built, level on following the brief, and behind on message economy and how firmly each claim is tied to its source.
| Block | Perceptis | Claude for PowerPoint |
|---|---|---|
Visual construction Built to a professional visual standard? |
STRONG 85% of charts titled · 8 visuals per deck |
MODERATE 50% of charts titled · 1.5 visuals per deck |
Structure & instructions Did it build what the brief asked for? |
STRONG 94.9% requirement coverage |
STRONG 94.8% requirement coverage |
Message & narrative Sharp slides, deck that holds together? |
MOD–STRONG specific findings, but runs long |
STRONG specific and concise |
Evidence & rigor Claims anchored, argument sound? |
MODERATE 90% of slides sourced · claim support varies |
MOD–STRONG 65% of slides sourced · 84% argument soundness |
The sections that follow take each block in turn, from the surface of the deck inward: first how it is built, then whether it did what was asked, then how it reads, and finally how well it holds up to scrutiny.
A chart without a title or units forces the reader to reconstruct what they are looking at. Finished visuals separate a deck that can be presented from one that needs a cleanup pass. This is the widest gap in the evaluation, and it favours Perceptis.
Before a deck can be judged on how it looks or reads, it has to deliver what was requested: the right sections, the right content, the right number of slides. This is the competence floor of the category. Both systems clear it, and clear it together.
There is no meaningful separation here. Both systems deliver the requested content at a high rate and match the requested structure.
This is also the bar that decided the field. The two systems set aside at screening could not carry a detailed brief and a source pack through to a finished deck. Clearing it reliably is what put Perceptis and Claude for PowerPoint into the deep evaluation.
The only visible movement is in conditional instructions, those that apply only when a stated condition holds, where Perceptis is slightly behind. The difference rests on a handful of cases and is not a stable gap.
Perceptis states real findings in its slide titles, at parity with the comparison system, and the slide beneath supports them. Message economy is the gap versus the frontier. Senior readers move through a deck in seconds per slide, and messages that run long tax exactly that audience.
There is a difference between a slide that carries a source note and a claim that survives being checked against that source. The first is a formatting discipline, and Perceptis leads on it. The second is an evidential discipline, and it is Perceptis's main gap versus the frontier.
On the harder check, whether each individual claim is fully supported by the source it cites, Perceptis scored lower, and the result varied widely from one brief to the next.
The Lab reports this as a direction rather than a single figure, because the spread across briefs is too wide to fix a defensible number. It is consistent with the economy pattern: generation runs ahead of verification.
One evidential strength is worth naming plainly. When Perceptis goes beyond its sources, it labels the addition rather than dressing it as a cited fact. On that honesty measure it scores highly, at 92%.
Caliper Lab is an independent AI capability measurement institution. Products make capability claims and buyers cannot check them, while generic benchmarks measure academic reasoning and say nothing about real professional work. The Lab builds the evidence base that closes the gap.
Independence is structural, not asserted. The Lab publishes findings regardless of any commercial relationship, grounds every finding in real workflows, and applies the same measurement to every system it tests. This assessment is one entry in that base, and the pipeline behind it keeps running after publication.
Every figure above resolves to a definition or table here.
Each system received the same three professional briefs, the same pack of public research reports as source material, and the same slide template. Every generated deck was rendered slide by slide and decomposed into a text layer and an image layer. A scoring harness judged each slide and each deck against a reference standard drawn from a library of professional decks from leading strategy firms, across thirteen measurement families grouped into four blocks. Where a check was purely mechanical, such as counting titled charts, it was computed directly. Quality, regeneration consistency, and cost were kept as separate concepts throughout.
Field screening. The evaluation opened with a wider field including Microsoft Copilot and a GPT-based PowerPoint plug-in. Neither reliably completed the end-to-end task, a detailed brief plus source pack executed to a finished deck, under a comparable harness. Both were excluded at screening, and the deep comparison covers the two systems that cleared that bar.
| Measure | Perceptis | Claude for PowerPoint |
|---|---|---|
| Charts carrying a title | 84.7% | 50.0% |
| Charts carrying a unit signal | 82.0% | 50.0% |
| Native charts per deck | 3.5 | 1.5 |
| Editable diagrams per deck | 4.5 | 0.0 |
| Editable diagram rate | 100% | — |
| Deck visual consistency | 97.2% | 94.4% |
| Source-note consistency | 100% | 91.7% |
| Template adherence | 83.3% | 75.0% |
Comparison chart-labelling figures come from the decks on which that system produced charts. Neither system generated axis titles.
Structure and instructions.
| Measure | Perceptis | Claude for PowerPoint |
|---|---|---|
| Requirement coverage | 94.9% | 94.8% |
| Structural conformance | 97.6% | 96.3% |
| Content delivery | 93.8% | 95.2% |
| Slide-count compliance | 100% | 100% |
Message and narrative.
| Measure | Perceptis | Claude for PowerPoint |
|---|---|---|
| States a finding | 99.2% | 100% |
| Message specificity | 99.2% | 99.2% |
| Message concision | 54.4% | 85.2% |
| Slide density presentable | 55.6% | 79.8% |
| Internal consistency | 58.3% | 91.7% |
Message concision is scored on 71 of 72 slides for Perceptis: 6 fully concise, 65 partial. Internal consistency rests on six decks and is reported as directional.
| Measure | Perceptis | Claude for PowerPoint |
|---|---|---|
| Slides carrying a source | 89.8% | 65.3% |
| Numeric mentions with units | 73.6% | 72.2% |
| Argument soundness | 73.1% | 84.1% |
| Off-source honesty | 92.2% | 96.7% |
Claim-to-source support and exact number preservation scored below the comparison system, with wide variation across briefs. Reported as directional rather than as a single rate, because the spread across scenarios is too wide to fix a defensible figure.
Scope. Findings are based on automated analysis of AI-generated outputs during the evaluation period. Results reflect the specific systems, configurations, prompts, source material and template used; different configurations may produce materially different results, and findings should not be generalised beyond the task types and system versions evaluated.
Evidence thresholds. Scored findings rest on measures observed across all six decks per system. Measures observed on partial samples are either excluded or labelled as directional in the text.
Source data. Data has been obtained from sources believed reliable. Caliper Lab makes no representation as to the accuracy or completeness of third-party data and has accepted it without independent verification of the underlying source claims.
Not investment advice. This assessment is not investment advice, a recommendation to buy or sell any product or service, or an opinion on the fairness of any transaction. It informs procurement and product evaluation decisions and should be read in that context.
Temporal validity. Findings are valid as of July 2026. AI capabilities and platform features change frequently, and Caliper Lab assumes no obligation to update this assessment for later changes.
Independence. Caliper Lab conducted this evaluation independently and without commercial bias. No fee arrangement with any evaluated vendor has influenced the findings, methodology, scoring or conclusions, and standards are applied consistently across vendors regardless of commercial relationship.