Vendor Radar Top
Caliper Lab.
Independent AI capability measurement
Platform assessment

Assessment of AI Presentation Platform - Perceptis

Visual constructionStructureMessageEvidence
Two systems · four measurement blocks · July 2026
01 · Method

A two-stage evaluation, judged against real consulting work.

The evaluation opened with a wider field. Every system was given the same task: take a detailed brief and a pack of source reports, and build a finished deck end to end. Microsoft Copilot and a GPT-based PowerPoint plug-in could not do this reliably, and were set aside at screening.

The two systems that cleared that bar, Perceptis and Claude for PowerPoint, then ran the same three briefs, six decks each, scored across thirteen measurement families. The metrics are drawn from a reference library of professional decks from leading strategy firms, not from opinion.

Exhibit 1 · How the field narrowed
STAGE 1 · VISUAL SCREEN Can it take a brief and source pack to a finished deck? Perceptis Claude for PowerPoint Copilot GPT slide plug-in Could not execute reliably · set aside STAGE 2 · DEEP EVALUATION Same three briefs · six decks each · 13 families Like-for-like comparison Perceptis vs Claude for PowerPoint
Three briefs · six decks per system · public source packs · July 2026
PerceptisAI presentation platform
Purpose-built deck generation. Brief plus source pack in, editable PPTX out.
Claude for PowerPointFrontier LLM · add-in
Latest model in the PowerPoint add-in. The closest capable comparable.
Microsoft CopilotScreened out
Could not carry the brief and source pack to a finished deck under a comparable setup.
GPT slide plug-inScreened out
Same limitation at the screening stage. Excluded from the deep comparison.
02 · Scorecard

Perceptis leads on visual construction and is at parity on structure. Its gaps are message economy and claim support.

Four blocks, two systems. Read across them, the two systems have distinct profiles. Perceptis behaves like a deck production system, building visual structure by default. Claude for PowerPoint behaves like a strong writer working inside PowerPoint. Perceptis is clearly ahead on how the deck is built, level on following the brief, and behind on message economy and how firmly each claim is tied to its source.

Exhibit 2 · Four-block scorecard
Block Perceptis Claude for PowerPoint
Visual construction
Built to a professional visual standard?
STRONG
85% of charts titled · 8 visuals per deck
MODERATE
50% of charts titled · 1.5 visuals per deck
Structure & instructions
Did it build what the brief asked for?
STRONG
94.9% requirement coverage
STRONG
94.8% requirement coverage
Message & narrative
Sharp slides, deck that holds together?
MOD–STRONG
specific findings, but runs long
STRONG
specific and concise
Evidence & rigor
Claims anchored, argument sound?
MODERATE
90% of slides sourced · claim support varies
MOD–STRONG
65% of slides sourced · 84% argument soundness

The sections that follow take each block in turn, from the surface of the deck inward: first how it is built, then whether it did what was asked, then how it reads, and finally how well it holds up to scrutiny.

03 · Visual construction

Visual construction is Perceptis's clearest strength: more visuals, and finished properly.

A chart without a title or units forces the reader to reconstruct what they are looking at. Finished visuals separate a deck that can be presented from one that needs a cleanup pass. This is the widest gap in the evaluation, and it favours Perceptis.

85%
of Perceptis charts carry a title, against 50% for Claude for PowerPoint
8.0
visuals per deck: 3.5 native charts plus 4.5 editable diagrams, against 1.5 charts and no diagrams
100%
of Perceptis diagrams built from editable shapes, not flat images
Exhibit 3 · Chart labelling
TitleUnit signal 84.7%82% 50%50% 0%50%100% SHARE OF GENERATED CHARTS CARRYING EACH ELEMENT Perceptis Claude
Comparison figures come from the decks on which that system produced charts. Neither system generated axis titles, a gap shared across the category.
Exhibit 4 · Visuals produced per deck
0369 3.54.51.5 PerceptisClaude Editable diagrams Native charts VISUAL OBJECTS PER DECK · SIX DECKS PER SYSTEM
Every Perceptis diagram is editable, letting a consultant adjust a deck in minutes rather than rebuild it. This reflects what Perceptis is: a system that builds presentation structure by default.
Exhibit 5 · Deck-level consistency
Deck visual consistencySource-note consistencyTemplate adherence 97.2%100%83.3% 94.4%91.7%75% 0%20%40%60%80%100% DECK-LEVEL CONSISTENCY · SIX DECKS PER SYSTEM · PERCEPTIS BLUE, CLAUDE GREY
Deck consistency combines typography, palette, spacing, components, template and source notes. Slide-level polish (alignment, contrast, readability) is high for both systems; full detail in A2.
04 · Structure and instructions

On doing what the brief asked, the two systems are at parity.

Before a deck can be judged on how it looks or reads, it has to deliver what was requested: the right sections, the right content, the right number of slides. This is the competence floor of the category. Both systems clear it, and clear it together.

Exhibit 6 · Following the brief
050100 94.997.693.8 94.896.395.2 CoverageConformanceContent Perceptis Claude for PowerPoint
Pooled across three briefs and six decks per system. Both systems hit the requested slide count on every deck.

There is no meaningful separation here. Both systems deliver the requested content at a high rate and match the requested structure.

This is also the bar that decided the field. The two systems set aside at screening could not carry a detailed brief and a source pack through to a finished deck. Clearing it reliably is what put Perceptis and Claude for PowerPoint into the deep evaluation.

The only visible movement is in conditional instructions, those that apply only when a stated condition holds, where Perceptis is slightly behind. The difference rests on a handful of cases and is not a stable gap.

05 · Message and narrative

The messages are specific and right. Messages can be made more tighter.

Perceptis states real findings in its slide titles, at parity with the comparison system, and the slide beneath supports them. Message economy is the gap versus the frontier. Senior readers move through a deck in seconds per slide, and messages that run long tax exactly that audience.

Exhibit 7 · Message sharpness
States a findingMessage is specific 99.2%99.2% 100%99.2% 0%60%100% PERCEPTIS BLUE · CLAUDE GREY · 72 SLIDES PER SYSTEM
Message-to-content alignment is perfect for both systems. What a Perceptis title claims, the slide underneath delivers.
Exhibit 8 · Message economy
Message is conciseSlide density presentable 54.4%55.6% 85.2%79.8% 0%60%100% PERCEPTIS BLUE · CLAUDE GREY · POOLED ACROSS ALL SLIDES
Almost nothing fails outright. Almost everything lands one edit short of tight: correct, specific, and longer than it should be. A compression pass is the single highest-return change identified.
06 · Evidence and rigor

The sources are on the slide. The gap is whether each claim is supported by them.

There is a difference between a slide that carries a source note and a claim that survives being checked against that source. The first is a formatting discipline, and Perceptis leads on it. The second is an evidential discipline, and it is Perceptis's main gap versus the frontier.

Exhibit 9 · Sourcing surface and argument
Slides carrying a sourceArgument soundness 89.8%73.1% 65.3%84.1% 0%60%100% PERCEPTIS BLUE · CLAUDE GREY · POOLED ACROSS THREE BRIEFS
Perceptis places a source on far more slides than the comparison system. On argument soundness, whether conclusions follow from premises without overclaiming, it trails by eleven points.

On the harder check, whether each individual claim is fully supported by the source it cites, Perceptis scored lower, and the result varied widely from one brief to the next.

The Lab reports this as a direction rather than a single figure, because the spread across briefs is too wide to fix a defensible number. It is consistent with the economy pattern: generation runs ahead of verification.

One evidential strength is worth naming plainly. When Perceptis goes beyond its sources, it labels the addition rather than dressing it as a cited fact. On that honesty measure it scores highly, at 92%.

07 · Caliper Lab

The Lab measures what AI products can actually do, and publishes it regardless of who benefits.

Caliper Lab is an independent AI capability measurement institution. Products make capability claims and buyers cannot check them, while generic benchmarks measure academic reasoning and say nothing about real professional work. The Lab builds the evidence base that closes the gap.

Independence is structural, not asserted. The Lab publishes findings regardless of any commercial relationship, grounds every finding in real workflows, and applies the same measurement to every system it tests. This assessment is one entry in that base, and the pipeline behind it keeps running after publication.

For vendors
Third-party evidence that a capability claim is real and measurable. Independent findings carry weight that self-reported benchmarks cannot.
For buyers
A repeatable basis for judging AI tools on real work, with results that compare across tools and tasks rather than vendor decks.
For investors
A way to test whether a capability premium priced into a valuation is real and durable, turning an assertion into evidence.
Appendix

Method, definitions, full tables and limits.

Every figure above resolves to a definition or table here.

A1 · Method in full

Each system received the same three professional briefs, the same pack of public research reports as source material, and the same slide template. Every generated deck was rendered slide by slide and decomposed into a text layer and an image layer. A scoring harness judged each slide and each deck against a reference standard drawn from a library of professional decks from leading strategy firms, across thirteen measurement families grouped into four blocks. Where a check was purely mechanical, such as counting titled charts, it was computed directly. Quality, regeneration consistency, and cost were kept as separate concepts throughout.

Field screening. The evaluation opened with a wider field including Microsoft Copilot and a GPT-based PowerPoint plug-in. Neither reliably completed the end-to-end task, a detailed brief plus source pack executed to a finished deck, under a comparable harness. Both were excluded at screening, and the deep comparison covers the two systems that cleared that bar.

A2 · Four-block definitions
Visual construction
Is the deck built to a professional visual standard? Chart labelling (title and unit signal), visual richness (native charts and editable diagrams per deck), and deck-level consistency across typography, palette, spacing, components, template and source notes.
Structure and instructions
Did it build what the brief asked for? Requirement coverage, structural conformance across counts and layout, and delivery of the requested content.
Message and narrative
Is each slide sharp and does the deck hold together? Whether messages state a finding, are specific, and are concise; slide density; and deck-level narrative coherence.
Evidence and rigor
Are the claims anchored and the argument sound? Source presence on slides, whether each claim is supported by its cited source, argument soundness, and honest labelling of off-source additions.
A3 · Visual construction detail
MeasurePerceptisClaude for PowerPoint
Charts carrying a title84.7%50.0%
Charts carrying a unit signal82.0%50.0%
Native charts per deck3.51.5
Editable diagrams per deck4.50.0
Editable diagram rate100%
Deck visual consistency97.2%94.4%
Source-note consistency100%91.7%
Template adherence83.3%75.0%

Comparison chart-labelling figures come from the decks on which that system produced charts. Neither system generated axis titles.

A4 · Structure and message detail

Structure and instructions.

MeasurePerceptisClaude for PowerPoint
Requirement coverage94.9%94.8%
Structural conformance97.6%96.3%
Content delivery93.8%95.2%
Slide-count compliance100%100%

Message and narrative.

MeasurePerceptisClaude for PowerPoint
States a finding99.2%100%
Message specificity99.2%99.2%
Message concision54.4%85.2%
Slide density presentable55.6%79.8%
Internal consistency58.3%91.7%

Message concision is scored on 71 of 72 slides for Perceptis: 6 fully concise, 65 partial. Internal consistency rests on six decks and is reported as directional.

A5 · Evidence and rigor detail
MeasurePerceptisClaude for PowerPoint
Slides carrying a source89.8%65.3%
Numeric mentions with units73.6%72.2%
Argument soundness73.1%84.1%
Off-source honesty92.2%96.7%

Claim-to-source support and exact number preservation scored below the comparison system, with wide variation across briefs. Reported as directional rather than as a single rate, because the spread across scenarios is too wide to fix a defensible figure.

A6 · Qualifications and limiting conditions

Scope. Findings are based on automated analysis of AI-generated outputs during the evaluation period. Results reflect the specific systems, configurations, prompts, source material and template used; different configurations may produce materially different results, and findings should not be generalised beyond the task types and system versions evaluated.

Evidence thresholds. Scored findings rest on measures observed across all six decks per system. Measures observed on partial samples are either excluded or labelled as directional in the text.

Source data. Data has been obtained from sources believed reliable. Caliper Lab makes no representation as to the accuracy or completeness of third-party data and has accepted it without independent verification of the underlying source claims.

Not investment advice. This assessment is not investment advice, a recommendation to buy or sell any product or service, or an opinion on the fairness of any transaction. It informs procurement and product evaluation decisions and should be read in that context.

Temporal validity. Findings are valid as of July 2026. AI capabilities and platform features change frequently, and Caliper Lab assumes no obligation to update this assessment for later changes.

Independence. Caliper Lab conducted this evaluation independently and without commercial bias. No fee arrangement with any evaluated vendor has influenced the findings, methodology, scoring or conclusions, and standards are applied consistently across vendors regardless of commercial relationship.

Caliper Lab · Independent AI capability measurement Assessment of AI Presentation Platform - Perceptis · July 2026 · thecaliperlab.com