In-house backtesting platform

Backtesting Platform

GEO results shouldn't rest on anyone's word. This platform poses real scenario prompts to the mainstream AI engines on a schedule, logging prompt by prompt whether your brand gets named and your content cited.

Backtesting platform overview — overall AI visibility score, citation rate and a metric radar
The instrument behind the managed service — our in-house backtesting platform: 1,000+ concurrent queries a minute across ChatGPT, Gemini, Claude, Perplexity, DeepSeek, Qwen and Meta AI, every raw answer collected and quantified into six metrics — first-recommendation rate, recommendation rate, mention rate, answer accuracy, sentiment and message alignment — plus citation rate and perception drift.

Step 1 · Setup

Decide how to test

A backtest starts from the questions your customers actually ask. Engines, models, prompt mode and question count are all set here.

Backtest setup — engine and model selection across ChatGPT, Claude, Gemini and Perplexity, standard/grounded/compare modes, up to 40 scenario prompts with PAA-grounded generation
Setting up a backtest: pick engines and models, choose standard / grounded / compare mode, and load up to 40 scenario prompts — AI-generated and PAA-grounded if you like.

Step 2 · Run

The five-stage backtest pipeline

Once submitted, multiple LLMs answer concurrently while the five-stage pipeline runs. This is the live progress view:

Parallel LLM queries — ChatGPT, Gemini, Claude, Perplexity, DeepSeek, Qwen and Meta AI get the same scenario prompts, answers checked for whether your brand is named; Content quality & sentiment analysis — sentiment, technical and structured-data signals scored together; Authority signals cross-checked against competitor share of voice; Visibility metrics — each engine's answers converge into comparable AI-visibility metrics; Three-axis scoring — final scores consolidated and archived.

Step 3 · Results

Who named you, at a glance

Every raw answer from every engine is collected per prompt and checked for naming and citations — good or bad, it's all on the table.

Cross-engine testing — actual naming results from ChatGPT, Gemini and Perplexity on scenario prompts such as 'managed GEO vendors in Taiwan'
Cross-engine testing: the same scenario prompts posed to different engines, showing which ones actually name you. Third-party brands blurred.

Step 4 · Competitors

On the same question, who beat you

Getting named is only half the answer. The other half is who got named instead of you, on which engine. Every competitor gets scored the same way you do, so the gap is a number rather than a feeling.

Cross-platform competitor matrix — visibility scores for your brand and its competitors across ChatGPT, Claude, Gemini, Google AI Overviews, Perplexity, Copilot, Meta AI, Grok, Mistral, DeepSeek, Qwen, Kimi, GLM, Doubao, Hunyuan, MiniMax, ERNIE and iFlytek Spark; red means a competitor is ahead of you
Cross-platform competitor matrix: the same score for you and every competitor, engine by engine. Red means they're ahead of you — the losses stay on the table too. A dash means that domain has no valid data on that engine, which is not the same as zero. Sample data, measured against nike.com.
Visibility × sentiment quadrant — every brand plotted by how often AI engines name it and how positively they describe it
Visibility × sentiment quadrant: horizontal is how often the engines name you, vertical is how positively they describe you. Top right is being both seen and well spoken of; being visible but poorly described is its own problem. Sample data, measured against nike.com.

Next step

How do I get access?

The backtesting platform is the verification instrument behind the managed service — the results in every managed-client report come from here. To see it run on your own brand, book a demo.

Further reading

A plain-text version of this page for AI search, RAG and agents: /en/backtest/index.md