2,612
evaluations have shaped AdAlign's scoring engine
Self-improving scoring, powered by every evaluation
Every evaluation makes AdAlign smarter
Before AdAlign, ad-to-landing-page alignment was a manual check — inconsistent, subjective, and impossible to scale. AdAlign replaces that with a structured scoring engine that improves with every analysis.
Every evaluation generates signal. Which dimensions matter most per vertical. Where the biggest alignment gaps occur by platform. What separates high-performing ad/page combinations from the rest. That data feeds a continuous research loop — an autonomous system that tests scoring variations against real performance data, keeps what improves accuracy, and discards the rest.
The result: a scoring engine calibrated across platforms and industries, refined by real-world data, not assumptions.

How It Works
Every ad-to-page pair is scored across four dimensions, weighted by platform, and refined by real-world feedback.
Visual Match
40% defaultCompares the visual language of the ad creative against the landing page — color palette, imagery style, layout density, and brand consistency.
Message Match
20% defaultEvaluates whether the headline, value proposition, and key claims in the ad carry through to the landing page without friction or disconnect.
Above the Fold Continuity
20% defaultMeasures how easily a user can follow the trail from ad promise to landing page fulfillment — CTA alignment, content hierarchy, and navigational clarity.
Tone Consistency
20% defaultChecks that the emotional register — formal vs. casual, urgent vs. relaxed, playful vs. serious — stays coherent from ad to page.
Platform-Specific Weights
Different platforms reward different strengths. AdAlign adjusts dimension weights per platform so scores reflect what actually drives performance.
| Surface | Visual | Message | Above fold | Tone |
|---|---|---|---|---|
| Default | ||||
| Default | 40% | 20% | 20% | 20% |
| Search | 10% | 35% | 35% | 20% |
| Shopping | 35% | 27.5% | 27.5% | 10% |
| Display & Demand Gen | 40% | 20% | 20% | 20% |
| Video | 35% | 17.5% | 17.5% | 30% |
| Meta | ||||
| Feed | 40% | 22.5% | 22.5% | 15% |
| Stories & Reels | 30% | 20% | 20% | 30% |
| TikTok | ||||
| TikTok | 30% | 17.5% | 17.5% | 35% |
| 40% | 22.5% | 22.5% | 15% | |
| 35% | 20% | 20% | 25% | |
Band-First Scoring
Instead of picking an arbitrary number, the AI first classifies each dimension into a qualitative band, then assigns a precise numeric score within that band's range. This prevents score drift and keeps results consistent.
Grounded in the AdAlign Brain
The scoring engine isn't a generic LLM wrapper. Every rubric, weight and edge-case rule lives in a curated knowledge graph — the AdAlign Brain — built from platform docs, academic research, ad-to-page teardowns and calibration sessions. Every recommendation cites at least one source.
This is where the flywheel lands. When an evaluation exposes a scoring mistake, the fix is written into the wiki as an edge case, compiled into the prompts, and applied to every score that follows.
Explore the AdAlign Brain- 40Edge cases
- 18Scoring rubrics
- 14Verticals
- 8Platform notes
- 6Concepts
How the wiki becomes the score
Sources
Platform docs from Meta and Google, academic papers, UX research, ad-to-page teardowns and calibration sessions on real ads. Never edited — only read.
Wiki
Each source is compiled into linked articles: concepts, scoring rubrics, platform notes, vertical profiles and edge cases. Every article cites the sources it came from.
Prompts
The rubrics and vertical profiles are compiled from the wiki into the scoring prompts at build time. Nothing is retrieved on the fly, so every score runs on the same reviewed rules.
Latest Calibration
Based on 34 anonymized evaluations · Oct 5, 2026
No adjustments needed — scoring is well-calibrated for current data.