Measured, not marketed.

The Standard for AI Video Quality

AQ Labs is an independent evaluation lab that benchmarks generative video models, builds bespoke training and eval datasets, and advises the teams shipping and buying them.

We score generative video models against a published, reproducible methodology — temporal coherence, prompt fidelity, artifact rate, motion realism — so your quality claims survive procurement, diligence, and the press. No vendor incentives. No cherry-picked reels. Just instrumented results you can cite.

41 models evaluated

Across text-to-video, image-to-video, and video-to-video pipelines.

128,400 clips scored

Every clip double-rated by trained human evaluators plus automated metrics.

6 benchmark dimensions

Weighted, documented, and versioned so results stay comparable over time.

Independent by design

We take no equity, revenue share, or placement fees from model providers.

Leaderboard · v4.2

The AQ Video Benchmark

A standing leaderboard of frontier and open-weight video models, refreshed each quarter under identical prompt sets, seeds, and rendering conditions. Scores are decomposed by dimension so you can see exactly where a model wins and where it breaks.

Temporal coherence evaluation framesTC · 87.4

Temporal Coherence

Object permanence, identity drift, and frame-to-frame stability across long generations.

Field median87.4 / 100
Prompt fidelity evaluation framesPF · 79.1

Prompt Fidelity

Semantic adherence to subject, action, camera language, and negative constraints.

Field median79.1 / 100
Artifact rate evaluation framesAR · 63.8

Artifact Rate

Frequency and severity of warping, ghosting, limb collapse, and text degradation.

Field median63.8 / 100
Motion realism evaluation framesMR · 71.5

Motion Realism

Physical plausibility of movement, contact, momentum, and scene dynamics.

Field median71.5 / 100
Aesthetic consistency evaluation framesAC · 82.0

Aesthetic Consistency

Lighting, color, and style continuity held across shots and durations.

Field median82.0 / 100
Safety and leakage evaluation framesSL · 91.2

Safety & Leakage

Rate of policy-violating output and training-data regurgitation under adversarial prompts.

Field median91.2 / 100
Engagements

Three ways to work with the lab

Every engagement starts with a scoped brief and ends with an artifact you can hand to a board, a buyer, or an engineering lead. Fixed deliverables, named timelines, senior staff on the work.

Video Benchmarking & Certification

Private model evaluation or public leaderboard entry, delivered as a signed audit report in 3–4 weeks.


  • Sample intake and prompt-set alignment
  • Standard protocols, fixed seeds, identical render conditions
  • Signed audit report plus raw scoring data
Custom experiments

À La Carte Dataset Creation

Bespoke training, fine-tuning, and eval sets — captured or curated, licensed clean, annotated to your schema.


  • Original capture or curated multi-source collection
  • Annotation to your rubric, with QA sampling at every tier
  • Provenance and licensing documented per asset
View pricing

AI Advisory

Retained expert guidance on eval strategy, model selection, build-versus-buy, and internal quality gates.


  • Eval strategy and internal quality-gate design
  • Model selection and build-versus-buy analysis
  • Dedicated senior evaluation-lead time
Check availability
Diligence review workstation

Buyer-Side Diligence

Independent validation of a vendor's claims before you sign a seven-figure contract.

À la carte

Configure your dataset

Specify modality, volume, annotation depth, and licensing terms and we assemble a scoped brief with indicative timeline and price band — typically returned within one business day. Contracts are priced per volume and annotation tier; enterprise SOWs are finalized offline.

Dataset
ModalityVideo, video-audio pairs, multi-camera, synthetic, or mixed-source collections.
VolumeFrom 2,000-clip evaluation sets to multi-million-clip training corpora.
Annotation DepthCaptioning, dense temporal labels, camera-motion tags, preference rankings, or custom rubrics.
LicensingFully cleared, model-release-backed, and indemnified — with provenance documented per asset.
Cross-references

Do I provide my own footage?

Either. We capture original material under model release, curate from cleared libraries, or annotate assets you already own — pipelines run on our calibrated evaluation stack in isolated client environments.


Milestones

  • Scoped brief, schema alignment, and pilot batch within one business day of intake.
  • Final delivery with raw and processed channels, per-asset provenance, and QA agreement scores.

Compare pipelines
Methodology

Why our numbers hold up

Our methodology is published in full, including prompt sets, weighting, rater qualification, and inter-rater agreement. Clients get the raw scoring data alongside the report, so any result can be reproduced or contested.

Open Methodology

Published & versioned

Versioned scoring rubrics and prompt suites available for public review.

Evaluation rater session

Rater Rigor

Qualified human evaluators with tracked agreement scores; disputed clips escalate to expert adjudication.

Secure evaluation environment

Security & Compliance

SOC 2 Type II controls, GDPR-aligned processing, NDA-first onboarding, and isolated client environments.

Senior evaluation bench

Senior Bench

Evaluation leads drawn from frontier research labs, computer-vision academia, and post-production supervision.

Proven Outcomes

Model-side engagement

A generative video lab cut artifact rate 34% in two fine-tune cycles after our diagnostic eval.

Procurement-Ready

Buyer-side engagement

An enterprise media team used our diligence report to reject two of three shortlisted vendors — before purchase.

Research

Research & next steps

Download State of AI Video, our quarterly analysis of model performance, cost-per-usable-second, and failure modes across the field. When you're ready for numbers on your own model, book a 30-minute scoping call — we respond to every qualified request within one business day.

State of AI Video report cover

State of AI Video — Q4

Full leaderboard, dimension breakdowns, and year-over-year quality trend lines.

Download the report
Eval design playbook

Eval Design Playbook

How to build an internal video quality gate that engineering and legal both accept.

Get the playbook
Scoping call

Scoping Call

30 minutes with an evaluation lead to define scope, timeline, and deliverables.

Book a scoping call
Direct line to the lab

Direct Line

hello@aqlabs.org — reviewed by a senior lead, not a queue.

hello@aqlabs.org

Measured, not marketed.

Bring us a model, a shortlist, or a dataset spec. We return a scoped brief with indicative timeline and price band within one business day.

Request an Evaluation