# Which video model should you trust? A vendor-neutral benchmark

Entry E17. Sector: Internal research. Status: internal. Delivered: 2026-06.
Source: https://ai.prospicience.in/work/video-model-benchmark

Prospicience ran the same approved keyframe through six AI video provider and model combinations and scored each on identity preservation, scene consistency, dialogue suitability and reliability. The result is a vendor-neutral routing guide: which model to use for dialogue close-ups, which for instructional action and long sequences, and which for cheap iteration.

Speed: Run and reported in a single day.

## The challenge

Every video model demo looks excellent, because the vendor picked the clip. Choosing between them for real production work needs the same input run through all of them.

## Why it mattered

Picking a video model from a demo reel means discovering its weak spots mid-production, when every regenerated shot is a real charge and a face that changes between shots breaks a whole sequence. The choice sets the cost and quality of every shot that follows.

## What we built

- Created one shared approval board: a reference sheet, character sheets, a location and an approved scene keyframe.
- Generated the same keyframe through six provider and model combinations, using timed prompts.
- Scored each on identity preservation, scene consistency, expression and dialogue suitability, and operational reliability.
- Wrote up every result, weak spots included, as a routing guide.

## The result

- A routing guide by need: one model family for dialogue and expression close-ups, another for instructional action and stitching longer sequences.
- All six outputs held the scene, none failed outright, and the differences that matter are on record.
- A confirmed prompting method for timed, image led video shots.
- A fixed baseline, so a new model release can be scored against the same input in an afternoon.

## Built for trust

Every model receives an identical approved input, so the comparison measures the model rather than the prompt. Weak spots are recorded at the same weight as strengths, which is what makes the routing guide safe to act on.

## AI at work

AI video models generate the same approved scene six ways, and people score each result on faces, room, expression and reliability. Every later production then sends each shot to the model that holds it best.

What it produces: Six clips from one approved keyframe, Scores on four dimensions, A routing guide by type of shot, A confirmed method for timed shots.

- Model choice is settled before production, so weak spots are not discovered on a paid shot.
- The fixed baseline means a new model release can be scored against the same input in an afternoon.
- Productions start from the routing guide rather than a fresh round of trials.

How it works day to day: One approved keyframe and a timed prompt go to each model. AI generates a short shot from it, a person reviews the six side by side and scores them, and the scores become the guide the studio routes by.

## Figures from the delivered system

- provider and model combinations: 6
- shared approved input: 1 keyframe
- scoring dimensions: 4
- additional consistency tests: 2
- outputs that failed outright: none of 6

## The technology

- One approved image, generated through six different services
- Scored on whether faces, rooms and expressions survive
- Reliability counted as well as quality
- The failures are written down too

AI models used: Seedance, Grok Imagine.

### Technical notes

- Shared approval board generated once and reused as the fixed input across all runs
- Identity preservation, scene consistency, expression and dialogue suitability, and operational reliability scored separately rather than as one subjective rating
- Separate character consistency and reference driven generation tests alongside the main comparison
- Results retained as artefacts so a later model can be scored against the same baseline

## Questions

### How do you choose an AI video model for production?

Run the same approved input through every candidate and score it on what actually matters: does the face survive, does the room stay the same, does the expression suit the line, and does the service respond reliably. That tells you more in a day than any vendor demo.

### Why does identity preservation matter so much?

Because it decides whether a sequence holds together. A slightly soft frame is acceptable. A character whose face changes between two consecutive shots breaks the film, so it is the first thing we score.

### Does the best model stay the best?

Models change every few months, which is why the inputs and results are kept. With a fixed baseline, re-scoring a new release takes an afternoon rather than starting the comparison from scratch.

Capabilities: ai-media-generation, ai-governance-and-guardrails.
