> For the complete documentation index, see [llms.txt](https://docs.fastrouter.ai/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://docs.fastrouter.ai/explore-features/video-evaluations.md).

# Video Evaluations

Evaluate AI-generated videos at scale using LLM-based judges, with automated scoring across motion, sync, quality, and prompt adherence.

## **Introduction**

FastRouter's Video Evaluations feature lets you assess the quality of AI-generated videos at scale. By importing video generation logs, defining LLM-based judging criteria, and running evaluations against those outputs, you can systematically measure video quality across dimensions like visual fidelity, scene flow, on-screen text accuracy, audio-visual sync, and adherence to the original prompt.

Video Evals work within the same Custom Evaluations infrastructure as text and image evals; the same judge configuration, the same scoring rubrics, and the same results dashboard extended to support multimodal video output.

***

## **Key Benefits**

* Evaluate AI-generated video outputs automatically using a multimodal LLM judge.
* Import video generation logs directly from your FastRouter activity.
* Use the same Auto grader and Custom grader setup as text and image evaluations.
* Catch failures that are hard to spot at a glance: a duplicated logo, a typo in the final frame, a scene that skips a beat with per-dimension judge reasoning for every video.

***

## **Before You Start**

Video evals run against videos you have already generated through FastRouter. You can generate them via the API, or interactively from **Model Playground** by switching to the **Video** tab and selecting a video model (e.g. `x-ai/grok-imagine-video-1.5`).

<figure><img src="https://2466471311-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FzZfZz8wlCHOmP1FU2BsK%2Fuploads%2FbysJIm7eaqDogGd6XUMt%2FVideo%20Playground.png?alt=media&amp;token=30d10338-abf5-438d-bf4f-51100bdc7930" alt=""><figcaption><p>Generating a video in Model Playground</p></figcaption></figure>

Video prompts are usually long and structured. The example used throughout this page:

```
Create a 8-second luxury jewelry ad for Star Diamonds, promoting a new diamond earring range.
Elegant, cinematic, premium aesthetic with black velvet, soft champagne-gold lighting, sparkling
highlights, and smooth transitions. Photorealistic diamonds, accurate earring symmetry, crisp
product details, refined typography, no distorted faces or jewelry.

0–2 sec: Macro shot of diamond earrings emerging from darkness, slowly rotating as light catches
each facet. On-screen copy: "A New Light Has Arrived"
2–4 sec: Elegant model wearing the earrings, turning her head subtly against a sophisticated
black-and-gold background. On-screen copy: "Made to Make You Shine"
4–6 sec: Fast, graceful montage of three distinct designs—classic studs, diamond drops, and
contemporary hoops—each sparkling in close-up. On-screen copy: "Discover the New Earring Collection"
6–8 sec: Hero product arrangement on black velvet, followed by the Star Diamonds logo and a soft
diamond-flare finish. On-screen copy: "Star Diamonds" CTA: "Find Your Perfect Pair"

Audio: Sophisticated cinematic shimmer with a gentle rising beat and a delicate crystal chime on
the final logo reveal.
```

Writing the prompt as timed beats pays off twice: the generation model has a clearer target, and you can reuse the same structure as your judge's checklist.

***

## **Creating a Video Evaluation**

Navigate to the **Evaluations** section in your FastRouter dashboard and click **Create Evaluation**.

<figure><img src="https://2466471311-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FzZfZz8wlCHOmP1FU2BsK%2Fuploads%2FSFqlg9HLnHZoTHqfb5g1%2FNew%20Evaluation.png?alt=media&amp;token=f16ae8a0-159b-4fb3-860e-129994628083" alt=""><figcaption><p>New Evaluation</p></figcaption></figure>

**Step 1 — Name Your Evaluation**

Provide a descriptive name (e.g. `Video Eval` or `Jewellery-Ad-Compliance-v2`).

**Step 2 — Import Video Logs**

Click **Import Data**. In the Import Test Data dialog, select the **Videos** tab.

<figure><img src="https://2466471311-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FzZfZz8wlCHOmP1FU2BsK%2Fuploads%2FxpDOIIM6G4V5pdeXk24B%2FImport%20Test%20Data.png?alt=media&amp;token=aab2d9c8-664b-4b46-9e5a-d76b4301f531" alt=""><figcaption><p>Import Test Data — Videos tab</p></figcaption></figure>

Configure the following fields:

* **Date Range** *(required)*: Select the date range covering the video generations you want to evaluate.
* **Model** *(required)*: Choose the video generation model whose outputs you want to import (e.g. `x-ai/grok-imagine-video-1.5`).
* **Project**: Optionally filter by project. Select a project to narrow the available API keys, or leave as "All Projects" to see all keys.
* **Key**: Optionally filter by a specific API key used during generation.
* **Input contains**: Search for specific text in the generation prompt to narrow down which logs are imported.
* **Sampling rate (%)**: Set a percentage of matching logs to import (1–100%). Useful for large log sets — start with a smaller sample to validate your setup before scaling.

> ℹ️ Video file logs are available for import approximately **2 hours** after they are generated. If your recent video logs aren't showing up, try again later.

The dialog shows a live **Total rows** count as you adjust the filters. Click **Import** to load the video generation logs as your evaluation dataset.

Once imported, the dataset appears as **Imported from Video logs**, and the source model is added automatically under **Models to Compare**. The **Data (preview)** panel on the right shows each row's prompt alongside its video output.

<figure><img src="https://2466471311-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FzZfZz8wlCHOmP1FU2BsK%2Fuploads%2FduApwrWwEjcDpxUUSTEH%2FSample%20Overview.png?alt=media&amp;token=29646faa-9fd3-4527-9a00-8139e2f7594d" alt=""><figcaption><p>New Evaluation with imported video logs</p></figcaption></figure>

**Step 3 — Add Evaluation Metrics**

Click **Add Metric** and configure your judge. Two presets are available:

* **Auto grader**: a general-purpose quality judge with a ready-made rubric — good for a quick first pass.
* **Custom grader**: your own prompt, model, and scoring scale — recommended when you have a specific brief to check the video against.

<figure><img src="https://2466471311-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FzZfZz8wlCHOmP1FU2BsK%2Fuploads%2FHwuARaR2UlrQQa2TN6Xn%2FCustom%20Grader.png?alt=media&amp;token=f2c3f765-e221-4ff3-a828-c75bdedeed0a" alt=""><figcaption><p>Configuring a Custom grader</p></figcaption></figure>

Configure the following:

* **Model** *(required)*: Select a multimodal model that accepts video input (e.g. `google/gemini-3.1-pro-preview`). A text-only or image-only model cannot score video.
* **System**: Describe the evaluator's role, the dimensions to check, and how to score them. The example used here:

  ```
  You are a video evaluator. Watch the video and check if it matches these requirements:
  1. Visual Quality: Is it a luxury jewelry ad with black velvet, gold lighting, and symmetrical,
     distortion-free diamonds?
  2. Scene Flow: Does it progress smoothly from a macro shot, to a model, to a product montage,
     and end on a hero product layout?
  3. Text & Typo Check: Read all on-screen text. Check for any typos, spelling mistakes, or
     grammatically incorrect copy.
  Score from 1 - 7. Give an overall fail score of 2 or below if any of the requirements fails to
  meet the standard.
  ```
* **Variables**: Use `{{curly braces}}` to interpolate values into your prompt. `{{item.input}}` passes the original generation prompt to the judge, which is what you need for prompt-adherence checks. The generated video itself is attached to the judge call automatically. The valid variable list is shown under each prompt box.
* **User** *(required)*: The user-turn template sent to the judge, typically the user input placeholder followed by the response to evaluate.
* **Scoring**: Define the numeric **Range** (e.g. `0–7`) and a **Pass Threshold** (e.g. `3.5`). Rows scoring at or above the threshold are marked Pass; everything below is Fail.

**Step 4 — Select Evaluation API Key**

Choose an API key from your account. This key is used for all LLM judge calls during the evaluation.

**Step 5 — Run**

Click **Run**. FastRouter first shows a **Credit Utilization Estimate** with the number of selected requests, the estimated cost, and your current balance.

Click **Proceed with Evaluation** to start the run. The judge is applied asynchronously to each video in the dataset and results are returned as they complete.

> ℹ️ The estimate is based on the selected judge model and an average prompt size — actual costs vary with clip length and resolution. If credits run out mid-run, the evaluation may not complete.

***

## **Viewing Results**

Access results from the **Evaluations** listing page by clicking your evaluation. Use the **Report** / **Data** toggle in the top right to switch views.

**Report view** — Aggregated metrics across all rows: the run and its Request ID, latency percentiles, generation cost, and the overall score for each grader. The **Test Criteria** panel below shows the judge prompt, its range and pass threshold, the judge model, and the total judge spend for the run.

<figure><img src="https://2466471311-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FzZfZz8wlCHOmP1FU2BsK%2Fuploads%2FJAjiLt3yx4DYwXpNhlyF%2FReport%20Details.png?alt=media&amp;token=2e3d438c-412d-4441-90aa-489d5d01a31d" alt=""><figcaption><p>Report Details</p></figcaption></figure>

> ℹ️ Latency shows **N/A** for imported video generations. Video models generate asynchronously, so per-request latency isn't captured the way it is for chat or image runs — use cost and score to compare video models.

**Data view** — One row per video with its generation input on the left and the graded output on the right, including the score, cost, and a video attachment. Use **Single View** to inspect one run. Search, sort, and fullscreen controls help with larger datasets.

<figure><img src="https://2466471311-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FzZfZz8wlCHOmP1FU2BsK%2Fuploads%2FmmeBaVpzJwJ2lPFXveK0%2FSample%20Overview.png?alt=media&amp;token=280e55c2-93b3-4b3a-9f43-52182b6fff43" alt=""><figcaption><p>Runs Overview — Data view</p></figcaption></figure>

**Content Preview** — Click the expand arrow on any row to open an inline video player next to the full input parameters, with the grader verdict and cost alongside. Play the clip here to check the judge's call yourself.

<figure><img src="https://2466471311-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FzZfZz8wlCHOmP1FU2BsK%2Fuploads%2FK5IOxHxdHnSthjABjVAz%2FContent%20Preview.png?alt=media&amp;token=8cd0ee74-2559-4579-9e3f-d94261c210ef" alt=""><figcaption><p>Content Preview</p></figcaption></figure>

**Judge Reasoning** — Click the score to open **Criteria Results** and see the judge's per-dimension breakdown, along with the grader configuration that produced it.

<figure><img src="https://2466471311-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FzZfZz8wlCHOmP1FU2BsK%2Fuploads%2FoHLj0ixKtoFRkK6Kj3MZ%2FJudge%20Reasoning.png?alt=media&amp;token=51aa5fb2-04d1-49d0-8e9f-4c1c9a78af83" alt=""><figcaption><p>Judge Feedback</p></figcaption></figure>

**Example output for the Star Diamonds ad eval (`x-ai/grok-imagine-video-1.5`):**

| Metric                   | Value                                                                                                                        |
| ------------------------ | ---------------------------------------------------------------------------------------------------------------------------- |
| Custom Grader            | Fail: 2 (Range 0–7, Pass ≥ 3.5)                                                                                              |
| Latency                  | N/A                                                                                                                          |
| Original Generation Cost | μ$1,120,000.0000 ($1.12). **Note:** This cost is only incurred during the original generation and not during the evaluation. |
| Total Judge Score Cost   | $0.0220                                                                                                                      |
| Pass Rate                | 0%                                                                                                                           |
| Video Length             | 8 seconds (720p, 16:9, audio on)                                                                                             |

Judge reasoning summary:

1. **Visual Quality** — Has the intended luxury aesthetic with velvet and gold lighting, but fails the distortion check due to visual artifacts overlapping the model's face.
2. **Scene Flow** — Passes. Follows the intended sequence from macro shot to model to montage to hero arrangement.
3. **Text Check** — Mostly spelled correctly, but fails due to poorly formatted and duplicated "Star Diamonds" text in the final scene.

Two of three dimensions failed, and the rubric instructed the judge to cap the overall score at 2 whenever any requirement misses the standard; so the row is marked **Fail** even though scene flow was clean.

***

## **Tips & Best Practices**

* **Start with a small sample**: Use the sampling rate slider to import 10–20% of your logs first. Validate your judge prompt on a handful of videos before scaling to the full dataset.
* **Use a capable multimodal judge**: Video evaluation requires a model that can process video frames and audio. Choose models that support video input explicitly.
* **Mirror your prompt structure in your rubric**: If the generation prompt is written as timed beats, make the judge check those same beats. Named dimensions (visual quality, scene flow, text accuracy) produce far more consistent scores than a single open-ended quality question.
* **Ask the judge to read on-screen text**: Typos, duplicated logos, and malformed CTAs in the final frame are the most common — and most embarrassing — failures in generated ad content, and they are easy to miss on a first watch.
* **Be deliberate with harsh aggregation rules**: "Fail the whole video if any requirement misses" is the right call for a compliance gate but collapses pass rates to 0%. For tracking quality over time, prefer an average across dimensions so you can see movement between runs.
* **Allow 2 hours post-generation**: Video logs take approximately 2 hours to become available for import. Plan your eval runs accordingly.

***

## **Relationship to Custom Evaluations**

Video Evals are an extension of FastRouter's [Custom Evaluations](/explore-features/custom-evaluations.md) feature. The same infrastructure — dataset management, run comparison, judge configuration, and results dashboard — applies to both. The key difference is the data source: instead of importing chat completion logs or CSV files, you import video generation logs via the **Videos** tab in the Import Test Data dialog.

All judge configuration options available for text evals (scoring rubrics, variable interpolation, multi-criteria graders) are fully supported for video evals.


---

# Agent Instructions
This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com.

## Querying This Documentation
If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question.

Perform an HTTP GET request on the current page URL with the `ask` query parameter, and the optional `goal` query parameter:

```
GET https://docs.fastrouter.ai/explore-features/video-evaluations.md?ask=<question>&goal=<endgoal>
```

`ask` is the immediate question: it should be specific, self-contained, and written in natural language.
`goal` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal.

The response will contain a direct answer to the question and relevant excerpts and sources from the documentation.

Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections.
