VGI-Bench (Video General Intelligence Bench, by Seldon) asks multiple-choice questions about long-form videos: what appeared first, what was on screen at the same time as something else, what a narrator said while a specific shot was up, and questions whose premise never happens. Each sample below is a real run through OpenRouter. The model received the full video and the question in one turn at temperature 0, and the graded answer is the letter on its final Answer: line.
Last benchmark run
Accuracy over the public split (439 questions, runs need at least 395 answered to be listed), with cost, time, and input and output tokens per question for each provider endpoint the model was run on.
| # | Model | Std dev | |||||
|---|---|---|---|---|---|---|---|
| 1 | Seed | 68.6% | -- | $0.066 | 2.1m | 228k | 4.55k |
| 2 | ByteDance Seed: Seed 1.6 Pareto | 67.3% | -- | $0.041 | 68s | 148k | 2.21k |
| 3 | Alibaba | 66.7% | -- | $0.058 | 4.7m | 413k | 11.3k |
| 4 | Alibaba | 63.6% | -- | $0.62 | 8.9m | 584k | 19.9k |
| 5 | 61.4% | ±0.0pp | $0.27 | 2.6m | 231k | 17.2k | |
| 6 | Google AI Studio | 60.7% | -- | $0.059 | 2.1m | 341 | 5.88k |
| 7 | Alibaba | 58.0% | -- | $0.33 | 8.2m | 579k | 13.6k |
| 8 | Google AI Studio | 55.4% | -- | $0.14 | 3.8m | 108k | 40.1k |
| 9 | 54.7% | -- | $0.25 | 2.9m | 152k | 22.8k | |
| 10 | Alibaba | 54.2% | -- | $0.26 | 5.8m | 1.48M | 11k |
| 11 | 51.6% | -- | $0.15 | 2.0m | 153k | 16.5k | |
| 12 | Alibaba | 51.0% | -- | $0.21 | 7.4m | 517k | 19.5k |
| 13 | Google: Gemini 2.5 Flash Pareto Google AI Studio | 49.8% | -- | $0.023 | 2.3m | 341 | 9.17k |
| 14 | 48.9% | ±1.4pp | $0.030 | 2.3m | 37.7k | 6.93k | |
| 15 | Alibaba | 47.4% | -- | $0.10 | 4.8m | 1.5M | 20.1k |
| 16 | Xiaomi: MiMo-V2.5 Pareto Xiaomi | 47.2% | ±2.0pp | $0.020 | 5.0m | 208k | 13.1k |
| 17 | Perceptron | 46.0% | -- | $0.008 | 1.9m | 45.8k | 1.01k |
| 18 | Alibaba | 46.0% | -- | $1.14 | 6.8m | 1.49M | 19.9k |
| 19 | 45.6% | -- | $0.14 | 5.6m | 368k | 16.8k | |
| 20 | 45.6% | -- | $0.027 | 4.6m | 77.4k | 14.7k | |
| 21 | 45.4% | -- | $0.034 | 5.2m | 337k | 11.9k | |
| 22 | 44.8% | -- | $0.17 | 8.4m | 517k | 21.4k | |
| 23 | Xiaomi | 44.3% | -- | $0.040 | 4.7m | 75.7k | 22.1k |
| 24 | 43.4% | -- | $0.34 | 3.2m | 77k | 21.1k | |
| 25 | 40.5% | -- | $0.024 | 1.9m | 211k | 7.3k | |
| 26 | 34.6% | ±2.7pp | $0.028 | 8.3m | 58.2k | 43.9k | |
| 27 | 28.3% | ±0.8pp | $0.10 | 13.4m | 52.4k | 97.3k | |
| 28 | Reka Edge Pareto Reka | 21.2% | -- | $0.00015 | 37s | 1.48k | 29.6 |
| 29 | 13.8% | -- | $0.025 | 14.9m | 11k | 25.1k | |
| 30 | 11.8% | -- | $0.084 | 10.3m | 114k | 2.33k | |
| 31 | Alibaba | 11.6% | ±0.5pp | $0.029 | 3.9m | 30.5k | 5.69k |
| 32 | 9.4% | -- | $0.00034 | 38s | 2.6k | 4.61k | |
| 33 | BaseTen | 5.3% | ±5.9pp | $0.028 | 36s | 4.49k | 1.46k |
| 34 | Darkbloom | 2.1% | -- | $0.003 | 11.4m | 612 | 27.9k |
| 35 | Z.AI | 1.4% | -- | $0.004 | 61s | 22k | 227 |
The reasoning trace is the model's perception: the moments it says it saw and the timestamps it attaches to them. Reading it against the video shows how the answer was reached, and on the failed runs, where the description parts from the footage.
Happy path: The model locates the moment and reads it correctly.
Question 1187 global-graph/interaction_partner Qwen3.8 27B
The full video the model was sent, served from the OpenRouter mirror used by the harness.
Which option does the mouse pointer hover over most often?
The trace tracks the pointer across the whole recording rather than one frame, names the Premiere Pro cameo near the end as a distractor, and lands on the window that dominates the session.
Failure mode: Salient object over the key
Question 1862 global-graph/co_presence Qwen3.8 27B
The full video the model was sent, served from the OpenRouter mirror used by the harness.
Which option is repeatedly visible at the same time as the wooden garden shed during the video?
The trace cites eight timestamps for a wagon wheel leaning against the shed and dismisses the presenter as incidental. The human-validated key is the presenter. The model chose the most distinctive object in the frame over the one the question was scored on.
Failure mode: Narrates a moment that never happens
Question 1096 contrastive-hard-negative/false-premise-visual-attribute Qwen3.8 27B
The full video the model was sent, served from the OpenRouter mirror used by the harness.
When the speaker says she used a full tablespoon of salt, where is the red measuring spoon positioned relative to the cheese curds?
The key says this moment does not occur. The trace narrates a 2:38 to 2:44 sequence with the spoon set down in front of the curds, rejects the 'does not happen' option outright, and answers with detail the key does not support.
Failure mode: Answers an audio question from world knowledge
Question 2852 contrastive-hard-negative/visual_audio Qwen3.8 27B
The full video the model was sent, served from the OpenRouter mirror used by the harness.
When the two M.2 slots are highlighted in purple, what click is heard as the top heatsink locks back into place?
The trace says 'I need to recall the audio' and then reasons from what PC build videos 'typically' sound like, never from this one. The key is 'does not happen'. The final response presents the guessed click as something heard in the video.
Failure mode: Locates a near-duplicate moment
Question 1539 near-duplicate-moments Qwen3.8 27B
The full video the model was sent, served from the OpenRouter mirror used by the harness.
When the hand inserts a black power cable into the Seasonic power supply at the rear of the black PC case, what surface is behind it?
The question pins a moment that recurs with small differences. The trace settles on a rear view at 15:30 to 15:40 with a white wall behind it. The key is a dark textured surface, so the model matched the action to a similar shot and read the background off that one.
Failure mode: Never locates the spoken cue
Question 2894 contrastive-hard-negative/transcript_visual Qwen3.8 27B
The full video the model was sent, served from the OpenRouter mirror used by the harness.
When he says, “This whole time he's back there drinking whiskey,” is the whiskey bottle to the left or right of the laptop on Craig’s bench?
The cue is a spoken line. The trace never finds it, says it 'probably occurs during the conversation at the bench', and answers from a frame at 60 to 70 seconds. The key is 'does not happen'. Without the transcript anchor, the model substituted the nearest matching scene.