HD-EPIC QUESTION PYRAMID

Four levels of audio-visual questions over egocentric kitchen video. Gold answers are re-derived from annotations, SLAM poses and 3D geometry, so the generator never writes the answer key. Every run on disk is below.

Why every question is asked twice

A multiple-choice score on its own cannot separate perception from prior knowledge. Four options put chance at 25%, and a plausible distractor set can hand the answer over on its own — the object nearest a butter knife is usually the butter.

So each question is also asked blind: same wording, same options, same gold answer, with the clip withheld. What the clip was worth is the difference.

VIDEO GAIN

WHAT THE MODEL GETS WHEN A QUESTION IS ASKED BLIND
 

No video, no audio, no annotations — not even which video or which 30 seconds it came from. The wording and the options are generated by the same code as the video run; only the clip is withheld.

Coverage

One row per seed × arm × input tier × model. A dash means that run was never executed.

Does difficulty rise with the level?

Every run drawn as one line: solid where the model watched the clip, dashed where the same questions were asked blind. If the level number tracked difficulty, the solid lines would fall left to right. The same numbers are in the coverage table above.

Blind is the identical question, options and gold answer with the clip withheld, so it scores what the wording and the option list give away on their own — four options already put chance at 25%, and a distractor set often leaks more. The dashed line is the floor a level cannot fall below; only the distance above it was earned by watching. Why every question is asked twice ↗

How much does the seed move the score?

Every seed draws a fresh question set from the same videos and annotations. If a result only holds for one draw, it is not a result.

Filled dots are the video tier, hollow dots the same questions with the clip withheld. The short bar is the mean of the seeds; the vertical line spans them. Curated-set runs are left out, since they use a different question set.

What each level asks

Levels encode evidence structure, not measured difficulty, so accuracy is not expected to fall monotonically. Per-subtype numbers pool every seed and arm of this version on the video tier.

The questions

Both models' picks are marked G and Q on the options. Blind asks the same question with the clip withheld; video gain is the difference.

Choose a question to inspect both models' answers.