HD-EPIC QUESTION PYRAMID ← RESULTS EXPLORER

Update, 29 September 2026. The levels have been re-sorted so that a question sits where its evidence structure puts it, and every question can now be asked in five answer formats instead of one.

The published runs on the results page (seeds 43–46) still use the previous layout. Runs on the new layout are in progress.

How is each level's question set now?

A level is defined by how much evidence a question needs, not by how often models get it right. Each subtype is assigned to the level whose evidence structure it has, and the answer key is always computed by a program from annotations, never written by the question generator.

LevelEvidence structureSubtypesWhere the answer comes from
L1 A single event answers it. Usually one modality; when sound and picture are both involved, they point to the same answer. easy_sound, easy_sound_object, easy_move, easy_spatial_near, easy_spatial_side, easy_goal NEW Sound labels, object movements, 3D positions, camera pose, recipe segments
L2
Cross-modal alignment
A sound fixes one moment; a second signal gives the answer at that moment. Neither is enough alone. low_sound_movement, low_spatial_near, low_spatial_side, low_purpose MOVED FROM L3 Sound timing with movement, 3D position, camera pose, or the narrated purpose of the action
L3
Inference
The whole 30-second window has to be combined. medium_count, medium_order, medium_spatial_extreme, medium_spatial_order, medium_goal NARROWED Event counts, narration order, final 3D positions, recipe segments
L4
Planning
Decide what to do next; some questions join two scenes. high_plan_next, high_plan_route The next recipe segment; camera pose and object position across two scenes

The two kinds of goal question

Both ask what activity the camera wearer is working on. They differ only in the wrong options.

The same idea on a second dataset

EPIC-Sounds has sound annotations only, with no objects, 3D or recipes. The levels are rebuilt there from sound evidence alone, on two 30-second clips (P02_133 15–45 s, P09_08 45–75 s), 27 questions in total.

LevelSubtypeQuestionsExample
L1easy_sound14Which of the following sounds can be heard during this clip?
L2low_sound_next NEW3Right after you hear metal and plastic knocking together, which of these sounds do you hear next?
L3medium_count7How many separate times is something being opened or closed heard in this clip?
L3medium_sound_order NEW3Put these three sounds in the order they are first heard in this clip.

There is no L4 on EPIC-Sounds: planning questions need annotations this dataset does not have.

What is different from before?

Four things changed. Two subtypes were sitting in L3 without needing the whole window, one L1 subtype had wrong options that gave it away in the wrong direction, and the way a question is answered is now a variable of its own.

Why this was needed. Across seeds 43–46, L3 was not harder than L2. Gemini 3.5 scored 56% on L2 and 58% on L3; Qwen3-Omni scored 48% and 53%. L3 came out higher than L2 in six of the eight seed-by-model runs.

ChangeBeforeNowReason
SPLIT
Goal questions
One subtype, medium_goal, in L3. Gemini 3.5 83%, Qwen3-Omni 73%. easy_goal in L1 and medium_goal in L3, separated by whether the wrong options can be ruled out by recognising objects. About half the old questions could be answered from the first few seconds. Near-duplicate options such as Tidy up against Clean up and Tidy are also removed.
MOVED
Purpose questions
medium_purpose in L3. Both models 70%. low_purpose in L2. The 67 questions are unchanged. The question is anchored on a single sound-timed moment, which is the L2 structure. It never needed the whole window.
FIXED
easy_sound wrong options
Wrong options were sounds not annotated within 2 seconds of the clip, taken from the most frequent sounds in the video. Wrong options must be absent for 30 seconds either side, must not be implied by an action in the clip, and are matched to the right answer in how common they are. Footsteps are no longer used. Asked without the clip, models scored 11–18%, below the 25% of guessing. The wrong options were more plausible than the right answer.
NEW
Answer format
Every question was four-option multiple choice. Each question is asked in five formats, described in the next section. A multiple-choice score mixes what the model perceived with what the options gave away.
UNCHANGED All other subtypes, their answer keys and their wrong options are identical to the published sets. L4 is unchanged.

A problem found but not yet fixed

In easy_sound_object, low_sound_movement and the purpose questions, none of the wrong options appear in the clip. The question reduces to “which of these four is in the clip”, and the sound is not needed. A first check supports this: with the audio track replaced by silence, Gemini 3.5 scored 75% on low_sound_movement, against 76% with sound. The planned fix is to draw wrong options from other moments of the same clip.

How is a multiple-choice question turned into a short answer?

Only if it still has one clear answer once the options are taken away. If it does, the options are dropped, the question states the output format it expects, and the reply is scored by rules first. If it does not, the options are kept and each one is judged true or false.

The conversion follows the scheme of Beyond Multiple Choice: Verifiable OpenQA for Robust Vision-Language RFT (Liu et al., 2025): sort questions by answer type, rewrite each type its own way, score by rules wherever possible, and judge each option true or false when a question cannot stand without its options.

Step 1 · Can the question stand without its options?

SubtypeWithout optionsFormat used
easy_sound_object, easy_move, low_sound_movementOne object, guaranteed unique by the question's own checksShort answer
medium_countOne integerShort answer
medium_order, medium_spatial_order, medium_sound_orderThe three items are listed in the question; one order is rightShort answer
easy_goal, low_purposeOne activity or purpose, phrased freelyShort answer
high_plan_routeOne directionShort answer
easy_sound, low_sound_next“Which of these sounds” has many valid answersTrue / false per option
All spatial_near, spatial_extreme, spatial_sideThe candidates are the tracked objects, which the model cannot knowTrue / false per option
medium_goal, high_plan_nextFree text cannot separate stages of the same recipe; “next step” has no single granularityTrue / false per option

Step 2 · Rewrite the question and state the output format

Answer typeAdded to the questionExample
NumberAnswer with a single integer.How many separate times is metal hitting ceramic heard in this clip? Answer with a single integer.
ObjectAnswer with the name of the object.At the moment you hear two plastic objects knocking together, which object is the camera wearer handling? Answer with the name of the object.
OrderList all three in order, separated by ' -> '.Put these three actions in the order they happen in this clip: (put the last now peeled potato; move the peeler; peel the potato). List all three in order, separated by ' -> '.
Purpose or activityAnswer in one short, direct phrase.At the moment you hear two plastic objects knocking together, what is the camera wearer trying to do? Answer in one short, direct phrase starting with 'to'.
DirectionAnswer with the direction they should head.… Which way should they head? Answer with the direction they should head.

Short-answer questions use fixed templates. Wording that depends on a list, such as “which of the following”, is rejected by the checker.

Step 3 · Score by rules, and use a judge only when rules cannot decide

Answer typeRule
NumberThe first integer in the reply equals the right count.
DirectionDirection words (left, right, straight, ahead, around, behind) map to one of four headings.
OrderThe position of each of the three items in the reply gives the order.
ObjectThe reply contains the object's name or its head noun, and no wrong option's name. “A lid” counts for “glass lid”.
Purpose or activityThe reply contains the key words of the right answer, and none that belong only to a wrong option.

When the rules give no verdict, for example “to clean up” for “to dry their hands”, Gemini 3.5 acts as judge. It sees only the question, the reply and the four known candidates in shuffled order, and returns which candidate the reply means, or none. It does not see the clip and is not told which candidate is right. Each row records whether it was scored by rule or by judge.

The five formats, on one question

Same clip, same right answer: roll of cling film.

1 · Yes / no

Chance 50%. Half the statements use the right answer, half a wrong option.

Statement: At the moment you hear two plastic objects knocking together, the camera wearer is handling the roll of cling film.

Answer: <yes/no>

Right answer: yes

2 · Single choice

Chance 25%. The existing format.

At the moment two plastic objects knocking together is heard, which object is the camera wearer handling?
A. roll of cling film
B. measuring jar
C. meat box
D. chunk of minced beef

Right answer: A

3 · Select all that apply

Chance 7%. One to three of four statements are true.

A. … the camera wearer is handling the chunk of minced beef.
B. … the camera wearer is handling the measuring jar.
C. … the camera wearer is handling the roll of cling film.
D. Rustling can be heard during this clip.

Right answer: C, D

4 · Short answer

Chance about 0%. No options shown.

At the moment you hear two plastic objects knocking together, which object is the camera wearer handling? Answer with the name of the object.

Answer: <your answer>

Right answer: roll of cling film

5 · Short answer with explanation

Chance about 0%. Only the answer line is scored.

(same question as 4)

Answer: <your answer>
Because: <one sentence on what you saw or heard that supports your answer>

Right answer: roll of cling film

4b · True / false per option

Used in place of 4 and 5 when the question needs its options. Chance 6% for all four right.

At the moment metal hitting wood is heard, which object is closest to the saucepan?
A. water bottle      A: <true/false>
B. yellow pepper     B: <true/false>
C. scissors          C: <true/false>
D. salt shaker       D: <true/false>

Right answer: true, false, false, false

Why format 4b exists, and whether it is easier

It exists because some questions cannot be asked without their options. “Which object is closest to the saucepan” has no single answer a model could be expected to name: the candidates are the objects the annotators tracked, and the model cannot know which those are. Forcing such a question into a short answer would make the scoring unfair, so the options stay and only the way of answering changes.

One true/false judgement is easier than one multiple-choice pick: a guess is right half the time. The question as a whole is not. It counts as right only when all four judgements are right.

Single choiceTrue / false per option
What the model decidesWhich one of four is bestWhether each of four is true, one at a time
What the model is toldExactly one option is rightNothing about how many are true
Can it answer by elimination?Yes. Ruling out three leaves the answer.Not by the instructions. Each option needs its own verdict.
Chance of a fully right answer25%6%

The same paper reports what models do in practice: once options are judged one at a time, the comparison between options is lost and models mark too many statements true. Among wrongly answered questions, 63–86% had more than one statement marked true, across seven models (Liu et al., 2025, Appendix D). The same ratio will be reported here.

A known limit. The four options are the same as in single choice, so exactly one is true, and they exclude each other: only one object can be the closest. A model that works this out can treat 4b as single choice again. So 4b is a fallback for questions that need options, not an equal of the short answer, and it may turn out no harder than single choice. Both the per-option accuracy and the all-four-right rate are reported so that this can be checked.

Where the extra true statements in format 3 come from

How formats are compared

Chance differs by format, so scores are compared after correcting for it: (accuracy − chance) ÷ (1 − chance). Every format is also asked without the clip. Because the five formats share the same base question, each comparison is paired question by question.

Results

Runs in progress. This section will be filled in when they finish.

  • Five formats on HD-EPIC: 40 questions per level, four levels, Gemini 3.5 and Qwen3-Omni, with and without the clip.
  • The sound-only levels on EPIC-Sounds: 27 questions, five formats, both models.
  • The silent-audio check on the three subtypes whose wrong options are all outside the clip.