Update, 29 September 2026. The levels have been re-sorted so that a question sits where its evidence structure puts it, and every question can now be asked in five answer formats instead of one.
The published runs on the results page (seeds 43–46) still use the previous layout. Runs on the new layout are in progress.
How is each level's question set now?
A level is defined by how much evidence a question needs, not by how often models get it right. Each subtype is assigned to the level whose evidence structure it has, and the answer key is always computed by a program from annotations, never written by the question generator.
| Level | Evidence structure | Subtypes | Where the answer comes from |
|---|---|---|---|
| L1 | A single event answers it. Usually one modality; when sound and picture are both involved, they point to the same answer. | easy_sound, easy_sound_object, easy_move, easy_spatial_near, easy_spatial_side, easy_goal NEW |
Sound labels, object movements, 3D positions, camera pose, recipe segments |
| L2 Cross-modal alignment |
A sound fixes one moment; a second signal gives the answer at that moment. Neither is enough alone. | low_sound_movement, low_spatial_near, low_spatial_side, low_purpose MOVED FROM L3 |
Sound timing with movement, 3D position, camera pose, or the narrated purpose of the action |
| L3 Inference |
The whole 30-second window has to be combined. | medium_count, medium_order, medium_spatial_extreme, medium_spatial_order, medium_goal NARROWED |
Event counts, narration order, final 3D positions, recipe segments |
| L4 Planning |
Decide what to do next; some questions join two scenes. | high_plan_next, high_plan_route |
The next recipe segment; camera pose and object position across two scenes |
The two kinds of goal question
Both ask what activity the camera wearer is working on. They differ only in the wrong options.
easy_goal(L1, recognition). The wrong options involve objects that are not in the clip. Recognising what is on screen is enough. Example: the answer is Prepare mango and a wrong option is Wash frying pan lid.medium_goal(L3, inference). Every wrong option shares objects with the right answer, so recognising objects cannot settle it; the sequence of actions has to. Example: Continue prepare puff pastry against Bake puff pastry and Filling Tarts.
The same idea on a second dataset
EPIC-Sounds has sound annotations only, with no objects, 3D or recipes. The levels are rebuilt there from sound evidence alone, on two 30-second clips (P02_133 15–45 s, P09_08 45–75 s), 27 questions in total.
| Level | Subtype | Questions | Example |
|---|---|---|---|
| L1 | easy_sound | 14 | Which of the following sounds can be heard during this clip? |
| L2 | low_sound_next NEW | 3 | Right after you hear metal and plastic knocking together, which of these sounds do you hear next? |
| L3 | medium_count | 7 | How many separate times is something being opened or closed heard in this clip? |
| L3 | medium_sound_order NEW | 3 | Put these three sounds in the order they are first heard in this clip. |
There is no L4 on EPIC-Sounds: planning questions need annotations this dataset does not have.
What is different from before?
Four things changed. Two subtypes were sitting in L3 without needing the whole window, one L1 subtype had wrong options that gave it away in the wrong direction, and the way a question is answered is now a variable of its own.
Why this was needed. Across seeds 43–46, L3 was not harder than L2. Gemini 3.5 scored 56% on L2 and 58% on L3; Qwen3-Omni scored 48% and 53%. L3 came out higher than L2 in six of the eight seed-by-model runs.
| Change | Before | Now | Reason |
|---|---|---|---|
| SPLIT Goal questions |
One subtype, medium_goal, in L3. Gemini 3.5 83%, Qwen3-Omni 73%. |
easy_goal in L1 and medium_goal in L3, separated by whether the wrong options can be ruled out by recognising objects. |
About half the old questions could be answered from the first few seconds. Near-duplicate options such as Tidy up against Clean up and Tidy are also removed. |
| MOVED Purpose questions |
medium_purpose in L3. Both models 70%. |
low_purpose in L2. The 67 questions are unchanged. |
The question is anchored on a single sound-timed moment, which is the L2 structure. It never needed the whole window. |
FIXEDeasy_sound wrong options |
Wrong options were sounds not annotated within 2 seconds of the clip, taken from the most frequent sounds in the video. | Wrong options must be absent for 30 seconds either side, must not be implied by an action in the clip, and are matched to the right answer in how common they are. Footsteps are no longer used. | Asked without the clip, models scored 11–18%, below the 25% of guessing. The wrong options were more plausible than the right answer. |
| NEW Answer format |
Every question was four-option multiple choice. | Each question is asked in five formats, described in the next section. | A multiple-choice score mixes what the model perceived with what the options gave away. |
| UNCHANGED | All other subtypes, their answer keys and their wrong options are identical to the published sets. L4 is unchanged. | ||
A problem found but not yet fixed
In easy_sound_object, low_sound_movement and the purpose questions, none of the wrong options appear in the clip. The question reduces to “which of these four is in the clip”, and the sound is not needed. A first check supports this: with the audio track replaced by silence, Gemini 3.5 scored 75% on low_sound_movement, against 76% with sound. The planned fix is to draw wrong options from other moments of the same clip.
How is a multiple-choice question turned into a short answer?
Only if it still has one clear answer once the options are taken away. If it does, the options are dropped, the question states the output format it expects, and the reply is scored by rules first. If it does not, the options are kept and each one is judged true or false.
The conversion follows the scheme of Beyond Multiple Choice: Verifiable OpenQA for Robust Vision-Language RFT (Liu et al., 2025): sort questions by answer type, rewrite each type its own way, score by rules wherever possible, and judge each option true or false when a question cannot stand without its options.
Step 1 · Can the question stand without its options?
| Subtype | Without options | Format used |
|---|---|---|
easy_sound_object, easy_move, low_sound_movement | One object, guaranteed unique by the question's own checks | Short answer |
medium_count | One integer | Short answer |
medium_order, medium_spatial_order, medium_sound_order | The three items are listed in the question; one order is right | Short answer |
easy_goal, low_purpose | One activity or purpose, phrased freely | Short answer |
high_plan_route | One direction | Short answer |
easy_sound, low_sound_next | “Which of these sounds” has many valid answers | True / false per option |
All spatial_near, spatial_extreme, spatial_side | The candidates are the tracked objects, which the model cannot know | True / false per option |
medium_goal, high_plan_next | Free text cannot separate stages of the same recipe; “next step” has no single granularity | True / false per option |
Step 2 · Rewrite the question and state the output format
| Answer type | Added to the question | Example |
|---|---|---|
| Number | Answer with a single integer. | How many separate times is metal hitting ceramic heard in this clip? Answer with a single integer. |
| Object | Answer with the name of the object. | At the moment you hear two plastic objects knocking together, which object is the camera wearer handling? Answer with the name of the object. |
| Order | List all three in order, separated by ' -> '. | Put these three actions in the order they happen in this clip: (put the last now peeled potato; move the peeler; peel the potato). List all three in order, separated by ' -> '. |
| Purpose or activity | Answer in one short, direct phrase. | At the moment you hear two plastic objects knocking together, what is the camera wearer trying to do? Answer in one short, direct phrase starting with 'to'. |
| Direction | Answer with the direction they should head. | … Which way should they head? Answer with the direction they should head. |
Short-answer questions use fixed templates. Wording that depends on a list, such as “which of the following”, is rejected by the checker.
Step 3 · Score by rules, and use a judge only when rules cannot decide
| Answer type | Rule |
|---|---|
| Number | The first integer in the reply equals the right count. |
| Direction | Direction words (left, right, straight, ahead, around, behind) map to one of four headings. |
| Order | The position of each of the three items in the reply gives the order. |
| Object | The reply contains the object's name or its head noun, and no wrong option's name. “A lid” counts for “glass lid”. |
| Purpose or activity | The reply contains the key words of the right answer, and none that belong only to a wrong option. |
When the rules give no verdict, for example “to clean up” for “to dry their hands”, Gemini 3.5 acts as judge. It sees only the question, the reply and the four known candidates in shuffled order, and returns which candidate the reply means, or none. It does not see the clip and is not told which candidate is right. Each row records whether it was scored by rule or by judge.
The five formats, on one question
Same clip, same right answer: roll of cling film.
1 · Yes / no
Chance 50%. Half the statements use the right answer, half a wrong option.
Statement: At the moment you hear two plastic objects knocking together, the camera wearer is handling the roll of cling film. Answer: <yes/no>
Right answer: yes
2 · Single choice
Chance 25%. The existing format.
At the moment two plastic objects knocking together is heard, which object is the camera wearer handling? A. roll of cling film B. measuring jar C. meat box D. chunk of minced beef
Right answer: A
3 · Select all that apply
Chance 7%. One to three of four statements are true.
A. … the camera wearer is handling the chunk of minced beef. B. … the camera wearer is handling the measuring jar. C. … the camera wearer is handling the roll of cling film. D. Rustling can be heard during this clip.
Right answer: C, D
4 · Short answer
Chance about 0%. No options shown.
At the moment you hear two plastic objects knocking together, which object is the camera wearer handling? Answer with the name of the object. Answer: <your answer>
Right answer: roll of cling film
5 · Short answer with explanation
Chance about 0%. Only the answer line is scored.
(same question as 4) Answer: <your answer> Because: <one sentence on what you saw or heard that supports your answer>
Right answer: roll of cling film
4b · True / false per option
Used in place of 4 and 5 when the question needs its options. Chance 6% for all four right.
At the moment metal hitting wood is heard, which object is closest to the saucepan? A. water bottle A: <true/false> B. yellow pepper B: <true/false> C. scissors C: <true/false> D. salt shaker D: <true/false>
Right answer: true, false, false, false
Why format 4b exists, and whether it is easier
It exists because some questions cannot be asked without their options. “Which object is closest to the saucepan” has no single answer a model could be expected to name: the candidates are the objects the annotators tracked, and the model cannot know which those are. Forcing such a question into a short answer would make the scoring unfair, so the options stay and only the way of answering changes.
One true/false judgement is easier than one multiple-choice pick: a guess is right half the time. The question as a whole is not. It counts as right only when all four judgements are right.
| Single choice | True / false per option | |
|---|---|---|
| What the model decides | Which one of four is best | Whether each of four is true, one at a time |
| What the model is told | Exactly one option is right | Nothing about how many are true |
| Can it answer by elimination? | Yes. Ruling out three leaves the answer. | Not by the instructions. Each option needs its own verdict. |
| Chance of a fully right answer | 25% | 6% |
The same paper reports what models do in practice: once options are judged one at a time, the comparison between options is lost and models mark too many statements true. Among wrongly answered questions, 63–86% had more than one statement marked true, across seven models (Liu et al., 2025, Appendix D). The same ratio will be reported here.
A known limit. The four options are the same as in single choice, so exactly one is true, and they exclude each other: only one object can be the closest. A model that works this out can treat 4b as single choice again. So 4b is a fallback for questions that need options, not an equal of the short answer, and it may turn out no harder than single choice. Both the per-option accuracy and the all-four-right rate are reported so that this can be checked.
Where the extra true statements in format 3 come from
- The right answer of the question itself is always one of the true statements.
- Further true statements come from other verified questions about the same 30-second window.
- If there are not enough of those, a statement about a sound that is present in the window is added.
- False statements come from the wrong options of the same questions, and from sounds that are absent.
How formats are compared
Chance differs by format, so scores are compared after correcting for it: (accuracy − chance) ÷ (1 − chance). Every format is also asked without the clip. Because the five formats share the same base question, each comparison is paired question by question.
Results
Runs in progress. This section will be filled in when they finish.
- Five formats on HD-EPIC: 40 questions per level, four levels, Gemini 3.5 and Qwen3-Omni, with and without the clip.
- The sound-only levels on EPIC-Sounds: 27 questions, five formats, both models.
- The silent-audio check on the three subtypes whose wrong options are all outside the clip.