How the questions are made. Every question is about one 30-second clip of first-person kitchen video with sound. A program decides what is asked and what the right answer is; Gemini 3.5 only writes the sentence.
From annotation to question
The same five steps produce every question at every level. Levels are independent: a higher-level question never uses the answer of a lower one.
1
Official annotations
HD-EPIC's human annotations: sounds, narrations, object movements, 3D positions, recipe steps, camera pose, eye-gaze priming.
Dataset
2
Mine a fact
A rule finds something that is true in one clip and unique there. It outputs the right answer and three wrong options.
Program
3
Write the sentence
Gemini 3.5 gets the fact and writes one question. It cannot change the answer or the options.
Gemini 3.5
4
Check
The answer is computed again on the same clip. Wording that leaks the answer or drops the moment is rejected.
Program
5
Assemble
100 questions per level. The right answer is spread evenly over A, B, C and D within each question type.
Program
If Gemini's sentence fails the check, a fixed template is used instead, for at most 20% of a set.
The four levels
A level is defined by how much evidence a question needs, not by how often models get it right. Each level has non-spatial and spatial questions. In every sample the right answer is marked.
A single event answers it. Usually one modality; when sound and picture are both involved they point to the same answer.
Facts come from
Sound events
Object movements
3D object positions
Camera pose
Recipe segments
The goal Gemini is given
“one piece of evidence, read off directly — either the sound and a visible object confirm each other, or a single placement answers it”
Rules that matter most here
For “sound present”: never name or hint at any sound in the question.
For spatial questions: keep the moment (“when the X is put down”) and never say where any object is.
Non-spatial
Sound presenteasy_sound
Which of the following sounds can be heard during this clip?
Ametal hitting wood
Bchopping
Cstirring or whisking
Dan electronic beep
Answer
A sound event lies fully inside the clip.
Wrong options
Sounds with no event within 30 s of the clip, similar in how common they are.
Sound to objecteasy_sound_object
You can hear sizzling or boiling in this clip. Which object visible on screen is producing it?
Abag of potatoes
Bsaucepan
Cpan
Dpotato
Answer
Exactly one visible object can make the sound that is heard.
Wrong options
Objects that can make the same sound but are not in the clip.
Object movedeasy_move
Which object is moved from the hob to the kitchen counter during the clip?
Aglass lid
Bcan of olive oil
Ccloth
Dolive oil can cap
Answer
The only movement in the clip that goes from fixture A to fixture B.
Wrong options
Objects that took the same route elsewhere in the video.
Activity, by recognitioneasy_goal
What activity is the camera wearer working on in this clip?
APrepare onion
BContinue preparing pizza dough
CSet the table
DPreheat oven
Answer
The recipe segment that contains the whole clip.
Wrong options
Other segments whose objects do not appear in the clip.
Spatial
Nearest objecteasy_spatial_near
At the moment the napkin is put down, which of these objects is closest to it?
Aspoon
Btumbler
Cstack of keep cups
Dblender
Answer
3D distance from the put-down point to every visible object.
Wrong options
Visible objects at least a set margin further away.
Left / right of the wearereasy_spatial_side
At the moment the napkin is put down, which of these objects is furthest to the camera wearer's left?
Afrying pan
Bbowl
Cmeasuring cup
Dgreen chopping board
Answer
Object positions rotated into the camera frame at the put-down moment.
Wrong options
Other visible objects, clearly less far to that side.
L2
Cross-modal alignment
A sound fixes one moment; a second signal gives the answer at that moment. Neither is enough alone.
Facts come from
Sound events
Object movements
Narrations
3D object positions
Camera pose
The goal Gemini is given
“two complementary cues: the sound gives the time anchor, the picture gives the answer”
Rules that matter most here
Refer to the sound only by its given phrase.
For purpose questions: point to the moment by the sound only; never name the action or the object.
For left / right questions: the reference is the camera wearer, and the side must not be flipped.
Non-spatial
Sound-timed handlinglow_sound_movement
At the moment two plastic objects knocking together is heard, which object is the camera wearer handling?
Aroll of cling film
Bmeasuring jar
Cmeat box
Dchunk of minced beef
Answer
The one movement within 1.5 s of a sound that occurs once.
Wrong options
Objects not in the clip.
Purpose of the actionlow_purpose
At the moment two plastic objects knocking together is heard, what is the camera wearer trying to do?
Ato wipe the counter surface
Bto pour out the water collected in it
Cto hold the sponge
Dto retrieve the bowl
Answer
The purpose clause of the narration at the sound-timed moment (“… to dry the hands”).
Wrong options
Purposes of similar actions elsewhere in the video.
Spatial
Sound-timed nearest objectlow_spatial_near
At the moment metal hitting wood is heard, which object is closest to the saucepan?
Awater bottle
Byellow pepper
Cscissors
Dsalt shaker
Answer
Position of the handled object at the sound, interpolated along its movement; then 3D distance.
Wrong options
Two visible objects further away, one object not in the clip.
Sound-timed left / rightlow_spatial_side
At the moment two metal objects knocking together is heard, which of the following is furthest to the camera wearer's right?
Abottle of cold water
Bice tray
Cspoon
Dblender
Answer
Camera pose at the sound; object positions rotated into the camera frame.
Wrong options
Other visible objects, clearly less far to that side.
L3
Inference
The whole 30-second clip has to be combined: how often, in what order, where things end up, what it all adds up to.
Facts come from
Sound events
Narrations
3D object positions
Recipe segments
The goal Gemini is given
“inference over the whole clip — what the overall activity is, how often something happens, the order of events, or where things end up”
Rules that matter most here
For order questions: list the three items in the given shuffled order.
For activity questions: never name, paraphrase or hint at any option.
For spatial questions: ask about where things are “by the end of the clip”.
Non-spatial
Counting a soundmedium_count
How many separate times is metal hitting ceramic heard in this clip?
A3
B2
C1
D4
Answer
Number of events of one sound, all fully inside the clip, 2 to 6.
Wrong options
Four consecutive counts; the true one in a random slot.
Order of actionsmedium_order
In what order do these events occur in the clip? (place the hand towel, stir the saucepan, pick up the hand rag)
Aplace the hand towel -> stir the saucepan -> pick up the hand rag
Bstir the saucepan -> place the hand towel -> pick up the hand rag
Cstir the saucepan -> pick up the hand rag -> place the hand towel
Dpick up the hand rag -> place the hand towel -> stir the saucepan
Answer
Start times of three distinct narrated actions.
Wrong options
Other orders of the same three actions.
Activity, by inferencemedium_goal
Taking the whole clip into account, what is the camera wearer mainly working on?
ADry hands
BStir saucepan and add water to it
CStir onions and transfer to bowl
DRetrieve garlic and ginger, add to saucepan and stir
Answer
The recipe segment that contains the whole clip.
Wrong options
Other segments that share objects or actions with the right one.
Spatial
Furthest / nearest at the endmedium_spatial_extreme
By the end of the clip, which object is closest to the large shopping bag?
Abag of mozzarella
Bbag of salad
Cbox of food
Dsingle cream pot
Answer
3D distances at the last instant of the clip.
Wrong options
The other visible objects.
Order by distancemedium_spatial_order
By the end of the clip, what is the order of these objects from closest to furthest to the napkin? (kettle, blender, measuring cup)
Ameasuring cup -> blender -> kettle
Bkettle -> blender -> measuring cup
Cmeasuring cup -> kettle -> blender
Dblender -> measuring cup -> kettle
Answer
Three objects with clearly separated distances at the end of the clip.
Wrong options
Other orders of the same three objects.
L4
Planning
Decide what happens next. The answer is never shown: the clip stops before it, or the question joins two scenes.
Facts come from
Recipe segments
Camera pose
3D object positions
Object movements
Eye-gaze priming
The goal Gemini is given
“planning — what the camera wearer will do next, where to go to fetch something seen earlier, or which way to head”
Rules that matter most here
For two-scene questions: open with the given sentence that says the video joins two scenes.
For route questions: the options are fixed headings; never describe where anything is.
For anticipation questions: never name a location or an object, and never say where the wearer is looking.
Non-spatial
Next stephigh_plan_next
What is the camera wearer planning to do next after completing the actions shown in the clip?
ATidy up
BAdd second slice of bread
CWashing the utensils
DBring tumbler and pour mango lassi into it
Answer
The recipe segment that follows; the clip is the last 30 s of the current one.
Wrong options
The current step as a trap, and steps from other videos.
Spatial
Route across two sceneshigh_plan_route
This video shows two scenes from the same kitchen, recorded about a minute apart. At the end of the second scene, which way should the camera wearer head to retrieve the mango bits seen in the first scene?
Aturn around
Bturn left
Cturn right
Dstraight ahead
Answer
Camera pose at the end of scene 2 and the 3D point where the object was left in scene 1.
Wrong options
The other three headings.
Gaze anticipation
Anticipate the put-downhigh_gaze_destination
At the end of this clip the camera wearer is holding the glass lid. Where are they about to put it down?
Athe oven
Bthe floor
Cthe hob
Dthe kitchen counter
Answer
Destination fixture of a put-down the wearer looked at 1.5 s or more in advance. The clip ends 1 s before it.
Wrong options
Where the object came from, as a trap, and other fixtures.
Anticipate the pick-uphigh_gaze_pickup
Which object is the camera wearer about to pick up right after this clip ends?
Aslice of pepper
Bspoon
Clid of pot
Dphone
Answer
The next object picked up, looked at 1.5 s or more in advance. The clip ends 1 s before it.
Wrong options
Visible objects that are not picked up in the next 10 s.
What Gemini is asked to do
One request covers ten facts of the same level. The request has five parts; only the first differs between levels.
1
Task and level goal
“For each fact packet below, write ONE four-choice question.” Followed by the goal sentence of the level, shown in each level above.
Differs by level
2
Hard rules
About 35 lines, the same for every level. The answer and options are fixed. The question must not contain the answer. Sounds are named only by the given phrase. No words such as “annotation” or “label”.
Same for all
3
Style examples
Three real questions of the same level, plus matching examples for spatial question types.
Same for all
4
How the model did one level below
From level 2 up: the score, accuracy by question type, and up to six missed questions. Gemini is told to word questions in the styles the model found hard, without changing any fact.
Level 2 and up
5
The facts
Ten packets. Each has the right answer, the three wrong options, and the fields its type needs.
{
"question": "At the moment two plastic objects
knocking together is heard, which object
is the camera wearer handling?",
"answer_echo": "roll of cling film"
}
The answer must be copied back unchanged. If it does not match the program's answer, the question is dropped. Settings: temperature 0, fixed seed, thinking off.
The fixed templates
Used when Gemini's sentence is rejected, and for every question in the newer formats. Words in braces are filled from the fact.
Level
Type
Template
L1
Sound to object
You can hear {sound} in this clip. Which object visible on screen is producing it?
L1
Nearest object
The camera wearer puts down {reference} during this clip. Which of these objects does it end up closest to?
L2
Sound-timed handling
At the moment you hear {sound}, which object is the camera wearer handling?
L2
Sound-timed left / right
At the moment you hear {sound}, which of these objects is furthest to the camera wearer's {side}?
L3
Counting a sound
How many separate times is {sound} heard in this clip?
L3
Furthest at the end
By the end of this clip, which of these objects ends up furthest away from {reference}?
L4
Next step
Judging from what the camera wearer is doing in this clip, what are they most likely to do next?
L4
Anticipate the put-down
At the end of this clip the camera wearer is holding {object}. Where are they about to put it down?
Is the Gemini that writes the question independent of the Gemini that answers it?
The calls are independent. The model is not: the same model writes, answers and judges.
What is independent
Gemini writing the question
Sees: the right answer, the three wrong options, and the fields of the fact
Does not see: the clip
Every request starts from nothing
Gemini answering the question
Sees: the question, the four options, and the 30-second clip
Does not see: which option is right, or anything from the writing step
Every question is a separate request; nothing carries over between questions
So the answer cannot reach the answering model through a shared conversation or cache.
What is not independent
WRITES THE SENTENCEGemini 3.5
PICKS THE FACTS (OPTIONAL)Gemini 3.5
ANSWERS, AS ONE OF TWO MODELS TESTEDGemini 3.5
JUDGES SHORT ANSWERSGemini 3.5
Qwen3-Omni, the other model tested, takes no part in writing or judging.
Possible bias
How it could arise
What limits it
Wording suits Gemini
The sentences are in Gemini's own style, which it may read more easily than Qwen3-Omni does.
The answer and options are fixed by the program. Four of the five formats use fixed templates, not Gemini's wording.
The writer sees earlier scores
From level 2 up, the writing request includes the tested models' score and missed questions from the level below.
The scores of both models are pooled, not Gemini's alone. Only the wording can change.
The judge is also a test-taker
Gemini judges its own short answers and Qwen3-Omni's, and may be more lenient to phrasing like its own.
The judge sees no clip and is not told which candidate is right. It is used only when rules cannot decide.
What the existing data can and cannot show
Question wording
Questions
Gemini 3.5
Qwen3-Omni
Written by Gemini
1,589
55%
45%
Fixed template
11
27%
36%
Seeds 43–46, with video. This cannot settle the question: only 11 questions use a template, and they are the ones whose Gemini wording was rejected, so they are not a fair sample.
Checks planned
1
Compare across formats
Single choice is worded by Gemini; the other four formats are templates. If Gemini's lead over Qwen3-Omni shrinks on the template formats, wording is part of the lead.
No extra calls
2
Check the judge by hand
Read 50 of the judge's verdicts, split between the two models' answers, and count disagreements.
No extra calls
3
Template-only rerun
Ask the same questions with template wording to both models. Run only if checks 1 and 2 show a difference.
About 800 calls per video
One question, five formats
The clip, the right answer and the wrong options stay the same. Only the way of answering changes, so a difference in score comes from the format alone.
One verified factright answer + three wrong options
1
Yes / no
One statement, built from the right answer or from a wrong option, half and half.
chance 50%
2
Single choice
The original four options.
chance 25%
3
Select all that apply
Four statements, one to three of them true. Extra true ones come from other facts about the same clip.
chance 7%
?
Does it still have one answer without the options?
Yes4
Short answer
No options. The question states the output format.
5
Short answer with explanation
Adds one sentence of support. Only the answer is scored.
chance about 0%
No4b
True / false per option
The options stay. Each is judged on its own, and all four must be right.
chance 6%
Which question types can drop their options
Short answer
One object: sound to object, object moved, sound-timed handling, anticipate the pick-up
One number: counting a sound
One order: order of actions, order by distance (the three items are listed in the question)
One phrase: activity by recognition, purpose of the action
One direction: route across two scenes
True / false per option
Sound present: many sounds are heard, so there is no single answer
All nearest, furthest and left / right questions: the candidates are the tracked objects, which the model cannot know
Activity by inference, next step: free text cannot separate stages of one recipe
Anticipate the put-down: several fixtures share a name
The same question in each format
1
Yes / no
Statement: At the moment you hear two plastic objects knocking together, the camera wearer is handling the roll of cling film.
yes
2
Single choice
At the moment two plastic objects knocking together is heard, which object is the camera wearer handling?
A. roll of cling film
B. measuring jar
C. meat box
D. chunk of minced beef
A
3
Select all that apply
A. … is handling the chunk of minced beef.
B. … is handling the measuring jar.
C. … is handling the roll of cling film.
D. Rustling can be heard during this clip.
C, D
4
Short answer
At the moment you hear two plastic objects knocking together, which object is the camera wearer handling? Answer with the name of the object.
roll of cling film
5
Short answer with explanation
(same question)
Answer: <your answer>
Because: <one sentence on what you saw or heard>
roll of cling film
4b
True / false per option
At the moment metal hitting wood is heard, which object is closest to the saucepan?
A. water bottle → true / false
B. yellow pepper → true / false
C. scissors → true / false
D. salt shaker → true / false
true, false, false, false
How a short answer is scored
1
Rules first
Number: the first integer. Direction: left, right, straight, around. Order: where each item appears in the reply. Object: its name or head noun, and no wrong option's name.
Program
2
Judge, only if rules cannot decide
Gemini 3.5 sees the question, the reply and the four candidates in shuffled order, and says which one the reply means, or none. It does not see the clip or know which is right.
Gemini 3.5
3
Record
Each answer is marked as scored by rule or by judge, so the two can be reported apart.
Runs finished on 29 September 2026: Gemini 3.5 and Qwen3-Omni, seed 43. Every question was also asked without the clip. Each cell is small, 40 questions per level and format, so read the pattern across cells, not single numbers.
1 · The answer format changes the score more than the level does
Select-all and short answer are far harder than single choice for both models. Yes/no is not reliably different from single choice once chance is taken into account.
Numbers by level
With the clip
Yes / no
Single choice
Select all
Short answer or true/false per option
Same, with explanation
Chance
50%
25%
7%
0–6%
0–6%
Gemini 3.5 · L1
62%
78%
32%
50%
55%
Gemini 3.5 · L2
70%
60%
38%
42%
42%
Gemini 3.5 · L3
68%
60%
28%
32%
38%
Gemini 3.5 · L4
62%
52%
22%
45%
42%
Qwen3-Omni · L1
68%
65%
20%
45%
45%
Qwen3-Omni · L2
62%
50%
15%
40%
45%
Qwen3-Omni · L3
80%
52%
20%
25%
28%
Qwen3-Omni · L4
42%
32%
5%
22%
8%
Mean over the four levels
Yes / no
Single choice
Select all
Short answer or true/false per option
Same, with explanation
Gemini 3.5, with the clip
66%
62%
30%
42%
44%
Gemini 3.5, without the clip
46%
32%
22%
20%
21%
Qwen3-Omni, with the clip
63%
50%
15%
33%
31%
Qwen3-Omni, without the clip
50%
36%
12%
20%
24%
Paired on the same 160 questions, select-all, short answer and short answer with explanation each score below single choice for both models (exact McNemar, p < 0.001 in all six comparisons with the clip).
Asking for a one-sentence explanation did not lower the score: 44% against 42% for Gemini 3.5, 31% against 33% for Qwen3-Omni.
The single-choice column reuses the published seed 43 answers to the same questions. Its wording was written by Gemini; the other columns used fixed templates in this run.
Short answers: rules scored 71% of Gemini's and 53% of Qwen3-Omni's; the judge scored the rest.
2 · True/false per option is treated like single choice
Both models mark exactly one option true most of the time, so this format does not behave as four independent judgements.
166 questions, with the clip
Per-option accuracy
All four right
Wrong answers with more than one “true”
Gemini 3.5
75%
48%
14%
Qwen3-Omni
67%
36%
8%
The last column is the over-true ratio of Liu et al. (2025), who report 63–86% on their benchmark. Here it is 8–14%: the options exclude each other, and the models act on that.
3 · The new question types
Splitting the activity question separated the two versions less than intended. The reworked “sound present” question is still hard, and still below chance without the clip for Gemini.
Numbers
64 questions each, single choice
Level
Gemini 3.5, clip / no clip
Qwen3-Omni, clip / no clip
Activity, by recognition easy_goal
L1
88% / 39%
80% / 34%
Activity, by inference medium_goal
L3
78% / 30%
77% / 30%
Sound present, new wrong options easy_sound
L1
50% / 14%
36% / 28%
Anticipate the put-down high_gaze_destination
L4
84% / 39%
69% / 44%
Anticipate the pick-up high_gaze_pickup (43 questions)
L4
42% / 16%
26% / 23%
The inference version of the activity question is 10 points harder than the recognition version for Gemini and 3 points for Qwen3-Omni. At 77–78% it is still among the easiest L3 questions.
“Sound present” without the clip scored 14% for Gemini, under the 25% of guessing. The wrong options were the most frequent sounds in each video, which look more plausible than the right answer. They are now matched to the right answer in frequency; that version has not been run yet.
Where the held object will be put down is far easier than expected for an L4 question, and is partly guessable: 39–44% without the clip, because “the kitchen counter” is right in 30% of cases.
Which object will be picked up next is hard, and at chance for Qwen3-Omni.
4 · Three “sound” question types do not need the sound
With the audio track replaced by silence, no score went down.
Numbers
Same questions, same clips
Model
Questions
With sound
Silent
Sound to object easy_sound_object
Gemini 3.5
64
66%
81%
Qwen3-Omni
64
53%
61%
Sound-timed handling low_sound_movement
Gemini 3.5
147
62%
75%
Qwen3-Omni
64
55%
56%
Purpose of the action low_purpose
Gemini 3.5
73
56%
62%
Qwen3-Omni
64
59%
67%
In these three types none of the wrong options appears in the clip, so the question reduces to which option is on screen. The fix is to draw wrong options from other moments of the same clip.
Silence cannot help a model, so the gains of 6–15 points for Gemini point to run-to-run variation between two days. The with-sound runs need repeating on the same day before the size of the effect can be stated.
5 · EPIC-Sounds: too few questions to conclude
Two 30-second clips gave 27 questions. On single choice, Gemini 3.5 scored 13 of 27 with the clip and 9 of 27 without; Qwen3-Omni scored 10 of 27 and 11 of 27. The pipeline runs end to end on a second dataset, but nothing about difficulty can be read from 27 questions.
What these runs do not yet show
One seed, temperature 0, one run per cell. Differences of a few points between cells are within noise.
The four template-worded formats have since been rewritten by Gemini (742 wordings, 0.7% rejected by the checker). Models have not answered that version yet.
The new level layout has only been run as pilot sets per question type, not as full 100-question levels.