HD-EPIC QUESTION PYRAMID ← RESULTS EXPLORER

How the questions are made. Every question is about one 30-second clip of first-person kitchen video with sound. A program decides what is asked and what the right answer is; Gemini 3.5 only writes the sentence.

From annotation to question

The same five steps produce every question at every level. Levels are independent: a higher-level question never uses the answer of a lower one.

  1. 1

    Official annotations

    HD-EPIC's human annotations: sounds, narrations, object movements, 3D positions, recipe steps, camera pose, eye-gaze priming.

    Dataset
  2. 2

    Mine a fact

    A rule finds something that is true in one clip and unique there. It outputs the right answer and three wrong options.

    Program
  3. 3

    Write the sentence

    Gemini 3.5 gets the fact and writes one question. It cannot change the answer or the options.

    Gemini 3.5
  4. 4

    Check

    The answer is computed again on the same clip. Wording that leaks the answer or drops the moment is rejected.

    Program
  5. 5

    Assemble

    100 questions per level. The right answer is spread evenly over A, B, C and D within each question type.

    Program

If Gemini's sentence fails the check, a fixed template is used instead, for at most 20% of a set.

The four levels

A level is defined by how much evidence a question needs, not by how often models get it right. Each level has non-spatial and spatial questions. In every sample the right answer is marked.

L1

Direct evidence

A single event answers it. Usually one modality; when sound and picture are both involved they point to the same answer.

Facts come from
  • Sound events
  • Object movements
  • 3D object positions
  • Camera pose
  • Recipe segments
The goal Gemini is given

“one piece of evidence, read off directly — either the sound and a visible object confirm each other, or a single placement answers it”

Rules that matter most here
  • For “sound present”: never name or hint at any sound in the question.
  • For spatial questions: keep the moment (“when the X is put down”) and never say where any object is.

Non-spatial

Sound presenteasy_sound

Which of the following sounds can be heard during this clip?

  • Ametal hitting wood
  • Bchopping
  • Cstirring or whisking
  • Dan electronic beep
Answer
A sound event lies fully inside the clip.
Wrong options
Sounds with no event within 30 s of the clip, similar in how common they are.
Fixed template
Sound to objecteasy_sound_object

You can hear sizzling or boiling in this clip. Which object visible on screen is producing it?

  • Abag of potatoes
  • Bsaucepan
  • Cpan
  • Dpotato
Answer
Exactly one visible object can make the sound that is heard.
Wrong options
Objects that can make the same sound but are not in the clip.
Worded by Gemini 3.5
Object movedeasy_move

Which object is moved from the hob to the kitchen counter during the clip?

  • Aglass lid
  • Bcan of olive oil
  • Ccloth
  • Dolive oil can cap
Answer
The only movement in the clip that goes from fixture A to fixture B.
Wrong options
Objects that took the same route elsewhere in the video.
Worded by Gemini 3.5
Activity, by recognitioneasy_goal

What activity is the camera wearer working on in this clip?

  • APrepare onion
  • BContinue preparing pizza dough
  • CSet the table
  • DPreheat oven
Answer
The recipe segment that contains the whole clip.
Wrong options
Other segments whose objects do not appear in the clip.
Fixed template

Spatial

Nearest objecteasy_spatial_near

At the moment the napkin is put down, which of these objects is closest to it?

  • Aspoon
  • Btumbler
  • Cstack of keep cups
  • Dblender
Answer
3D distance from the put-down point to every visible object.
Wrong options
Visible objects at least a set margin further away.
Worded by Gemini 3.5
Left / right of the wearereasy_spatial_side

At the moment the napkin is put down, which of these objects is furthest to the camera wearer's left?

  • Afrying pan
  • Bbowl
  • Cmeasuring cup
  • Dgreen chopping board
Answer
Object positions rotated into the camera frame at the put-down moment.
Wrong options
Other visible objects, clearly less far to that side.
Worded by Gemini 3.5
L2

Cross-modal alignment

A sound fixes one moment; a second signal gives the answer at that moment. Neither is enough alone.

Facts come from
  • Sound events
  • Object movements
  • Narrations
  • 3D object positions
  • Camera pose
The goal Gemini is given

“two complementary cues: the sound gives the time anchor, the picture gives the answer”

Rules that matter most here
  • Refer to the sound only by its given phrase.
  • For purpose questions: point to the moment by the sound only; never name the action or the object.
  • For left / right questions: the reference is the camera wearer, and the side must not be flipped.

Non-spatial

Sound-timed handlinglow_sound_movement

At the moment two plastic objects knocking together is heard, which object is the camera wearer handling?

  • Aroll of cling film
  • Bmeasuring jar
  • Cmeat box
  • Dchunk of minced beef
Answer
The one movement within 1.5 s of a sound that occurs once.
Wrong options
Objects not in the clip.
Worded by Gemini 3.5
Purpose of the actionlow_purpose

At the moment two plastic objects knocking together is heard, what is the camera wearer trying to do?

  • Ato wipe the counter surface
  • Bto pour out the water collected in it
  • Cto hold the sponge
  • Dto retrieve the bowl
Answer
The purpose clause of the narration at the sound-timed moment (“… to dry the hands”).
Wrong options
Purposes of similar actions elsewhere in the video.
Worded by Gemini 3.5

Spatial

Sound-timed nearest objectlow_spatial_near

At the moment metal hitting wood is heard, which object is closest to the saucepan?

  • Awater bottle
  • Byellow pepper
  • Cscissors
  • Dsalt shaker
Answer
Position of the handled object at the sound, interpolated along its movement; then 3D distance.
Wrong options
Two visible objects further away, one object not in the clip.
Worded by Gemini 3.5
Sound-timed left / rightlow_spatial_side

At the moment two metal objects knocking together is heard, which of the following is furthest to the camera wearer's right?

  • Abottle of cold water
  • Bice tray
  • Cspoon
  • Dblender
Answer
Camera pose at the sound; object positions rotated into the camera frame.
Wrong options
Other visible objects, clearly less far to that side.
Worded by Gemini 3.5
L3

Inference

The whole 30-second clip has to be combined: how often, in what order, where things end up, what it all adds up to.

Facts come from
  • Sound events
  • Narrations
  • 3D object positions
  • Recipe segments
The goal Gemini is given

“inference over the whole clip — what the overall activity is, how often something happens, the order of events, or where things end up”

Rules that matter most here
  • For order questions: list the three items in the given shuffled order.
  • For activity questions: never name, paraphrase or hint at any option.
  • For spatial questions: ask about where things are “by the end of the clip”.

Non-spatial

Counting a soundmedium_count

How many separate times is metal hitting ceramic heard in this clip?

  • A3
  • B2
  • C1
  • D4
Answer
Number of events of one sound, all fully inside the clip, 2 to 6.
Wrong options
Four consecutive counts; the true one in a random slot.
Worded by Gemini 3.5
Order of actionsmedium_order

In what order do these events occur in the clip?
(place the hand towel, stir the saucepan, pick up the hand rag)

  • Aplace the hand towel -> stir the saucepan -> pick up the hand rag
  • Bstir the saucepan -> place the hand towel -> pick up the hand rag
  • Cstir the saucepan -> pick up the hand rag -> place the hand towel
  • Dpick up the hand rag -> place the hand towel -> stir the saucepan
Answer
Start times of three distinct narrated actions.
Wrong options
Other orders of the same three actions.
Worded by Gemini 3.5
Activity, by inferencemedium_goal

Taking the whole clip into account, what is the camera wearer mainly working on?

  • ADry hands
  • BStir saucepan and add water to it
  • CStir onions and transfer to bowl
  • DRetrieve garlic and ginger, add to saucepan and stir
Answer
The recipe segment that contains the whole clip.
Wrong options
Other segments that share objects or actions with the right one.
Fixed template

Spatial

Furthest / nearest at the endmedium_spatial_extreme

By the end of the clip, which object is closest to the large shopping bag?

  • Abag of mozzarella
  • Bbag of salad
  • Cbox of food
  • Dsingle cream pot
Answer
3D distances at the last instant of the clip.
Wrong options
The other visible objects.
Worded by Gemini 3.5
Order by distancemedium_spatial_order

By the end of the clip, what is the order of these objects from closest to furthest to the napkin?
(kettle, blender, measuring cup)

  • Ameasuring cup -> blender -> kettle
  • Bkettle -> blender -> measuring cup
  • Cmeasuring cup -> kettle -> blender
  • Dblender -> measuring cup -> kettle
Answer
Three objects with clearly separated distances at the end of the clip.
Wrong options
Other orders of the same three objects.
Worded by Gemini 3.5
L4

Planning

Decide what happens next. The answer is never shown: the clip stops before it, or the question joins two scenes.

Facts come from
  • Recipe segments
  • Camera pose
  • 3D object positions
  • Object movements
  • Eye-gaze priming
The goal Gemini is given

“planning — what the camera wearer will do next, where to go to fetch something seen earlier, or which way to head”

Rules that matter most here
  • For two-scene questions: open with the given sentence that says the video joins two scenes.
  • For route questions: the options are fixed headings; never describe where anything is.
  • For anticipation questions: never name a location or an object, and never say where the wearer is looking.

Non-spatial

Next stephigh_plan_next

What is the camera wearer planning to do next after completing the actions shown in the clip?

  • ATidy up
  • BAdd second slice of bread
  • CWashing the utensils
  • DBring tumbler and pour mango lassi into it
Answer
The recipe segment that follows; the clip is the last 30 s of the current one.
Wrong options
The current step as a trap, and steps from other videos.
Worded by Gemini 3.5

Spatial

Route across two sceneshigh_plan_route

This video shows two scenes from the same kitchen, recorded about a minute apart. At the end of the second scene, which way should the camera wearer head to retrieve the mango bits seen in the first scene?

  • Aturn around
  • Bturn left
  • Cturn right
  • Dstraight ahead
Answer
Camera pose at the end of scene 2 and the 3D point where the object was left in scene 1.
Wrong options
The other three headings.
Worded by Gemini 3.5

Gaze anticipation

Anticipate the put-downhigh_gaze_destination

At the end of this clip the camera wearer is holding the glass lid. Where are they about to put it down?

  • Athe oven
  • Bthe floor
  • Cthe hob
  • Dthe kitchen counter
Answer
Destination fixture of a put-down the wearer looked at 1.5 s or more in advance. The clip ends 1 s before it.
Wrong options
Where the object came from, as a trap, and other fixtures.
Fixed template
Anticipate the pick-uphigh_gaze_pickup

Which object is the camera wearer about to pick up right after this clip ends?

  • Aslice of pepper
  • Bspoon
  • Clid of pot
  • Dphone
Answer
The next object picked up, looked at 1.5 s or more in advance. The clip ends 1 s before it.
Wrong options
Visible objects that are not picked up in the next 10 s.
Fixed template

What Gemini is asked to do

One request covers ten facts of the same level. The request has five parts; only the first differs between levels.

1

Task and level goal

“For each fact packet below, write ONE four-choice question.” Followed by the goal sentence of the level, shown in each level above.

Differs by level
2

Hard rules

About 35 lines, the same for every level. The answer and options are fixed. The question must not contain the answer. Sounds are named only by the given phrase. No words such as “annotation” or “label”.

Same for all
3

Style examples

Three real questions of the same level, plus matching examples for spatial question types.

Same for all
4

How the model did one level below

From level 2 up: the score, accuracy by question type, and up to six missed questions. Gemini is told to word questions in the styles the model found hard, without changing any fact.

Level 2 and up
5

The facts

Ten packets. Each has the right answer, the three wrong options, and the fields its type needs.

Same for all
WHAT GEMINI RECEIVES FOR ONE FACT
{
  "subtype": "low_sound_movement",
  "clip_length_sec": 30,
  "correct_answer": "roll of cling film",
  "distractors": ["measuring jar", "meat box",
                  "chunk of minced beef"],
  "sound_natural_phrase":
      "two plastic objects knocking together"
}
WHAT GEMINI RETURNS
{
  "question": "At the moment two plastic objects
      knocking together is heard, which object
      is the camera wearer handling?",
  "answer_echo": "roll of cling film"
}

The answer must be copied back unchanged. If it does not match the program's answer, the question is dropped. Settings: temperature 0, fixed seed, thinking off.

The fixed templates

Used when Gemini's sentence is rejected, and for every question in the newer formats. Words in braces are filled from the fact.

LevelTypeTemplate
L1Sound to objectYou can hear {sound} in this clip. Which object visible on screen is producing it?
L1Nearest objectThe camera wearer puts down {reference} during this clip. Which of these objects does it end up closest to?
L2Sound-timed handlingAt the moment you hear {sound}, which object is the camera wearer handling?
L2Sound-timed left / rightAt the moment you hear {sound}, which of these objects is furthest to the camera wearer's {side}?
L3Counting a soundHow many separate times is {sound} heard in this clip?
L3Furthest at the endBy the end of this clip, which of these objects ends up furthest away from {reference}?
L4Next stepJudging from what the camera wearer is doing in this clip, what are they most likely to do next?
L4Anticipate the put-downAt the end of this clip the camera wearer is holding {object}. Where are they about to put it down?

Is the Gemini that writes the question independent of the Gemini that answers it?

The calls are independent. The model is not: the same model writes, answers and judges.

What is independent

Gemini writing the question

  • Sees: the right answer, the three wrong options, and the fields of the fact
  • Does not see: the clip
  • Every request starts from nothing

Gemini answering the question

  • Sees: the question, the four options, and the 30-second clip
  • Does not see: which option is right, or anything from the writing step
  • Every question is a separate request; nothing carries over between questions

So the answer cannot reach the answering model through a shared conversation or cache.

What is not independent

WRITES THE SENTENCEGemini 3.5
PICKS THE FACTS (OPTIONAL)Gemini 3.5
ANSWERS, AS ONE OF TWO MODELS TESTEDGemini 3.5
JUDGES SHORT ANSWERSGemini 3.5

Qwen3-Omni, the other model tested, takes no part in writing or judging.

Possible biasHow it could ariseWhat limits it
Wording suits GeminiThe sentences are in Gemini's own style, which it may read more easily than Qwen3-Omni does.The answer and options are fixed by the program. Four of the five formats use fixed templates, not Gemini's wording.
The writer sees earlier scoresFrom level 2 up, the writing request includes the tested models' score and missed questions from the level below.The scores of both models are pooled, not Gemini's alone. Only the wording can change.
The judge is also a test-takerGemini judges its own short answers and Qwen3-Omni's, and may be more lenient to phrasing like its own.The judge sees no clip and is not told which candidate is right. It is used only when rules cannot decide.

What the existing data can and cannot show

Question wordingQuestionsGemini 3.5Qwen3-Omni
Written by Gemini1,58955%45%
Fixed template1127%36%

Seeds 43–46, with video. This cannot settle the question: only 11 questions use a template, and they are the ones whose Gemini wording was rejected, so they are not a fair sample.

Checks planned

  1. 1

    Compare across formats

    Single choice is worded by Gemini; the other four formats are templates. If Gemini's lead over Qwen3-Omni shrinks on the template formats, wording is part of the lead.

    No extra calls
  2. 2

    Check the judge by hand

    Read 50 of the judge's verdicts, split between the two models' answers, and count disagreements.

    No extra calls
  3. 3

    Template-only rerun

    Ask the same questions with template wording to both models. Run only if checks 1 and 2 show a difference.

    About 800 calls per video

One question, five formats

The clip, the right answer and the wrong options stay the same. Only the way of answering changes, so a difference in score comes from the format alone.

One verified factright answer + three wrong options
1

Yes / no

One statement, built from the right answer or from a wrong option, half and half.

chance 50%
2

Single choice

The original four options.

chance 25%
3

Select all that apply

Four statements, one to three of them true. Extra true ones come from other facts about the same clip.

chance 7%
?

Does it still have one answer without the options?

Yes4

Short answer

No options. The question states the output format.

5

Short answer with explanation

Adds one sentence of support. Only the answer is scored.

chance about 0%
No4b

True / false per option

The options stay. Each is judged on its own, and all four must be right.

chance 6%

Which question types can drop their options

Short answer

  • One object: sound to object, object moved, sound-timed handling, anticipate the pick-up
  • One number: counting a sound
  • One order: order of actions, order by distance (the three items are listed in the question)
  • One phrase: activity by recognition, purpose of the action
  • One direction: route across two scenes

True / false per option

  • Sound present: many sounds are heard, so there is no single answer
  • All nearest, furthest and left / right questions: the candidates are the tracked objects, which the model cannot know
  • Activity by inference, next step: free text cannot separate stages of one recipe
  • Anticipate the put-down: several fixtures share a name

The same question in each format

1

Yes / no

Statement: At the moment you hear two plastic objects knocking together, the camera wearer is handling the roll of cling film.

yes

2

Single choice

At the moment two plastic objects knocking together is heard, which object is the camera wearer handling?
A. roll of cling film
B. measuring jar
C. meat box
D. chunk of minced beef

A

3

Select all that apply

A. … is handling the chunk of minced beef.
B. … is handling the measuring jar.
C. … is handling the roll of cling film.
D. Rustling can be heard during this clip.

C, D

4

Short answer

At the moment you hear two plastic objects knocking together, which object is the camera wearer handling? Answer with the name of the object.

roll of cling film

5

Short answer with explanation

(same question)
Answer: <your answer>
Because: <one sentence on what you saw or heard>

roll of cling film

4b

True / false per option

At the moment metal hitting wood is heard, which object is closest to the saucepan?
A. water bottle     → true / false
B. yellow pepper    → true / false
C. scissors         → true / false
D. salt shaker      → true / false

true, false, false, false

How a short answer is scored

  1. 1

    Rules first

    Number: the first integer. Direction: left, right, straight, around. Order: where each item appears in the reply. Object: its name or head noun, and no wrong option's name.

    Program
  2. 2

    Judge, only if rules cannot decide

    Gemini 3.5 sees the question, the reply and the four candidates in shuffled order, and says which one the reply means, or none. It does not see the clip or know which is right.

    Gemini 3.5
  3. 3

    Record

    Each answer is marked as scored by rule or by judge, so the two can be reported apart.

    Program

The scheme follows Beyond Multiple Choice: Verifiable OpenQA for Robust Vision-Language RFT (Liu et al., 2025). More detail, and what changed from the published sets, is on the update page.

Results from the first runs

Runs finished on 29 September 2026: Gemini 3.5 and Qwen3-Omni, seed 43. Every question was also asked without the clip. Each cell is small, 40 questions per level and format, so read the pattern across cells, not single numbers.

1 · The answer format changes the score more than the level does

Select-all and short answer are far harder than single choice for both models. Yes/no is not reliably different from single choice once chance is taken into account.

Numbers by level
With the clipYes / noSingle choiceSelect allShort answer or true/false per optionSame, with explanation
Chance50%25%7%0–6%0–6%
Gemini 3.5 · L162%78%32%50%55%
Gemini 3.5 · L270%60%38%42%42%
Gemini 3.5 · L368%60%28%32%38%
Gemini 3.5 · L462%52%22%45%42%
Qwen3-Omni · L168%65%20%45%45%
Qwen3-Omni · L262%50%15%40%45%
Qwen3-Omni · L380%52%20%25%28%
Qwen3-Omni · L442%32%5%22%8%
Mean over the four levelsYes / noSingle choiceSelect allShort answer or true/false per optionSame, with explanation
Gemini 3.5, with the clip66%62%30%42%44%
Gemini 3.5, without the clip46%32%22%20%21%
Qwen3-Omni, with the clip63%50%15%33%31%
Qwen3-Omni, without the clip50%36%12%20%24%

2 · True/false per option is treated like single choice

Both models mark exactly one option true most of the time, so this format does not behave as four independent judgements.

166 questions, with the clipPer-option accuracyAll four rightWrong answers with more than one “true”
Gemini 3.575%48%14%
Qwen3-Omni67%36%8%

The last column is the over-true ratio of Liu et al. (2025), who report 63–86% on their benchmark. Here it is 8–14%: the options exclude each other, and the models act on that.

3 · The new question types

Splitting the activity question separated the two versions less than intended. The reworked “sound present” question is still hard, and still below chance without the clip for Gemini.

Numbers
64 questions each, single choiceLevelGemini 3.5, clip / no clipQwen3-Omni, clip / no clip
Activity, by recognition easy_goalL188% / 39%80% / 34%
Activity, by inference medium_goalL378% / 30%77% / 30%
Sound present, new wrong options easy_soundL150% / 14%36% / 28%
Anticipate the put-down high_gaze_destinationL484% / 39%69% / 44%
Anticipate the pick-up high_gaze_pickup (43 questions)L442% / 16%26% / 23%

4 · Three “sound” question types do not need the sound

With the audio track replaced by silence, no score went down.

Numbers
Same questions, same clipsModelQuestionsWith soundSilent
Sound to object easy_sound_objectGemini 3.56466%81%
Qwen3-Omni6453%61%
Sound-timed handling low_sound_movementGemini 3.514762%75%
Qwen3-Omni6455%56%
Purpose of the action low_purposeGemini 3.57356%62%
Qwen3-Omni6459%67%

5 · EPIC-Sounds: too few questions to conclude

Two 30-second clips gave 27 questions. On single choice, Gemini 3.5 scored 13 of 27 with the clip and 9 of 27 without; Qwen3-Omni scored 10 of 27 and 11 of 27. The pipeline runs end to end on a second dataset, but nothing about difficulty can be read from 27 questions.

What these runs do not yet show