← Back
rhfeiyang

rhfeiyang/GEB

View on GitHub ↗
Stars
45
Forks
0
Watchers
45
Open issues
2
Contributors
1
Language
—
License
—
Default branch
main
Created Sep 28, 2026Updated Sep 30, 2026

Star growth

Today—
This week—
This month—

Star history will appear here once this repo has been tracked for a couple of days.

README

Beyond the Timeline:
Augmenting Long-Video Memory with Grounded Entity Biographies

Project page arXiv Hugging Face Paper Artifacts

Hui Ren1, Lei Fan2, Henry Pao2, Han Guo2, Zeeshan Zia2, Ying Chen2, Alexander G. Schwing1, Gang Hua2
1University of Illinois Urbana-Champaign    2Amazon.com, Inc.

Two red mugs, two biographies. A memory of moments cannot tell which mug went into the dishwasher; a memory of entities can.

Overview

Answering questions about long videos often requires connecting events involving the same objects across hours or days. Chronological descriptions and text-derived entities can leave physical identity unresolved: different objects may share a description, while observations of the same object remain disconnected across events. Retrieving relevant events therefore does not necessarily recover the biography of the particular entity a question concerns.

History is written in two ways, and long-video memory needs both:

  • Chronicle: follows events through time and recalls what happened at a moment. Two accurate descriptions of "a red mug" still cannot tell whether they are the same mug.
  • Biography: follows one subject through those events and recalls what happened to this mug. The coffee mug never reaches the dishwasher.

Animation: five moments in time order (the chronicle) are regrouped into two biographies, one per red mug. The striped mug is filled with coffee and returns to the counter; the solid red mug is picked up and goes into the dishwasher.

Grounded Entity Biographies (GEB) is a long-video memory framework that groups visually grounded observations of the same physical instance across clips into retrievable biographies while preserving the context of each moment. During question answering, the biography is retrieved alongside episodic evidence, allowing the model to follow an entity through events using identity links established during memory construction.

Overview of GEB: a grounded observation of a blue hand mixer on Day 1 is linked to observations on Days 3 to 6; a retrieval controller reads the question, issues searches, and passes a biography excerpt and linked episode context to the answer model.

The memory is written in two steps and read in a third:

  1. Ground. Each tracked subject in a clip becomes an observation, described from its own crops, the scene frames and the dialogue of that moment.
  2. Associate. An observation joins an existing biography only if it matches the entity's recent references and is never seen apart from them in a shared frame; otherwise it starts a new one.
  3. Read. Retrieval enters through a matched moment, follows same-instance edges to the rest of the biography, and reaches the episodes around each encounter. The biography excerpt also lists the appearances not yet inspected, giving the controller concrete targets for further search.

Results

GEB is evaluated on four benchmarks over week-long and day-long recordings, with multiple-choice and open-ended questions: EgoLifeQA, Ego-R1-Bench, MM-Lifelong (Test@Week and Test@Day) and MultiHop-EgoQA. On EgoLifeQA it reaches 72.0% accuracy, 4.4 points above the best published result, with the same controller, answer model and retrieval limits as the strongest baseline.

Main results table: accuracy of sixteen systems on EgoLifeQA, Ego-R1-Bench and MM-Lifelong Test@Week and Test@Day. GEB has the best overall score in every column.

Full tables, ablations and the evidence-access analysis are on the project page and in the paper.

Release

The code and the artifacts are coming soon. Stay tuned!

Citation

@misc{ren2026GEB,
      title={Beyond the Timeline: Augmenting Long-Video Memory with Grounded Entity Biographies}, 
      author={Hui Ren and Lei Fan and Henry Pao and Han Guo and Zeeshan Zia and Ying Chen and Alexander Schwing and Gang Hua},
      year={2026},
      eprint={2609.38155},
      archivePrefix={arXiv},
      primaryClass={cs.CV},
      url={https://arxiv.org/abs/2609.38155}, 
}