R4DSG: Relative 4D Scene Graph Memory for Object-Centric Question Answering in Long Egocentric Video
- Problem: Object-centric question answering in long egocentric video requires preserving persistent object identity and structured spatial change, which caption-based methods fail to capture.
- Model: R4DSG: Relative 4D scene graph memory that indexes video by time, place, persistent objects, and anchor-relative transitions without requiring global world coordinates.
- Code: https://dualtransparency.github.io/R4DSG/
