EDBT 2026 Demo / reviewers in the wild / expert
Jacob Chalk
dblp:339/6566
· DBLP profile ↗
5ranked-venue papers
1as first author
5since 2021 · last 2025
0000-0002-9751-8660ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Video understanding and tracking · 93% Vision and language · 7% | |
| Computer graphics and multimedia
2 papers |
Audio and music processing · 50% Multimedia analysis and retrieval · 50% |
Topics — the 6 heaviest of 7, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Video understanding and tracking
egocentric video understanding |
1.7 | 2 | 2025 | EPIC-SOUNDS: A Large-Scale Dataset of Actions That Sound · IEEE Trans. Pattern Anal. Mach. Intell. 2025 HD-EPIC: A Highly-Detailed Egocentric Video Dataset · CVPR 2025 |
Multimedia analysis and retrieval
audio-visual learning |
0.9 | 1 | 2025 | EPIC-SOUNDS: A Large-Scale Dataset of Actions That Sound · IEEE Trans. Pattern Anal. Mach. Intell. 2025 |
Audio and music processing › sound event detection
sound event recognition |
0.9 | 1 | 2025 | EPIC-SOUNDS: A Large-Scale Dataset of Actions That Sound · IEEE Trans. Pattern Anal. Mach. Intell. 2025 |
Computer vision › Video understanding and tracking
action recognition |
0.8 | 1 | 2024 | TIM: A Time Interval Machine for Audio-Visual Action Recognition · CVPR 2024 |
Computer vision › Video understanding and tracking › action recognition › multimodal action recognition
audio-visual action recognition |
0.8 | 1 | 2024 | TIM: A Time Interval Machine for Audio-Visual Action Recognition · CVPR 2024 |
Computer vision › Vision and language
visual question answering |
0.3 | 1 | 2025 | HD-EPIC: A Highly-Detailed Egocentric Video Dataset · CVPR 2025 |
Methods — techniques the papers use, named apart from their topics
audio-visual models · 1.7audio recognition models · 1.7transformer encoder · 1.5time interval queries · 1.5multimodal fusion · 1.5gaze priming · 0.9digital twinning · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Spatial Cognition from Egocentric Video: Out of Sight, Not Out of MindabstractAs humans move around, performing their daily tasks, they are able to recall where they have positioned objects in their environment, even if these objects are currently out of their sight. In this paper, we aim to mimic this spatial cognition ability. We thus formulate the task of Out of Sight, Not Out of Mind - 3D tracking active objects using observations captured through an egocentric camera. We introduce a simple but effective approach to address this challenging problem, called Lift, Match, and Keep (LMK). LMK lifts partial 2D observations to 3D world coordinates, matches them over time using visual appearance, 3D location and interactions to form object tracks, and keeps these object tracks even when they go out-of-view of the camera. We benchmark LMK on 100 long videos from EPICKITCHENS. Our results demonstrate that spatial cognition is critical for correctly locating objects over short and long time scales. E.g., for one long egocentric video, we estimate the 3D location of 50 active objects. After 120 seconds, 57 % of the objects are correctly localised by LMK, compared to just 33% by a recent 3D method for egocentric videos and 17 % by a general 2D tracking method. Chiara Plizzari, Shubham Goel 0001, Toby Perrett, Jacob Chalk, Angjoo Kanazawa, Dima Damen |
3DV | 4 |
| 2025 | HD-EPIC: A Highly-Detailed Egocentric Video DatasetabstractWe present a validation dataset of newly-collected kitchen-based egocentric videos, manually annotated with highly detailed and interconnected ground-truth labels covering: recipe steps, fine-grained actions, ingredients with nutritional values, moving objects, and audio annotations. Importantly, all annotations are grounded in 3D through digital twinning of the scene, fixtures, object locations, and primed with gaze. Footage is collected from unscripted recordings in diverse home environments, making HD-EPIC the first dataset collected in-the-wild but with detailed annotations matching those in controlled lab environments.We show the potential of our highly-detailed annotations through a challenging VQA benchmark of 26K questions assessing the capability to recognise recipes, ingredients, nutrition, fine-grained actions, 3D perception, object motion, and gaze direction. The powerful long-context Gemini Pro only achieves 37.6% on this benchmark, showcasing its difficulty and highlighting shortcomings in current VLMs. We additionally assess action recognition, sound recognition, and long-term video-object segmentation on HD-EPIC.HD-EPIC is 41 hours of video in 9 kitchens with digital twins of 413 kitchen fixtures, capturing 69 recipes, 59K fine-grained actions, 51K audio events, 20K object movements and 37K object masks lifted to 3D. On average, we have 263 annotations per minute of our unscripted videos. Toby Perrett, Ahmad Darkhalil, Saptarshi Sinha, Omar Emara, Sam Pollard, Kranti Kumar Parida, Kaiting Liu, Prajwal Gatti, Siddhant Bansal, Kevin Flanagan, Jacob Chalk, Zhifan Zhu 0001, Rhodri Guerrier, Fahd Abdelazim, Bin Zhu 0006, Davide Moltisanti, Michael Wray, Hazel Doughty, Dima Damen |
CVPR | 11 |
| 2025 | EPIC-SOUNDS: A Large-Scale Dataset of Actions That SoundabstractWe introduce EPIC-SOUNDS, a large-scale dataset of audio annotations capturing temporal extents and class labels within the audio stream of the egocentric videos. We propose an annotation pipeline where annotators temporally label distinguishable audio segments and describe the action that could have caused this sound. We identify actions that can be discriminated purely from audio, through grouping these free-form descriptions of audio into classes. For actions that involve objects colliding, we collect human annotations of the materials of these objects (e.g., a glass object being placed on a wooden surface), which we verify from video, discarding ambiguities. Overall, EPIC-SOUNDS includes 78.4 k categorised segments of audible events and actions, distributed across 44 classes as well as 39.2 k non-categorised segments. We train and evaluate state-of-the-art audio recognition and detection models on our dataset, for both audio-only and audio-visual methods. We also conduct analysis on: the temporal overlap between audio events, the temporal and label correlations between audio and visual modalities, the ambiguities in annotating materials from audio-only input, the importance of audio-only labels and the limitations of current models to understand actions that sound. Jaesung Huh, Jacob Chalk, Evangelos Kazakos, Dima Damen, Andrew Zisserman |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2024 | TIM: A Time Interval Machine for Audio-Visual Action RecognitionabstractDiverse actions give rise to rich audio-visual signals in long videos. Recent works showcase that the two modalities of audio and video exhibit different temporal extents of events and distinct labels. We address the interplay between the two modalities in long videos by explicitly modelling the temporal extents of audio and visual events. We propose the Time Interval Machine (TIM) where a modality-specific time interval poses as a query to a transformer encoder that ingests a long video input. The encoder then attends to the specified interval, as well as the surrounding context in both modalities, in order to recognise the ongoing action. We test TIM on three long audio-visual video datasets: EPIC-KITCHENS, Perception Test, and AVE, reporting state-of-the-art (SOTA) for recognition. On EPIC-KITCHENS, we beat previous SOTA that utilises LLMs and significantly larger pre-training by 2.9% top-1 action recognition accuracy. Additionally, we show that TIM can be adapted for action detection, using dense multi-scale inter-val queries, outperforming SOTA on EPIC-KITCHENS-IOO for most metrics, and showing strong performance on the Perception Test. Our ablations show the critical role of in-tegrating the two modalities and modelling their time inter-vals in achieving this performance. Code and models at: https://github.com/JacobChalk/TIM. Jacob Chalk, Jaesung Huh, Evangelos Kazakos, Andrew Zisserman, Dima Damen |
CVPR | 1 |
| 2023 | Epic-Sounds: A Large-Scale Dataset of Actions that SoundabstractWe introduce EPIC-SOUNDS, a large-scale dataset of audio annotations capturing temporal extents and class labels within the audio stream of the egocentric videos from EPIC-KITCHENS-100. We propose an annotation pipeline where annotators temporally label distinguishable audio segments and describe the action that could have caused this sound. We identify actions that can be discriminated purely from audio, through grouping free-form descriptions into classes. For actions that involve objects colliding, we collect human annotations of the materials of these objects (e.g. a glass object being placed on a wooden surface), which we verify from visual labels, discarding ambiguities. Overall, EPIC-SOUNDS includes 78.4k categorised segments of audible events and actions, distributed across 44 classes, as well as 39.2k non-categorised segments, totalling 117.6k segments spanning 100 hours of audio, capturing diverse actions that sound in home kitchens. We train and evaluate two state-of-the-art audio recognition models on our dataset, highlighting the importance of audio-only labels and the limitations of current models to recognise actions that sound.EPIC-SOUNDS and baseline source code is available from: https://epic-kitchens.github.io/epic-sounds. Jaesung Huh, Jacob Chalk, Evangelos Kazakos, Dima Damen, Andrew Zisserman |
ICASSP | 2 |