Haozhe Qi

dblp:266/3261 · DBLP profile ↗
← Back
5ranked-venue papers
2as first author
4since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
3D vision · 52% Video understanding and tracking · 24% Vision and language · 10%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Environmental and earth informatics · 100%
Computer graphics and multimedia
1 paper
Multimedia analysis and retrieval · 100%

Topics — the 17 heaviest of 19, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Video understanding and tracking
action segmentation
0.912025
EPFL-Smart-Kitchen: An Ego-Exo Multi-Modal Dataset for Challenging Action and Motion Understanding in Video-Language Models · NeurIPS 2025
Computer vision › Video understanding and tracking › activity recognition
animal behavior recognition
0.912025
MammAlps: A Multi-view Video Behavior Monitoring Dataset of Wild Mammals in the Swiss Alps · CVPR 2025
Machine learning › Generative modeling › diffusion model › human motion generation
text-to-motion generation
0.912025
EPFL-Smart-Kitchen: An Ego-Exo Multi-Modal Dataset for Challenging Action and Motion Understanding in Video-Language Models · NeurIPS 2025
Computer vision › Vision and language
video-language model
0.912025
EPFL-Smart-Kitchen: An Ego-Exo Multi-Modal Dataset for Challenging Action and Motion Understanding in Video-Language Models · NeurIPS 2025
Environmental and earth informatics
ecological monitoring
0.912025
MammAlps: A Multi-view Video Behavior Monitoring Dataset of Wild Mammals in the Swiss Alps · CVPR 2025
Multimedia analysis and retrieval
multimodal action understanding
0.912025
EPFL-Smart-Kitchen: An Ego-Exo Multi-Modal Dataset for Challenging Action and Motion Understanding in Video-Language Models · NeurIPS 2025
Computer vision › 3D vision
3d shape reconstruction
0.812024
HOISDF: Constraining 3D Hand-Object Pose Estimation with Global Signed Distance Fields · CVPR 2024
Computer vision › 3D vision › object pose estimation
hand-object pose estimation
0.812024
HOISDF: Constraining 3D Hand-Object Pose Estimation with Global Signed Distance Fields · CVPR 2024
Computer vision › 3D vision › 3d shape representation
implicit function
0.812024
HOISDF: Constraining 3D Hand-Object Pose Estimation with Global Signed Distance Fields · CVPR 2024
Computer vision › 3D vision › 3d reconstruction
signed distance field representation
0.812024
HOISDF: Constraining 3D Hand-Object Pose Estimation with Global Signed Distance Fields · CVPR 2024
Computer vision › 3D vision
3d object detection
0.412020
P2B: Point-to-Box Network for 3D Object Tracking in Point Clouds · CVPR 2020
Computer vision › Video understanding and tracking › object tracking
3d object tracking
0.412020
P2B: Point-to-Box Network for 3D Object Tracking in Point Clouds · CVPR 2020
Computer vision › 3D vision › 3d object detection
point cloud object detection
0.412020
P2B: Point-to-Box Network for 3D Object Tracking in Point Clouds · CVPR 2020
Computer vision › 3D vision › point cloud processing › point cloud video understanding
point cloud tracking
0.412020
P2B: Point-to-Box Network for 3D Object Tracking in Point Clouds · CVPR 2020
Computer vision › 3D vision › 3d object detection › proposal-based 3d object detection
vote-based 3d detection
0.412020
P2B: Point-to-Box Network for 3D Object Tracking in Point Clouds · CVPR 2020
Computer vision › Segmentation and scene understanding
video segmentation
0.312025
MammAlps: A Multi-view Video Behavior Monitoring Dataset of Wild Mammals in the Swiss Alps · CVPR 2025
Computer vision › Face, body and person analysis › human pose estimation › articulated pose estimation
hand pose estimation
0.212024
HOISDF: Constraining 3D Hand-Object Pose Estimation with Global Signed Distance Fields · CVPR 2024

Methods — techniques the papers use, named apart from their topics

multimodal learning · 4.4vision-language model · 2.6audio-visual learning · 1.7signed distance field · 0.8graph neural network · 0.8pointnet++ · 0.4permutation-invariant feature augmentation · 0.4hough voting · 0.4
YearPublicationVenuePosition
2026 Proto-Former: Unified Facial Landmark Detection by Prototype Transformer
abstract
Recent advances in deep learning have significantly improved facial landmark detection. However, existing facial landmark detection datasets often define different numbers of landmarks, and most mainstream methods can only be trained on a single dataset. This limits the model generalization to different datasets and hinders the development of a unified model. To address this issue, we propose Proto-Former, a unified, adaptive, end-to-end facial landmark detection framework that explicitly enhances dataset-specific facial structural representations (i.e., prototype). Proto-Former overcomes the limitations of single-dataset training by enabling joint training across multiple datasets within a unified architecture. Specifically, Proto-Former comprises two key components: an Adaptive Prototype-Aware Encoder (APAE) that performs adaptive feature extraction and learns prototype representations, and a Progressive Prototype-Aware Decoder (PPAD) that refines these prototypes to generate prompts that guide the model's attention to key facial regions. Furthermore, we introduce a novel Prototype-Aware (PA) loss, which achieves optimal path finding by constraining the selection weights of prototype experts. This loss function effectively resolves the problem of prototype expert addressing instability during multi-dataset training, alleviates gradient conflicts, and enables the extraction of more accurate facial structure features. Extensive experiments on widely used benchmark datasets demonstrate that our Proto-Former achieves superior performance compared to existing state-of-the-art methods. The code is publicly available at:https://github.com/Husk021118/Proto-Former.
Shengkai Hu, Haozhe Qi, Jun Wan 0005, Jiaxing Huang 0001, Lefei Zhang, Dacheng Tao
IEEE Trans. Multim.2
2025 MammAlps: A Multi-view Video Behavior Monitoring Dataset of Wild Mammals in the Swiss Alps
abstract
Monitoring wildlife is essential for ecology and ethology, especially in light of the increasing human impact on ecosystems. Camera traps have emerged as habitat-centric sensors enabling the study of wildlife populations at scale with minimal disturbance. However, the lack of annotated video datasets limits the development of powerful video understanding models needed to process the vast amount of fieldwork data collected. To advance research in wild animal behavior monitoring we present MammAlps, a multi-modal and multi-view dataset of wildlife behavior monitoring from 9 camera-traps in the Swiss National Park. Mam-mAlps contains over 14 hours of video with audio, 2D segmentation maps and 8.5 hours of individual tracks densely labeled for species and behavior. Based on 6‘135 single animal clips, we propose the first hierarchical and multi-modal animal behavior recognition benchmark using audio, video and reference scene segmentation maps as inputs. Furthermore, we also propose a second ecology-oriented benchmark aiming at identifying activities, species, number of individuals and meteorological conditions from 397 multi-view and long-term ecological events, including false positive triggers. We advocate that both tasks are complementary and contribute to bridging the gap between machine learning and ecology. Code and data are available at https://github.com/eceo-epfl/MammAlps.
Valentin Gabeff, Haozhe Qi, Brendan Flaherty, Gencer Sumbul, Alexander Mathis, Devis Tuia
CVPR2
2025 EPFL-Smart-Kitchen: An Ego-Exo Multi-Modal Dataset for Challenging Action and Motion Understanding in Video-Language Models
abstract
Understanding behavior requires datasets that capture humans while carrying out complex tasks. The kitchen is an excellent environment for assessing human motor and cognitive function, as many complex actions are naturally exhibited in kitchens from chopping to cleaning. Here, we introduce the EPFL-Smart-Kitchen-30 dataset, collected in a noninvasive motion capture platform inside a kitchen environment. Nine static RGB-D cameras, inertial measurement units (IMUs) and one head-mounted HoloLens~2 headset were used to capture 3D hand, body, and eye movements. The EPFL-Smart-Kitchen-30 dataset is a multi-view action dataset with synchronized exocentric, egocentric, depth, IMUs, eye gaze, body and hand kinematics spanning 29.7 hours of 16 subjects cooking four different recipes. Action sequences were densely annotated with 33.78 action segments per minute. Leveraging this multi-modal dataset, we propose four benchmarks to advance behavior understanding and modeling through 1) a vision-language benchmark, 2) a semantic text-to-motion generation benchmark, 3) a multi-modal action recognition benchmark, 4) a pose-based action segmentation benchmark. We expect the EPFL-Smart-Kitchen-30 dataset to pave the way for better methods as well as insights to understand the nature of ecologically-valid human behavior. Code and data are available at https://amathislab.github.io/EPFL-Smart-Kitchen
Andy Bonnetto, Haozhe Qi, Franklin Leong, Matea Tashkovska, Mahdi Rad, Solaiman Shokur, Friedhelm Hummel, Silvestro Micera, Marc Pollefeys, Alexander Mathis
NeurIPS2
2024 HOISDF: Constraining 3D Hand-Object Pose Estimation with Global Signed Distance Fields
abstract
Human hands are highly articulated and versatile at handling objects. Jointly estimating the 3D poses of a hand and the object it manipulates from a monocular camera is challenging due to frequent occlusions. Thus, existing methods often rely on intermediate 3D shape representations to increase performance. These representations are typically explicit, such as 3D point clouds or meshes, and thus provide information in the direct surroundings of the intermediate hand pose estimate. To address this, we in-troduce HOISDF, a Signed Distance Field (SDF) guided hand-object pose estimation network, which jointly exploits hand and object SDFs to provide a global, implicit repre-sentation over the complete reconstruction volume. Specif-ically, the role of the SDFs is threefold: equip the visual encoder with implicit shape information, help to encode hand-object interactions, and guide the hand and object pose regression via SDF-based sampling and by augmenting the feature representations. We show that HOISDF achieves state-of-the-art results on hand-object pose esti-mation benchmarks (DexYCB and H03Dv2). Code is avail-able at https://github.com/amathislabIHOISDF.
Haozhe Qi, Chen Zhao 0025, Mathieu Salzmann, Alexander Mathis
CVPR1
2020 P2B: Point-to-Box Network for 3D Object Tracking in Point Clouds
abstract
Towards 3D object tracking in point clouds, a novel point-to-box network termed P2B is proposed in an end-to-end learning manner. Our main idea is to first localize potential target centers in 3D search area embedded with target information. Then point-driven 3D target proposal and verification are executed jointly. In this way, the time-consuming 3D exhaustive search can be avoided. Specifically, we first sample seeds from the point clouds in template and search area respectively. Then, we execute permutation-invariant feature augmentation to embed target clues from template into search area seeds and represent them with target-specific features. Consequently, the augmented search area seeds regress the potential target centers via Hough voting. The centers are further strengthened with seed-wise targetness scores. Finally, each center clusters its neighbors to leverage the ensemble power for joint 3D target proposal and verification. We apply PointNet++ as our backbone and experiments on KITTI tracking dataset demonstrate P2B's superiority (~10%'s improvement over state-of-the-art). Note that P2B can run with 40FPS on a single NVIDIA 1080Ti GPU. Our code and model are available at https://github.com/HaozheQi/P2B.
Haozhe Qi, Chen Feng 0002, Zhiguo Cao 0001, Yang Xiao 0007
CVPR1