EDBT 2026 Demo / reviewers in the wild / expert
Yurong Fu
dblp:369/9461
· DBLP profile ↗
2ranked-venue papers
1as first author
2since 2021 · last 2026
—ORCID · unresolved
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Face, body and person analysis · 35% Vision and language · 17% Robot navigation and mapping · 17% |
Topics — the 6 heaviest of 6, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Face, body and person analysis › human pose estimation › 3d pose estimation
egocentric pose estimation |
1.0 | 1 | 2026 | SAME: Spatial-Aware Multimodal Egocentric Human Pose Estimation · AAAI 2026 |
Computer vision › Face, body and person analysis
human pose estimation |
1.0 | 1 | 2026 | SAME: Spatial-Aware Multimodal Egocentric Human Pose Estimation · AAAI 2026 |
Computer vision › Vision and language
multimodal fusion |
1.0 | 1 | 2026 | SAME: Spatial-Aware Multimodal Egocentric Human Pose Estimation · AAAI 2026 |
Robotics › Robot navigation and mapping › sensor fusion
visual-inertial fusion |
1.0 | 1 | 2026 | SAME: Spatial-Aware Multimodal Egocentric Human Pose Estimation · AAAI 2026 |
Computer vision › 3D vision › motion capture
human motion capture |
0.9 | 1 | 2025 | Motions as Queries: One-Stage Multi-Person Holistic Human Motion Capture · CVPR 2025 |
Computer vision › Video understanding and tracking › multi-object tracking
multi-person tracking |
0.9 | 1 | 2025 | Motions as Queries: One-Stage Multi-Person Holistic Human Motion Capture · CVPR 2025 |
Methods — techniques the papers use, named apart from their topics
temporal transformer · 1.0dual coordinate frame · 1.0deformable stereo attention · 1.0temporal cross attention · 0.9end-to-end training · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SAME: Spatial-Aware Multimodal Egocentric Human Pose EstimationabstractEgocentric human pose estimation (HPE) plays a crucial role in immersive applications such as virtual and augmented reality. However, existing methods relying on either visual or sparse inertial data alone often suffer from occlusion or ill-posed problems. In this work, we propose SAME, a novel spatial-aware multimodal fusion framework combining the complementary signals from the stereo images and sparse IMUs for accurate and robust egocentric HPE. It adopts a two-stage network based on a dual coordinate frame to mitigate the coordinate inconsistencies among the stereo cameras and the IMUs. In the first stage, the IMU signals are transformed into the local frame and iteratively fused with the stereo images for estimating 3D poses in the local frame. In the second stage, the local poses are transformed into the global frame with the 6DOF head poses provided by the head-mounted display's (HMD) SLAM algorithm and then temporally aggregated via a temporal Transformer network. Meanwhile, to achieve geometric and semantic alignment among multi-modal features, we present a depth-guided spatial-aware deformable stereo attention network and a modality-aware Transformer decoder for cross-view and cross-modal feature fusion. Extensive experiments demonstrate that our approach achieves state-of-the-art performance on the public EMHI multi-modal egocentric pose estimation benchmark. Yurong Fu, Yiqiang Feng, Haoqian Wang |
AAAI | 1 |
| 2025 | Motions as Queries: One-Stage Multi-Person Holistic Human Motion CaptureabstractExisting methods for capturing multi-person holistic human motions from a monocular video usually involve integrating the detector, the tracker, and the human pose & shape estimator into a cascaded system. Differently, we develop a one-stage multi-person holistic human motion capture system, which 1) employs only one network, enabling significant benefits from the end-to-end training on a large-scale dataset; 2) enables performance improving of the tracking module during training, avoiding being limited by a pre-trained tracker; 3) captures the motions of all individuals within a single shot, rather than tracking and estimating each person sequentially. In this system, each query within a temporal cross-attention module is responsible for the long motion of a specific individual, implicitly aggregating individual-specific information throughout the entire video. To further boost the proposed system from end-to-end training, we also construct a synthetic human video dataset, with multi-person and whole-body annotations. Extensive experiments across different datasets demonstrate both the efficacy and the efficiency of both the proposed method and the dataset. Codes are avaiable at https://github.com/KenkunLiu/MaQ. Kenkun Liu, Yurong Fu, Weihao Yuan 0001, Peihao Li 0003, Xiaodong Gu 0004, Lingteng Qiu, Haoqian Wang, Zilong Dong, Xiaoguang Han 0001 |
CVPR | 2 |