EDBT 2026 Demo / reviewers in the wild / expert
Yiqun Zhao
dblp:317/6895
· DBLP profile ↗
5ranked-venue papers
2as first author
5since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer graphics and multimedia
3 papers |
Rendering · 52% Visual content generation and editing · 30% Geometric modeling and processing · 17% | |
| Artificial intelligence
4 papers |
Video understanding and tracking · 52% 3D vision · 48% |
Topics — the 13 heaviest of 14, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › 3D vision › 3d human reconstruction
human avatar reconstruction |
0.9 | 1 | 2025 | Surfel-Based Gaussian Inverse Rendering for Fast and Relightable Dynamic Human Reconstruction From Monocular Videos · IEEE Trans. Pattern Anal. Mach. Intell. 2025 |
Geometric modeling and processing
3d reconstruction |
0.9 | 1 | 2025 | 3D StreetUnveiler with Semantic-aware 2DGS - a simple baseline · ICLR 2025 |
Rendering
gaussian splatting |
0.9 | 1 | 2025 | 3D StreetUnveiler with Semantic-aware 2DGS - a simple baseline · ICLR 2025 |
Rendering
inverse rendering |
0.9 | 1 | 2025 | Surfel-Based Gaussian Inverse Rendering for Fast and Relightable Dynamic Human Reconstruction From Monocular Videos · IEEE Trans. Pattern Anal. Mach. Intell. 2025 |
Rendering
physically based rendering |
0.9 | 1 | 2025 | Surfel-Based Gaussian Inverse Rendering for Fast and Relightable Dynamic Human Reconstruction From Monocular Videos · IEEE Trans. Pattern Anal. Mach. Intell. 2025 |
Visual content generation and editing
video editing |
0.8 | 1 | 2024 | HeroMaker: Human-centric Video Editing with Motion Priors · ACM Multimedia 2024 |
Visual content generation and editing
video generation |
0.8 | 1 | 2024 | HeroMaker: Human-centric Video Editing with Motion Priors · ACM Multimedia 2024 |
Computer vision › Video understanding and tracking › human action analysis › action understanding
action counting |
0.6 | 1 | 2022 | TransRAC: Encoding Multi-scale Temporal Correlation with Transformers for Repetitive Action Counting · CVPR 2022 |
Computer vision › Video understanding and tracking › human action analysis › action understanding › action counting
repetitive action counting |
0.6 | 1 | 2022 | TransRAC: Encoding Multi-scale Temporal Correlation with Transformers for Repetitive Action Counting · CVPR 2022 |
Computer vision › Video understanding and tracking › human action analysis › action understanding
temporal action analysis |
0.6 | 1 | 2022 | TransRAC: Encoding Multi-scale Temporal Correlation with Transformers for Repetitive Action Counting · CVPR 2022 |
Computer vision › 3D vision › 3d reconstruction
single-view 3d reconstruction |
0.3 | 1 | 2025 | Surfel-Based Gaussian Inverse Rendering for Fast and Relightable Dynamic Human Reconstruction From Monocular Videos · IEEE Trans. Pattern Anal. Mach. Intell. 2025 |
Computer vision › 3D vision › 3d scene reconstruction
street scene reconstruction |
0.3 | 1 | 2025 | 3D StreetUnveiler with Semantic-aware 2DGS - a simple baseline · ICLR 2025 |
Computer vision › 3D vision › 3d motion analysis
human motion modeling |
0.2 | 1 | 2024 | HeroMaker: Human-centric Video Editing with Motion Priors · ACM Multimedia 2024 |
Methods — techniques the papers use, named apart from their topics
time-reversal framework · 1.7surfel-based gaussian splatting · 1.7semantic-aware inpainting · 1.7preintegration · 1.7occlusion approximation · 1.7image-based lighting · 1.72d gaussian splatting · 1.7neural deformation · 1.5diffusion model · 1.5canonical field learning · 1.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | 3D StreetUnveiler with Semantic-aware 2DGS - a simple baselineabstractUnveiling an empty street from crowded observations captured by in-car cameras is crucial for autonomous driving. However, removing all temporarily static objects, such as stopped vehicles and standing pedestrians, presents a significant challenge. Unlike object-centric 3D inpainting, which relies on thorough observation in a small scene, street scene cases involve long trajectories that differ from previous 3D inpainting tasks. The camera-centric moving environment of captured videos further complicates the task due to the limited degree and time duration of object observation. To address these obstacles, we introduce StreetUnveiler to reconstruct an empty street. StreetUnveiler learns a 3D representation of the empty street from crowded observations. Our representation is based on the hard-label semantic 2D Gaussian Splatting (2DGS) for its scalability and ability to identify Gaussians to be removed. We inpaint rendered image after removing unwanted Gaussians to provide pseudo-labels and subsequently re-optimize the 2DGS. Given its temporal continuous movement, we divide the empty street scene into observed, partial-observed, and unobserved regions, which we propose to locate through a rendered alpha map. This decomposition helps us to minimize the regions that need to be inpainted. To enhance the temporal consistency of the inpainting, we introduce a novel time-reversal framework to inpaint frames in reverse order and use later frames as references for earlier frames to fully utilize the long-trajectory observations. Our experiments conducted on the street scene dataset successfully reconstructed a 3D representation of the empty street. The mesh representation of the empty street can be extracted for further applications. Yikai Wang 0002, Yiqun Zhao, Yanwei Fu 0001, Shenghua Gao |
ICLR | 3 |
| 2025 | Surfel-Based Gaussian Inverse Rendering for Fast and Relightable Dynamic Human Reconstruction From Monocular VideosabstractEfficient and accurate reconstruction of a relightable, dynamic clothed human avatar from a monocular video is crucial for the entertainment industry. This article presents SGIA (Surfel-based Gaussian Inverse Avatar), which introduces efficient training and rendering for relightable dynamic human reconstruction. SGIA advances previous Gaussian Avatar methods by comprehensively modeling Physically-Based Rendering (PBR) properties for clothed human avatars, allowing for the manipulation of avatars into novel poses under diverse lighting conditions. Specifically, our approach integrates pre-integration and image-based lighting for fast light calculations that surpass the performance of existing implicit-based techniques. To address challenges related to material lighting disentanglement and accurate geometry reconstruction, we propose an innovative occlusion approximation strategy and a progressive training approach. Extensive experiments demonstrate that SGIA not only achieves highly accurate physical properties but also significantly enhances the realistic relighting of dynamic human avatars, providing a substantial speed advantage. Yiqun Zhao, Chenming Wu, Binbin Huang 0004, Yihao Zhi, Chen Zhao 0011, Jingdong Wang 0001, Shenghua Gao |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2024 | RoomDesigner: Encoding Anchor-latents for Style-consistent and Shape-compatible Indoor Scene GenerationabstractIndoor scene generation aims at creating shape-compatible, style-consistent furniture arrangements within a spatially reasonable layout. However, most existing approaches primarily focus on generating plausible furniture layouts without incorporating specific details related to individual furniture. To address this limitation, we propose a two-stage model integrating shape priors into the indoor scene generation by encoding furniture as anchor latent representations. In the first stage, we employ discrete vector quantization to encode each piece of furniture as anchor-latents. Based on the anchor-latents representation, the shape and location information of furniture was characterized by a concatenation of location, size, orientation, class, and our anchor latent. In the second stage, we leverage a transformer model to predict indoor scenes configuration autoregressively. Thanks to the proposed anchor-latents representations, our generative model can synthesis furniture in diverse shapes and produce physically plausible arrangements with shape-compatible and style-consistent furniture. Furthermore, our method facilitates various human interaction applications, such as style-consistent scene completion, object mismatch correction, and controllable object-level editing. Experimental results on the 3D-Front dataset demonstrate that our approach can generate more consistent and compatible indoor scenes compared to existing methods, even without shape retrieval. Additionally, extensive ablation studies confirm the effectiveness of our design choices in the indoor scene generation model. Yiqun Zhao, Zibo Zhao 0001, Jing Li 0117, Sixun Dong, Shenghua Gao |
3DV | 1 |
| 2024 | HeroMaker: Human-centric Video Editing with Motion PriorsabstractVideo generation and editing, particularly human-centric video editing, has seen a surge of interest in its potential to create immersive and dynamic content. A fundamental challenge is ensuring temporal coherence and visual harmony across frames, especially in handling large-scale human motion and maintaining consistency over long sequences. The previous methods, such as zero-shot text-to-video methods with diffusion model, struggle with flickering and length limitations. In contrast, methods employing Video-2D representations grapple with accurately capturing complex structural relationships in large-scale human motion. Simultaneously, some patterns on the human body appear intermittently throughout the video, posing a knotty problem in identifying visual correspondence. To address the above problems, we present HeroMaker. This human-centric video editing framework manipulates the person's appearance within the input video and achieves consistent results across frames. Specifically, we propose to learn the motion priors, which represent the correspondences between dual canonical fields and each video frame, by leveraging the body mesh-based human motion warping and neural deformation-based margin refinement in the video reconstruction framework to ensure the semantic correctness of canonical fields. HeroMaker performs human-centric video editing by manipulating the dual canonical fields and combining them with motion priors to synthesize temporally coherent and visually plausible results. Comprehensive experiments demonstrate that our approach surpasses existing methods regarding temporal consistency, visual quality, and semantic coherence. Zibo Zhao 0001, Yihao Zhi, Yiqun Zhao, Binbin Huang 0004, Ruoyu Wang 0014, Michael Xuan, Shenghua Gao |
ACM Multimedia | 4 |
| 2022 | TransRAC: Encoding Multi-scale Temporal Correlation with Transformers for Repetitive Action CountingabstractCounting repetitive actions are widely seen in human activities such as physical exercise. Existing methods focus on performing repetitive action counting in short videos, which is tough for dealing with longer videos in more realistic scenarios. In the data-driven era, the degradation of such generalization capability is mainly attributed to the lack of long video datasets. To complement this margin, we introduce a new large-scale repetitive action counting dataset covering a wide variety of video lengths, along with more realistic situations where action interruption or action inconsistencies occur in the video. Besides, we also provide a fine-grained annotation of the action cycles instead of just counting annotation along with a numerical value. Such a dataset contains 1,451 videos with about 20,000 annotations, which is more challenging. For repetitive action counting towards more realistic scenarios, we further propose encoding multi-scale temporal correlation with transformers that can take into account both performance and efficiency. Furthermore, with the help of fine-grained annotation of action cycles, we propose a density map regression-based method to predict the action period, which yields better performance with sufficient interpretability. Our proposed method outperforms state-of-the-art methods on all datasets and also achieves better performance on the unseen dataset without fine-tuning. The dataset and code are available11https://svip-lab.github.io/dataset/RepCount_dataset.html. Huazhang Hu, Sixun Dong, Yiqun Zhao, Dongze Lian, Shenghua Gao |
CVPR | 3 |