Shaoxiang Wang

dblp:331/5700 · DBLP profile ↗
← Back
5ranked-venue papers
2as first author
5since 2021 · last 2026
0009-0003-6274-2958ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
3D vision · 48% Vision and language · 37% Information extraction and text analysis · 15%
Databases, data mining, and information retrieval
1 paper
Machine learning and data management · 100%

Topics — the 6 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning and data management
active learning
1.012026
DALL: Data Labeling via Data Programming and Active Learning Enhanced by Large Language Models · CHI 2026
Machine learning and data management
data annotation
1.012026
DALL: Data Labeling via Data Programming and Active Learning Enhanced by Large Language Models · CHI 2026
Computer vision › Vision and language
3d vision and language
0.812024
MiKASA: Multi-Key-Anchor & Scene-Aware Transformer for 3D Visual Grounding · CVPR 2024
Computer vision › 3D vision › 3d scene understanding
3d visual grounding
0.812024
MiKASA: Multi-Key-Anchor & Scene-Aware Transformer for 3D Visual Grounding · CVPR 2024
Natural language and speech › Information extraction and text analysis › data annotation
LLM-based annotation
0.312026
DALL: Data Labeling via Data Programming and Active Learning Enhanced by Large Language Models · CHI 2026
Computer vision › 3D vision
3d scene understanding
0.212024
MiKASA: Multi-Key-Anchor & Scene-Aware Transformer for 3D Visual Grounding · CVPR 2024

Methods — techniques the papers use, named apart from their topics

large language model · 2.0data programming · 2.0active learning · 2.0transformer · 0.8self-attention · 0.8multi-key-anchor · 0.8
YearPublicationVenuePosition
2026 DALL: Data Labeling via Data Programming and Active Learning Enhanced by Large Language Models
abstract
Deep learning models for natural language processing rely heavily on high-quality labeled datasets. However, existing labeling approaches often struggle to balance label quality with labeling cost. To address this challenge, we propose DALL, a text labeling framework that integrates data programming, active learning, and large language models. DALL introduces a structured specification that allows users and large language models to define labeling functions via configuration, rather than code. Active learning identifies informative instances for review, and the large language model analyzes these instances to help users correct labels and to refine or suggest labeling functions. We implement DALL as an interactive labeling system for text labeling tasks. Comparative, ablation, and usability studies demonstrate DALL’s efficiency, the effectiveness of its modules, and its usability.
Guozheng Li 0002, Shaoxiang Wang, Yu Zhang 0043, Pengcheng Cao, Chi Harold Liu
CHI3
2026 TalkingPose: Efficient Face and Gesture Animation with Feedback-guided Diffusion Model
abstract
Recent advancements in diffusion models have significantly improved the realism and generalizability of character-driven animation, enabling the synthesis of high-quality motion from just a single RGB image and a set of driving poses. Nevertheless, generating temporally coherent long-form content remains challenging. Existing approaches are constrained by computational and memory limitations, as they are typically trained on short video segments, thus performing effectively only over limited frame lengths and hindering their potential for extended coherent generation. To address these constraints, we propose TalkingPose, a novel diffusion-based framework specifically designed for producing long-form, temporally consistent human upper-body animations. TalkingPose leverages driving frames to precisely capture expressive facial and hand movements, transferring these seamlessly to a target actor through a stable diffusion backbone. To ensure continuous motion and enhance temporal coherence, we introduce a feedback-driven mechanism built upon image-based diffusion models. Notably, this mechanism does not incur additional computational costs or require secondary training stages, enabling the generation of animations with unlimited duration. Additionally, we introduce a comprehensive, large-scale dataset to serve as a new benchmark for human upper-body animation. Project page: https://dfki-av.github.io/TalkingPose
Alireza Javanmardi, Pragati Jaiswal, Tewodros Habtegebrial, Christen Millerdurai, Shaoxiang Wang, Alain Pagani, Didier Stricker
WACV5
2026 Inpaint360GS: Efficient Object-Aware 3D Inpainting via Gaussian Splatting for 360° Scenes
abstract
Despite recent advances in single-object front-facing inpainting using NeRF and 3D Gaussian Splatting (3DGS), inpainting in complex 360° scenes remains largely underexplored. This is primarily due to three key challenges: (i) identifying target objects in the 3D field of 360° environments, (ii) dealing with severe occlusions in multi-object scenes, which makes it hard to define regions to inpaint, and (iii) maintaining consistent and high-quality appearance across views effectively.To tackle these challenges, we propose Inpaint360GS, a flexible 360° editing framework based on 3DGS that supports multi-object removal and high-fidelity inpainting in 3D space. By distilling 2D segmentation into 3D and leveraging virtual camera views for contextual guidance, our method enables accurate object-level editing and consistent scene completion. We further introduce a new dataset tailored for 360° inpainting, addressing the lack of ground truth object-free scenes. Experiments demonstrate that Inpaint360GS outperforms existing baselines and achieves state-of-the-art performance. Project page: https://dfki-av.github.io/inpaint360gs/
Shaoxiang Wang, Shihong Zhang, Christen Millerdurai, Rüdiger Westermann, Didier Stricker, Alain Pagani
WACV1
2025 Uni-SLAM: Uncertainty-Aware Neural Implicit SLAM for Real-Time Dense Indoor Scene Reconstruction
abstract
Neural implicit fields have recently emerged as a powerful representation method for multi-view surface reconstruction due to their simplicity and state-of-the-art performance. However, reconstructing thin structures of indoor scenes while ensuring real-time performance remains a challenge for dense visual SLAM systems. Previous methods do not consider varying quality of input RGB-D data and employ fixed-frequency mapping process to reconstruct the scene, which could result in the loss of valuable information in some frames. In this paper, we propose Uni-SLAM, a decoupled 3D spatial representation based on hash grids for indoor reconstruction. We introduce a novel defined predictive uncertainty to reweight the loss function, along with strategic local-to-global bundle adjustment. Experiments on synthetic and real-world datasets demonstrate that our system achieves state-of-the-art tracking and mapping accuracy while maintaining real-time performance. It significantly improves over current methods with a 25% reduction in depth L1 error and a 66.86% completion rate within 1 cm on the Replica dataset, reflecting a more accurate reconstruction of thin structures. Project page: https://shaoxiang777.github.io/project/uni-slam/
Shaoxiang Wang, Yaxu Xie, Chun-Peng Chang, Christen Millerdurai, Alain Pagani, Didier Stricker
WACV1
2024 MiKASA: Multi-Key-Anchor & Scene-Aware Transformer for 3D Visual Grounding
abstract
3D visual grounding involves matching natural language descriptions with their corresponding objects in 3D spaces. Existing methods often face challenges with accuracy in object recognition and struggle in interpreting complex linguistic queries, particularly with descriptions that involve multiple anchors or are view-dependent. In response, we present the MiKASA (Multi-Key-Anchor Scene-Aware) Transformer. Our novel end-to-end trained model integrates a self-attention-based scene-aware object encoder and an original multi-key-anchor technique, enhancing object recognition accuracy and the understanding of spatial relationships. Furthermore, MiKASA improves the explainability of decision-making, facilitating error diagnosis. Our model achieves the highest overall accuracy in the Referit3D challenge for both the Sr3D and Nr3D datasets, particularly excelling by a large margin in categories that require viewpoint-dependent descriptions. The source code and additional resources for this project are available on GitHub: https://github.com/dfki-av/MiKASA-3DVG
Chun-Peng Chang, Shaoxiang Wang, Alain Pagani, Didier Stricker
CVPR2