EDBT 2026 Demo / reviewers in the wild / expert
Xiao-Ming Wu 0002
dblp:98/2898-2
· DBLP profile ↗
21ranked-venue papers
3as first author
21since 2021 · last 2025
0000-0003-1115-8551ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 18 · 3 first-author · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 16 · 3 first-author · 16 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Dynamic Spectral Graph Anomaly DetectionabstractGraph anomaly detection is crucial for identifying anomalous nodes within graphs and addressing applications like financial fraud detection and social spam detection. Recent spectral graph neural network methods advance graph anomaly detection by focusing on anomalies that notably affect the distribution of graph spectral energy. Such spectrum-based methods rely on two steps: graph wavelet extraction and feature fusion. However, both steps are hand-designed, capturing incomprehensive anomaly information of wavelet-specific features and resulting in their inconsistent feature fusion. To address these problems, we propose a dynamic spectral graph anomaly detection framework DSGAD to adaptively capture comprehensive anomaly information and perform consistent feature fusion. DSGAD introduces dynamic wavelets, consisting of trainable wavelets to adaptively learn anomalous patterns and capture wavelet-specific features with comprehensive anomaly information. Furthermore, the consistent fusion of wavelet-specific features achieves dynamic fusion by combining wavelet-specific feature extraction with energy difference and channel convolution fusion using location correlation. Experimental results on four datasets substantiate the efficacy of our DSGAD method, surpassing state-of-the-art methods in both homogeneous and heterogeneous graphs. Jianbo Zheng, Chao Yang 0015, Tairui Zhang, Longbing Cao, Bin Jiang 0006, Xuhui Fan 0001, Xiao-Ming Wu 0002, Xianxun Zhu |
AAAI | 7 |
| 2025 | Panorama Generation From NFoV Image Done RightabstractGenerating 360-degree panoramas from narrow field of view (NFoV) image is a promising computer vision task for Virtual Reality (VR) applications. Existing methods mostly assess the generated panoramas with InceptionNet or CLIP based metrics, which tend to perceive the image quality and is not suitable for evaluating the distortion. In this work, we first propose a distortion-specific CLIP, named Distort-CLIP to accurately evaluate the panorama distortion and discover the "visual cheating" phenomenon in previous works (i.e., tending to improve the visual results by sacrificing distortion accuracy). This phenomenon arises because prior methods employ a single network to learn the distinct panorama distortion and content completion at once, which leads the model to prioritize optimizing the latter. To address the phenomenon, we propose PanoDecouple, a decoupled diffusion model framework, which decouples the panorama generation into distortion guidance and content completion, aiming to generate panoramas with both accurate distortion and visual appeal. Specifically, we design a DistortNet for distortion guidance by imposing panorama-specific distortion prior and a modified condition registration mechanism; and a ContentNet for content completion by imposing perspective image information. Additionally, a distortion correction loss function with Distort-CLIP is introduced to constrain the distortion explicitly. The extensive experiments validate that PanoDecouple surpasses existing methods both in distortion and visual metrics. Dian Zheng, Xiao-Ming Wu 0002, Cao Li, Chengfei Lv, Jianfang Hu, Wei-Shi Zheng 0001 |
CVPR | 3 |
| 2025 | Rethinking Bimanual Robotic Manipulation: Learning with Decoupled Interaction FrameworkabstractBimanual robotic manipulation is an emerging and critical topic in the robotics community. Previous works primarily rely on integrated control models that take the perceptions and states of both arms as inputs to directly predict their actions. However, we think bimanual manipulation involves not only coordinated tasks but also various uncoordinated tasks that do not require explicit cooperation during execution, such as grasping objects with the closest hand, which integrated control frameworks ignore to consider due to their enforced cooperation in the early inputs. In this paper, we propose a novel decoupled interaction framework that considers the characteristics of different tasks in bimanual manipulation. The key insight of our framework is to assign an independent model to each arm to enhance the learning of uncoordinated tasks, while introducing a selective interaction module that adaptively learns weights from its own arm to improve the learning of coordinated tasks. Extensive experiments on seven tasks in the RoboTwin dataset demonstrate that: (1) Our framework achieves outstanding performance, with a 23.5% boost over the SOTA method. (2) Our framework is flexible and can be seamlessly integrated into existing methods. (3) Our framework can be effectively extended to multi-agent manipulation tasks, achieving a 28% boost over the integrated control SOTA. (4) The performance boost stems from the decoupled design itself, surpassing the SOTA by 16.5% in success rate with only 1/6 of the model size. Jian-Jian Jiang, Xiao-Ming Wu 0002, Yi-Xiang He, Ling-An Zeng, Yi-Lin Wei, Wei-Shi Zheng 0001 |
ICCV | 2 |
| 2025 | AffordDexGrasp: Open-Set Language-Guided Dexterous Grasp With Generalizable-Instructive Affordance
Yi-Lin Wei, Mu Lin, Jian-Jian Jiang, Xiao-Ming Wu 0002, Ling-An Zeng, Wei-Shi Zheng 0001 |
ICCV | 5 |
| 2025 | iManip: Skill-Incremental Learning for Robotic Manipulation
Zexin Zheng, Jia-Feng Cai, Xiao-Ming Wu 0002, Yi-Lin Wei, Yu-Ming Tang, Ancong Wu, Wei-Shi Zheng 0001 |
ICCV | 3 |
| 2025 | Supplementary Material for "NoiseActor: A Noise-Action Collaborative Framework for Privacy-Preserving Action Recognition without Privacy Labels"abstractOur supplementary material is organized into the following sections: • Section II provides the details of the LOCATION module. • Section III provides the evaluation protocol for the SBU dataset. • Section IV provides the evaluation protocol for cross-dataset experiment on the UCF101 dataset and the VISPR dataset. • Section V provides implementation details for each datasets. • Section VI provides more anonymized frames on the SBU dataset. • Section VII provides more anonymized frames on the UCF101 dataset. • Section VIII provides details about anonymized video file. • Section IX provides ablation studies of the adapter. Xiao Li 0074, Xiao-Ming Wu 0002, Delong Zhang, Kun-Yu Lin, Yi-Xing Peng, Ling-An Zeng, Wei-Shi Zheng 0001 |
ICME | 2 |
| 2025 | UGotMe: An Embodied System for Affective Human-Robot InteractionabstractEquipping humanoid robots with the capability to understand emotional states of human interactants and express emotions appropriately according to situations is essential for affective human-robot interaction. However, enabling current vision-aware multimodal emotion recognition models for affective human-robot interaction in the real-world raises embodiment challenges: addressing the environmental noise issue and meeting real-time requirements. First, in multi-party conversation scenarios, the noises inherited in the visual observation of the robot, which may come from either 1) distracting objects in the scene or 2) inactive speakers appearing in the field of view of the robot, hinder the models from extracting emotional cues from vision inputs. Secondly, real-time response, a desired feature for an interactive system, is also challenging to achieve. To tackle both challenges, we introduce an affective human-robot interaction system called UGotMe designed specifically for multiparty conversations. Two denoising strategies are proposed and incorporated into the system to solve the first issue. Specifically, to filter out distracting objects in the scene, we propose extracting face images of the speakers from the raw images and introduce a customized active face extraction strategy to rule out inactive speakers. As for the second issue, we employ efficient data transmission from the robot to the local server to improve real-time response capability. We deploy UGotMe on a human robot named Ameca to validate its real-time inference capabilities in practical scenarios. Videos demonstrating real-world deployment are available at https://lipzh5.github.io/HumanoidVLE/ Pei-Zhen Li, Longbing Cao, Xiao-Ming Wu 0002, Xiaohan Yu 0001 |
ICRA | 3 |
| 2025 | Deep Concept Forgetting in Text-to-Image Diffusion Models
Dian Zheng, Xiao-Ming Wu 0002, Wei-Shi Zheng 0001 |
PRCV (2) | 3 |
| 2025 | DiffuVolume: Diffusion Model for Volume based Stereo Matching
Dian Zheng, Xiao-Ming Wu 0002, Zuhao Liu 0002, Jingke Meng, Wei-Shi Zheng 0001 |
Int. J. Comput. Vis. | 2 |
| 2024 | Single-View Scene Point Cloud Human Grasp GenerationabstractIn this work, we explore a novel task of generating human grasps based on single-view scene point clouds, which more accurately mirrors the typical real-world situation of observing objects from a single viewpoint. Due to the incompleteness of object point clouds and the presence of numerous scene points, the generated hand is prone to penetrating into the invisible parts of the object and the model is easily affected by scene points. Thus, we introduce S2HGrasp, a framework composed of two key modules: the Global Perception module that globally perceives partial object point clouds, and the DiffuGrasp module designed to generate high-quality human grasps based on complex inputs that include scene points. Additionally, we introduce S2HGD dataset, which comprises approximately 99,000 single-object single-view scene point clouds of 1,668 unique objects, each annotated with one human grasp. Our extensive experiments demonstrate that S2HGrasp can not only generate natural human grasps regardless of scene points, but also effectively prevent penetration between the hand and invisible parts of the object. Moreover, our model showcases strong generalization capability when applied to unseen objects. Our code and dataset are available at https://github.com/iSEE-Laboratory/S2HGrasp. Yan-Kang Wang, Chengyi Xing, Yi-Lin Wei, Xiao-Ming Wu 0002, Wei-Shi Zheng 0001 |
CVPR | 4 |
| 2024 | Dexterous Grasp TransformerabstractIn this work, we propose a novel discriminative frame-work for dexterous grasp generation, named Dexterous Grasp TRansformer (DGTR), capable of predicting a di-verse set of feasible grasp poses by processing the object point cloud with only one forward pass. We formulate dex-terous grasp generation as a set prediction task and design a transformer-based grasping model for it. However, we identify that this set prediction paradigm encounters sev-eral optimization challenges in the field of dexterous grasping and results in restricted performance. To address these issues, we propose progressive strategies for both the training and testing phases. First, the dynamic-static matching training (DSMT) strategy is presented to enhance the opti-mization stability during the training phase. Second, we in-troduce the adversarial-balanced test-time adaptation (AB-TTA) with a pair of adversarial losses to improve grasping quality during the testing phase. Experimental results on the DexGraspNet dataset demonstrate the capability of DGTR to predict dexterous grasp poses with both high quality and diversity. Notably, while keeping high qual-ity, the diversity of grasp poses predicted by DGTR sig-nificantly outperforms previous works in multiple metrics without any data pre-processing. Codes are available at https://github.com/iSEE-Laboratory/DGTR. Guo-Hao Xu, Yi-Lin Wei, Dian Zheng, Xiao-Ming Wu 0002, Wei-Shi Zheng 0001 |
CVPR | 4 |
| 2024 | Selective Hourglass Mapping for Universal Image Restoration Based on Diffusion ModelabstractUniversal image restoration is a practical and poten-tial computer vision task for real-world applications. The main challenge of this task is handling the different degra-dation distributions at once. Existing methods mainly utilize task-specific conditions (e.g., prompt) to guide the model to learn different distributions separately, named multi-partite mapping. However, it is not suitable for universal model learning as it ignores the shared information between different tasks. In this work, we propose an advanced selective hourglass mapping strategy based on diffusion model, termed DiffUIR. Two novel considerations make our Dif-fUIR non-trivial. Firstly, we equip the model with strong condition guidance to obtain accurate generation direction of diffusion model (selective). More importantly, DiffUIR integrates a flexible shared distribution term (SDT) into the diffusion algorithm elegantly and naturally, which gradually maps different distributions into a shared one. In the reverse process, combined with SDT and strong condition guidance, DiffUIR iteratively guides the shared distribution to the task-specific distribution with high image quality (hourglass). Without bells and whistles, by only modifying the mapping strategy, we achieve state-of-the-art performance on five image restoration tasks, 22 benchmarks in the universal setting and zero-shot generalization setting. Surprisingly, by only using a lightweight model (only 0.89M), we could achieve outstanding performance. The source code and pre-trained models are available at https://github.com/iSEE-Laboratory/DiffUIR. Dian Zheng, Xiao-Ming Wu 0002, Shuzhou Yang, Jianfang Hu, Wei-Shi Zheng 0001 |
CVPR | 2 |
| 2024 | An Economic Framework for 6-DoF Grasp Detection
Xiao-Ming Wu 0002, Jia-Feng Cai, Jian-Jian Jiang, Dian Zheng, Yi-Lin Wei, Wei-Shi Zheng 0001 |
ECCV (27) | 1 |
| 2024 | iGrasp: An Interactive 2D-3D Framework for 6-DoF Grasp Detection
Jian-Jian Jiang, Xiao-Ming Wu 0002, Zibo Chen, Yi-Lin Wei, Wei-Shi Zheng 0001 |
ICPR (30) | 2 |
| 2024 | PixelFade: Privacy-preserving Person Re-identification with Noise-guided Progressive Replacement
Delong Zhang, Yi-Xing Peng, Xiao-Ming Wu 0002, Ancong Wu, Wei-Shi Zheng 0001 |
ACM Multimedia | 3 |
| 2024 | Grasp as You Say: Language-guided Dexterous Grasp GenerationabstractThis paper explores a novel task "Dexterous Grasp as You Say'' (DexGYS), enabling robots to perform dexterous grasping based on human commands expressed in natural language. However, the development of this field is hindered by the lack of datasets with natural human guidance; thus, we propose a language-guided dexterous grasp dataset, named DexGYSNet, offering high-quality dexterous grasp annotations along with flexible and fine-grained human language guidance. Our dataset construction is cost-efficient, with the carefully-design hand-object interaction retargeting strategy, and the LLM-assisted language guidance annotation system. Equipped with this dataset, we introduce the DexGYSGrasp framework for generating dexterous grasps based on human language instructions, with the capability of producing grasps that are intent-aligned, high quality and diversity. To achieve this capability, our framework decomposes the complex learning process into two manageable progressive objectives and introduce two components to realize them. The first component learns the grasp distribution focusing on intention alignment and generation diversity. And the second component refines the grasp quality while maintaining intention consistency. Extensive experiments are conducted on DexGYSNet and real world environments for validation. Yi-Lin Wei, Jian-Jian Jiang, Chengyi Xing, Xiantuo Tan, Xiao-Ming Wu 0002, Hao Li 0076, Mark R. Cutkosky, Wei-Shi Zheng 0001 |
NeurIPS | 5 |
| 2023 | Generating Anomalies for Video Anomaly Detection with Prompt-based Feature MappingabstractAnomaly detection in surveillance videos is a challenging computer vision task where only normal videos are available during training. Recent work released the first virtual anomaly detection dataset to assist real-world detection. However, an anomaly gap exists because the anomalies are bounded in the virtual dataset but unbounded in the real world, so it reduces the generalization ability of the virtual dataset. There also exists a scene gap between virtual and real scenarios, including scene-specific anomalies (events that are abnormal in one scene but normal in another) and scene-specific attributes, such as the viewpoint of the surveillance camera. In this paper, we aim to solve the problem of the anomaly gap and scene gap by proposing a prompt-based feature mapping framework (PFMF). The PFMF contains a mapping network guided by an anomaly prompt to generate unseen anomalies with unbounded types in the real scenario, and a mapping adaptation branch to narrow the scene gap by applying domain classifier and anomaly classifier. The proposed framework outperforms the state-of-the-art on three benchmark datasets. Extensive ablation experiments also show the effectiveness of our framework design. Zuhao Liu 0002, Xiao-Ming Wu 0002, Dian Zheng, Kun-Yu Lin, Wei-Shi Zheng 0001 |
CVPR | 2 |
| 2023 | Estimator Meets Equilibrium Perspective: A Rectified Straight Through Estimator for Binary Neural Networks TrainingabstractBinarization of neural networks is a dominant paradigm in neural networks compression. The pioneering work BinaryConnect uses Straight Through Estimator (STE) to mimic the gradients of the sign function, but it also causes the crucial inconsistency problem. Most of the previous methods design different estimators instead of STE to mitigate it. However, they ignore the fact that when reducing the estimating error, the gradient stability will decrease concomitantly. These highly divergent gradients will harm the model training and increase the risk of gradient vanishing and gradient exploding. To fully take the gradient stability into consideration, we present a new perspective to the BNNs training, regarding it as the equilibrium between the estimating error and the gradient stability. In this view, we firstly design two indicators to quantitatively demonstrate the equilibrium phenomenon. In addition, in order to balance the estimating error and the gradient stability well, we revise the original straight through estimator and propose a power function based estimator, Rectified Straight Through Estimator (ReSTE for short). Comparing to other estimators, ReSTE is rational and capable of flexibly balancing the estimating error with the gradient stability. Extensive experiments on CIFAR-10 and ImageNet datasets show that ReSTE has excellent performance and surpasses the state-of-the-art methods without any auxiliary modules or losses. Xiao-Ming Wu 0002, Dian Zheng, Zuhao Liu 0002, Wei-Shi Zheng 0001 |
ICCV | 1 |
| 2022 | Online Enhanced Semantic Hashing: Towards Effective and Efficient Retrieval for Streaming Multi-Modal DataabstractWith the vigorous development of multimedia equipments and applications, efficient retrieval of large-scale multi-modal data has become a trendy research topic. Thereinto, hashing has become a prevalent choice due to its retrieval efficiency and low storage cost. Although multi-modal hashing has drawn lots of attention in recent years, there still remain some problems. The first point is that existing methods are mainly designed in batch mode and not able to efficiently handle streaming multi-modal data. The second point is that all existing online multi-modal hashing methods fail to effectively handle unseen new classes which come continuously with streaming data chunks. In this paper, we propose a new model, termed Online enhAnced SemantIc haShing (OASIS). We design novel semantic-enhanced representation for data, which could help handle the new coming classes, and thereby construct the enhanced semantic objective function. An efficient and effective discrete online optimization algorithm is further proposed for OASIS. Extensive experiments show that our method can exceed the state-of-the-art models. For good reproducibility and benefiting the community, our code and data are already publicly available. Xiao-Ming Wu 0002, Xin Luo 0006, Yu-Wei Zhan, Chenlu Ding, Zhen-Duo Chen 0001, Xin-Shun Xu |
AAAI | 1 |
| 2022 | Weakly-Supervised Online Hashing with Refined Pseudo TagsabstractWith the rapid development of social media, various types of tags uploaded by social users are attached to the images. Compared to clean labels marked by experts, although user-provided tags are imperfect, e.g., wrong tags, reduplicative tags, or missing tags, they are more diverse, fine-grained, and informative. Currently, there exist several weakly-supervised hashing methods attempting to learn hash codes using tags as supervision. Although they could benefiting from the rich information contained in tags, most of them may defy the nature of social media data. In real scenarios, social media data appears in streaming fashion, but most weakly-supervised hashing methods are just batch-based which cannot effectively handle streaming data. To this end, only one weakly-supervised online hashing method has been proposed, but it is still far from enough to alleviate the negative effects of tags. Chenlu Ding, Xin Luo 0006, Xiao-Ming Wu 0002, Yu-Wei Zhan, Rui Li 0090, Xin-Shun Xu |
CIKM | 3 |
| 2022 | Discrete online cross-modal hashing
Yu-Wei Zhan, Yongxin Wang 0001, Xiao-Ming Wu 0002, Xin Luo 0006, Xin-Shun Xu |
Pattern Recognit. | 4 |