EDBT 2026 Demo / reviewers in the wild / expert
Xihao Wang
dblp:212/6309
· DBLP profile ↗
14ranked-venue papers
3as first author
12since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 2 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 4 since 2021Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Rethinking Surgical Smoke: A Smoke-Type-Aware Laparoscopic Video Desmoking Method and DatasetabstractElectrocautery or lasers will inevitably generate surgical smoke, which hinders the visual guidance of laparoscopic videos for surgical procedures. The surgical smoke can be classified into different types based on its motion patterns, leading to distinctive spatio-temporal characteristics across smoky laparoscopic videos. However, existing desmoking methods fail to account for such smoke-type-specific distinctions. Therefore, we propose the first Smoke-Type-Aware Laparoscopic Video Desmoking Network (STANet) by introducing two smoke types: Diffusion Smoke and Ambient Smoke. Specifically, a smoke mask segmentation sub-network is designed to jointly conduct smoke mask and smoke type predictions based on the attention-weighted mask aggregation, while a smokeless video reconstruction sub-network is proposed to perform specially desmoking on smoky features guided by two types of smoke mask. To address the entanglement challenges of two smoke types, we further embed a coarse-to-fine disentanglement module into the mask segmentation sub-network, which yields more accurate disentangled masks through the smoke-type-aware cross attention between non-entangled and entangled regions. In addition, we also construct the first large-scale synthetic video desmoking dataset with smoke type annotations. Extensive experiments demonstrate that our method not only outperforms state-of-the-art approaches in quality evaluations, but also exhibits superior generalization across multiple downstream surgical tasks. Qifan Liang, Zhen Han 0002, Xihao Wang, Zhongyuan Wang 0001, Bin Mei |
AAAI | 4 |
| 2026 | An angle-guided bidirectional feature transformation network for multi-frame tilt-angle face recognition
Wenqin Song, Xihao Wang, Zhen Han 0002, Kangli Zeng, Zhongyuan Wang 0001 |
Expert Syst. Appl. | 2 |
| 2026 | Gait Planning and Adaptive Impedance Control for Turning Walk Based on a Self-Balancing Lower Limb ExoskeletonabstractSelf-balancing lower limb exoskeletons (SBLLE) hold the promise of helping individuals with diverse mobility impairments regain their ability to walk. Equipping exoskeletons with stable turning capabilities is crucial for their real-world application. This paper presents an effective control framework that enables stable turning motions in SBLLE without relying on external assistive devices. A turn-specific gait planner has been designed to generate stable reference gaits. Moreover, an adaptive variable impedance controller (AVIC) is introduced, which adjusts the center of mass (CoM) motion in real time to maintain balance during walking, even when carrying different patients. The proposed approach is validated through simulations, comparative studies, and multi-stage human-subject experiments with ten participants, ranging from healthy users to individuals with severe lower-limb motor impairment. Results consistently demonstrate stable turning performance, with the Zero Moment Point (ZMP) remaining within the support polygon during the entire walking process. Xihao Wang, Shisheng Zhang, Weijie Sun 0001, Zengle Ren, Wujing Cao, Meng Yin, Xinyu Wu 0001 |
IEEE Trans Autom. Sci. Eng. | 1 |
| 2025 | A Dual-Task Learning Model for Temporal Knowledge Graph Entity Alignment
Jingwei Cheng, Xihao Wang, Fu Zhang 0001 |
DASFAA (3) | 2 |
| 2025 | Multimodal Re-Ranking for Heterogeneous Face Re-IdentificationabstractHeterogeneous face re-identification (Re-ID), aiming to match low-quality faces captured by disjoint visible light (VIS) and near-infrared (NIR) cameras, has become a critical application in video surveillance. However, the domain discrepancy between the NIR-VIS faces degrades the Re-ID performance. To solve this problem, this paper proposes a multimodal re-ranking method including two stages. Firstly, we utilize the VIS-NIR face bi-directional modality transformation based on the positive and negative samples separate training strategy to reduce domain discrepancy and generate the multimodal ranking lists of face Re-ID with complementarities. Secondly, we propose linear and nonlinear multimodal ranking lists fusion strategies based on single-modal and multi-modal k-reciprocal nearest neighbors (K-RNNs) to obtain a more accurate fused ranking list for face Re-ID. Extensive experiments on heterogeneous face datasets demonstrate the superior performance of our method over existing methods. Wenqin Song, Jiawei Zhang 0002, Zhen Han 0002, Yunfeng Xue, Xihao Wang, Zhongyuan Wang 0001 |
ICIP | 6 |
| 2025 | CL2OD: Class-Incremental Continual Learning for Outdoor 3D Object Detection in Autonomous DrivingabstractThe classic learning paradigm of deep neural networks is becoming increasingly difficult to handle the evergrowing real-world traffic data in the intelligent transportation domain. To manage more complex and constantly growing traffic data, continual learning has attracted significant attention as a promising ability. This ability allows models to study new tasks sequentially while maintaining the performance of previously learned tasks. However, due to the characteristics of LiDAR data being different from the indoor point clouds, most existing research on continual learning has limited exploration of the outdoor 3D point cloud scenarios. The feasibility of directly applying continual learning techniques to real-world intelligent traffic scenarios remains unclear. To narrow down this gap, we propose a class-incremental continual learning strategy for LiDAR point cloud object detection via knowledge distillation. To enhance the adaptability of dynamic traffic scenarios, we further extract the equivariant features of the point cloud. In the experiment section, we evaluate our approach on the nuScenes dataset and achieve promising results. These findings demonstrate the effectiveness of our method and highlight its potential for advancing continual learning in autonomous driving applications. Xihao Wang, Xiangzhong Liu |
IJCNN | 1 |
| 2025 | Cross-Level Fusion: Integrating Object Lists with Raw Sensor Data for 3D Object TrackingabstractSmart sensors and Vehicle-To-Everything (V2X) modules are commonly utilized in automotive perception systems, which primarily provide processed object lists rather than raw data. However, high-level fusion approaches suffer from significant information loss and representational misalignment due to the inherently abstract and sparse nature of these high-level outputs. We propose a novel cross-level fusion paradigm that enables bidirectional information flow between object lists and raw vision features within an end-to-end Transformer framework for 3D object detection and tracking. Our approach extracts inherent positional and dimensional cues from object lists to generate two outputs: structured query features that are fused with the initial learnable queries in the Transformer decoder, and soft Gaussian attention masks that guide feature extraction. This integrated mechanism not only improves tracking accuracy by synergistically combining object priors with fine-grained vision data but also promotes hardware economy and AI model sustainability by adapting legacy sensors to evolving sensor setups. To overcome the lack of dedicated datasets, we develop a pseudo object list generation pipeline that simulates realistic sensor tracking behavior. Experiments on the nuScenes dataset demonstrate significant performance gains over vision-only baselines and robust generalization across diverse noise levels, validating the efficacy of our cross-level fusion strategy. The code is available at: https://github.com/CesarLiu/DNF.git. Xiangzhong Liu, Xihao Wang, Hao Shen 0002 |
IROS | 2 |
| 2024 | Attr-Int: A Simple and Effective Entity Alignment Framework for Heterogeneous Knowledge GraphsabstractEntity alignment (EA) refers to the task of linking entities in different knowledge graphs (KGs). Existing EA methods rely heavily on structural isomorphism. However, in real-world KGs, aligned entities usually have non-isomorphic neighborhood structures, which paralyses the application of these structure-dependent methods. In this paper, we investigate and tackle the problem of entity alignment between heterogeneous KGs. First, we propose two new benchmarks to closely simulate real-world EA scenarios of heterogeneity. Then we conduct extensive experiments to evaluate the performance of representative EA methods on the new benchmarks. Finally, we propose a simple and effective entity alignment framework called Attr-Int, in which innovative attribute information interaction methods can be seamlessly integrated with any embedding encoder for entity alignment, improving the performance of existing entity alignment techniques. Experiments demonstrate that our framework outperforms the state-of-the-art approaches on two new benchmarks. Linyan Yang, Jingwei Cheng, Chuanhao Xu, Xihao Wang, Fu Zhang 0001 |
ICASSP | 4 |
| 2024 | ProEqBEV: Product Group Equivariant BEV Network for 3D Object Detection in Road Scenes of Autonomous DrivingabstractWith the rapid development of autonomous driving systems, 3D object detection based on Bird’s Eye View (BEV) in road scenes has witnessed great progress over the past few years. As a road scene exhibits a part-whole hierarchy between the within objects and the scene itself, simple parts (e.g., roads, lane lines, vehicles and pedestrians) can be assembled into progressively more complex shapes to form a BEV representation of the whole road scene. Therefore, a BEV often has multiple levels of freedom on motion, i.e., the rotation and the moving shift of the whole BEV, and the random movements of objects (e.g., pedestrians and vehicles) inside the BEV. However, most of the current single-sensor or multi-sensor fusion-based BEV object detection methods have not yet taken into account capturing such multi-level motion in a BEV. To address this problem, we propose a product group equivariant object detection network framework that is equivariant with respect to multiple levels of symmetry groups based on multi-sensor fusion. The proposed framework extracts local equivariant features of objects in point clouds, while global equivariant features are extracted in both point clouds and images. Furthermore, the network learns diverse rotation-equivariant features and mitigates a significant amount of detection errors caused by rotations of BEV and objects inside a BEV, thereby further enhancing the performance of object detection. The experiment results show that the network architecture significantly improves object detection on mAP and NDS, respectively. In addition, in order to demonstrate the effectiveness of the proposed local-multi-global equivariant components, we conduct sufficient ablation experiments. The results show that the individual components are indispensable for the object detection performance improvement of the overall network architecture. Jian Yang 0034, Ke Li 0005, Jianzhang Zheng, Xihao Wang, Mingsong Chen 0001, Xiong You, Xian Wei |
ICRA | 6 |
| 2023 | DuEqNet: Dual-Equivariance Network in Outdoor 3D Object Detection for Autonomous DrivingabstractOutdoor 3D object detection has played an essential role in the environment perception of autonomous driving. In complicated traffic situations, precise object recognition provides indispensable information for prediction and planning in the dynamic system, improving self-driving safety and reliability. However, with the vehicle's veering, the constant rotation of the surrounding scenario makes a challenge for the perception systems. Yet most existing methods have not focused on alleviating the detection accuracy impairment brought by the vehicle's rotation, especially in outdoor 3D detection. In this paper, we propose DuEqNet, which first introduces the concept of equivariance into 3D object detection network by leveraging a hierarchical embedded framework. The dual-equivariance of our model can extract the equivariant features at both local and global levels, respectively. For the local feature, we utilize the graph-based strategy to guarantee the equivariance of the feature in point cloud pillars. In terms of the global feature, the group equivariant convolution layers are adopted to aggregate the local feature to achieve the global equivariance. In the experiment part, we evaluate our approach with different baselines in 3D object detection tasks and obtain State-Of-The-Art performance. According to the results, our model presents higher accuracy on orientation and better prediction efficiency. Moreover, our dual-equivariance strategy exhibits the satisfied plug-and-play ability on various popular object detection frameworks to improve their performance. Xihao Wang, JiaMing Lei, Arafat Al-Jawari, Xian Wei |
ICRA | 1 |
| 2023 | Couplformer: Rethinking Vision Transformer with Coupling AttentionabstractWith the development of the self-attention mechanism, the Transformer model has demonstrated its outstanding performance in the computer vision domain. However, the massive computation brought from the full attention mechanism became a heavy burden for memory consumption. Sequentially, the limitation of memory consumption hinders the deployment of the Transformer model on the embedded system where the computing resources are limited. To remedy this problem, we propose a novel memory economy attention mechanism named Couplformer, which decouples the attention map into two sub-matrices and generates the alignment scores from spatial information. Our method enables the Transformer model to improve time and memory efficiency while maintaining expressive power. A series of different scale image classification tasks are applied to evaluate the effectiveness of our model. The result of experiments shows that on the ImageNet-1K classification task, the Couplformer can significantly decrease 42% memory consumption compared with the regular Transformer. Meanwhile, it accesses sufficient accuracy requirements, which outperforms 0.56% on Top-1 accuracy and occupies the same memory footprint. Besides, the Couplformer achieves state-of-art performance in MS COCO 2017 object detection and instance segmentation tasks. As a result, the Couplformer can serve as an efficient backbone in visual tasks and provide a novel perspective on deploying attention mechanisms for researchers. Xihao Wang, Hao Shen 0002, Peidong Liang, Xian Wei |
WACV | 2 |
| 2022 | Geodesic Self-Attention for 3D Point CloudsabstractDue to the outstanding competence in capturing long-range relationships, self-attention mechanism has achieved remarkable progress in point cloud tasks. Nevertheless, point cloud object often has complex non-Euclidean spatial structures, with the behavior changing dynamically and unpredictably. Most current self-attention modules highly rely on the dot product multiplication in Euclidean space, which cannot capture internal non-Euclidean structures of point cloud objects, especially the long-range relationships along the curve of the implicit manifold surface represented by point cloud objects. To address this problem, in this paper, we introduce a novel metric on the Riemannian manifold to capture the long-range geometrical dependencies of point cloud objects to replace traditional self-attention modules, namely, the Geodesic Self-Attention (GSA) module. Our approach achieves state-of-the-art performance compared to point cloud Transformers on object classification, few-shot classification and part segmentation benchmarks. Zihao Xu 0002, Xihao Wang, Mingsong Chen 0001, Xian Wei |
NeurIPS | 4 |
| 2018 | Constructing High Quality Sense-specific Corpus and Word Embedding via Unsupervised Elimination of Pseudo Multi-sense
Freda Shi, Xihao Wang |
LREC | 2 |
| 2017 | Off-Topic Spoken Response Detection Using Siamese Convolutional Neural Networks
Chong Min Lee, Su-Youn Yoon, Xihao Wang, Matthew Mulholland, Ikkyu Choi, Keelan Evanini |
INTERSPEECH | 3 |