EDBT 2026 Demo / reviewers in the wild / expert
Xu Wang 0053
dblp:181/2815-53
· DBLP profile ↗
13ranked-venue papers
6as first author
13since 2021 · last 2026
0000-0001-6852-1740ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-author · 8 since 2021Artificial intelligence and machine learning · 4 · 3 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Beyond Single-Speed Reasoning: Coordinating Fast and Slow Dynamics for Efficient World ModelingabstractModel-based reinforcement learning (MBRL) enables efficient decision-making by learning predictive world modelsof environment dynamics. Despite recent advances, existingmodels often struggle to reconcile accurate short-term transitions with coherent long-term planning, especially in partially observable or long-horizon settings. We argue that thislimitation often stems from modeling all transitions at a single temporal resolution, which makes it challenging to simultaneously capture fine-grained local dynamics and abstractglobal structures. To this end, we propose SF-RSSM (Slow-Fast Recurrent State-Space Model), a novel method that decouples short-term and long-term dynamics via a dualbranchdesign. The fast branch captures short-horizon transitions using residual prediction, while the slow branch models long-range dependencies with a GRU-based recurrent pathway.A distillation mechanism is developed to enable cooperationacross timescales, with the slow model providing soft targetsto guide the fast model. Additionally, a curiosity module encourages exploration by promoting learning in regions wherethe fast and slow branches exhibit divergent dynamics. Experiments on CARLA, DMControl and Atari benchmarks showthat SF-RSSM outperforms strong baselines in policy performance. Yangru Huang, Xu Wang 0053, Yi Jin 0001 |
AAAI | 4 |
| 2026 | Adaptive spatial-temporal graph ODE networks for traffic flow forecasting
Shixiang Han, Xu Wang 0053, Yi Jin 0001, Songhe Feng, Congyan Lang, Yidong Li |
Multim. Syst. | 2 |
| 2026 | Visual perception-inspired 3D point cloud samplingabstractTask-oriented sampling aims to predict the importance of points of a point cloud to better serve downstream tasks, which has attracted increasing attention in the fields of computer vision and visualization in recent years. However, existing methods cannot sufficiently leverage both global saliency and local saliency cues, resulting in suboptimal performance that requires further improvement. To tackle this challenge, we propose a novel 3D point cloud sampling method inspired by the human visual perception mechanism in this study, which can effectively extract important point cloud subsets from critical regions to better adapt to downstream tasks, thereby maintaining superior sampling performance. The proposed Visual Perception-inspired 3D Point Cloud Sampling (VPI-3DPS) method simulates the human visual system’s dynamic attention-shifting strategy by combining coarse-grained attention-driven sampling with fine-grained detail preservation. This allows our approach to adaptively capture both global context and local details within point cloud data, safeguarding downstream task performance. By leveraging Gated Recurrent Units (GRUs) for long-term dependency modeling and integrating Graph Convolutional Networks (GCNs) to capture local structures, VPI-3DPS obtains an integrated representation of regional correlation and detail awareness. Extensive experiments show that VPI-3DPS outperforms existing methods. Compared to the best-performing approaches, it achieves an average increase of 1.29% in classification accuracy, an average reduction of 13.20% in registration MRE, and an average decrease of 4.29% in Chamfer Distance for reconstruction. Xu Wang 0053, Yi Jin 0001, Hui Yu 0001, Yi-Gang Cen, Yidong Li |
Pattern Recognit. | 1 |
| 2025 | Federated Privacy Re-identification via Frequency Domain Splitting
Xuanwen Su, Xu Wang 0053, Tengfei Liang, Yi Jin 0001, Yidong Li |
ICIG (3) | 2 |
| 2025 | Generative Adversarial Network-based Image and Tabular Data Generation with Differential PrivacyabstractMachine learning and artificial intelligence technologies have become integral to various industries, driven by the availability of large-scale data. However, the use of sensitive industrial and personal data introduces significant privacy risks. Generative Adversarial Networks (GANs) are employed to generate synthetic data, thereby rendering them a feasible privacy-preserving technique. Despite their potential, existing private GANs face two major limitations: they are restricted to single-modal data generation or fail to ensure strict privacy guarantees. To tackle these issues, we propose Differential Privacy GAN of Image and Table—DPGAN-IT, which is capable of generating both image and tabular data simultaneously while enforcing robust privacy protection. Extensive experiments validate high utility of multi-modal synthetic data and demonstrate an effective balance between privacy and data utility. Jiming Yang, Xu Wang 0053, Yi Jin 0001, Yidong Li, Hui Yu 0001 |
ICME | 2 |
| 2025 | Open World Adaptive Pseudo Contrastive Learning for Generalized Category DiscoveryabstractIn this work, we investigate the challenging task of Generalized Category Discovery (GCD). Given datasets collected from open-world scenarios comprising both labeled and unlabeled images, GCD aims to classify all unlabeled images while simultaneously identifying unlabeled novel categories. The fundamental challenge in GCD tasks stems from inherent annotation discrepancies between seen and novel classes within the dataset. The lack of reliable label supervision for novel classes in unlabeled data leads to significant disparities in the model’s learning between old and novel classes, which is termed the bias risk. Recent advancements in GCD have employed the entropy maximization algorithm to alleviate the bias risk. However, they fail to provide debiased optimization for unlabeled data, leading to models that struggle with extracting discriminative features from such data. To address these challenges, we have created an Open-world pseudo-contrastive learning framework named OpcGCD. Our OpcGCD framework implements a dynamic category-wise threshold mechanism, which employs a parametric prototype classifiers to generate debiased pseudo-labels for unlabeled samples. To facilitate the learning of discriminative feature representations, our proposed OpcGCD employs debiased pseudo-labels in the formulation of a contrastive learning loss. Extensive evaluations conducted on multiple GCD benchmark datasets demonstrate the robustness and effectiveness of the approach. Yiqing Hao, Xu Wang 0053, Yi Jin 0001, Tao Wang 0011, Yidong Li, Shuoyan Liu, Chao Li 0026, Hui Yu 0001 |
SMC | 2 |
| 2025 | V2PNet: A Voxel-to-Point Network Framework for Task-Oriented Point Cloud SamplingabstractTask-oriented point cloud sampling is a fundamental technique in 3D computer vision and has become a crucial step in numerous 3D applications. However, most state-of-the-art task-oriented sampling methods adopt a point-wise analysis strategy, making them susceptible to data redundancy. Taking inspiration from the abstract-to-detailed recognition process of the human visual system, we propose a novel voxel-to-point network framework called V2PNet for task-oriented point cloud sampling. Specifically, we first design a lightweight coarse-grained sampling module named Important Voxel Prediction (IMVP). This module adaptively outputs points from significant regions of the point cloud by explicitly modeling inter-region relationships, thereby reducing interference from redundant points. Then, the V2PNet framework seamlessly integrates the IMVP module with existing point-wise and task-oriented sampling networks, enabling joint training with downstream tasks. This creates a task-oriented coarse-to-fine-grained sampling pipeline that effectively samples representative and informative points from significant regions to represent the original point cloud. Moreover, to mitigate disturbances across similar regions, we introduce a voxel simplification loss function to enhance the discriminative voxel prediction. Extensive experiments demonstrate that V2PNet improves the performance of existing state-of-the-art task-oriented sampling models. Xu Wang 0053, Yi Jin 0001, Yi-Gang Cen, Yidong Li, Hui Yu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | LighTN: Light-Weight Transformer Network for Performance-Overhead Tradeoff in Point Cloud DownsamplingabstractDownsampling is a crucial task for processing large scale and/or dense point clouds with limited resources. Owing to the development of deep learning, approaches of task-oriented point cloud downsampling have significant performance gains in preserving geometric information. However, most downsamling methods are limited by the disordered and unstructured point cloud data, making it difficult to continually improve the performance. To address this issue, we propose a light-weight Transformer network (LighTN) for the task-oriented point cloud downsampling as an end-to-end solution. In LighTN, we design an energy-efficient and permutation invariant single-head self-correlation module to extract refined global geometric features. Moreover, we present a novel sampling loss function to guide LighTN to focus on critical point cloud regions with more uniform distributions and prominent point coverage. Extensive experiments on classification, registration, and reconstruction tasks demonstrate that LighTN can achieve the state-of-the-art performance-overhead tradeoff and high-quality qualitative results. Xu Wang 0053, Yi Jin 0001, Yi-Gang Cen, Tao Wang 0011, Bowen Tang 0001, Yidong Li |
IEEE Trans. Multim. | 1 |
| 2024 | Enhancing Point Cloud Sampling Quality with Dual-Branch Fusion NetworksabstractTask-oriented point cloud sampling methods have attracted considerable attention for their ability to adaptively select important point sets based on downstream tasks, achieving an excellent balance between data simplification and task performance. However, existing task-oriented sampling models, primarily based on single-branch designs, struggle to fully extract features from input point clouds that comprehensively reflect multi-dimensional key information, thus limiting their sampling performance. In this paper, we introduce a dual-branch sampling network, named DBS-NET, which conducts crucial point sampling from both the global and local importance perspectives separately before merging them, thereby preserving multi-dimensional key information of the input data during the sampling process. Qualitative and quantitative experimental results demonstrate the competitive performance of DBS-NET on the classification benchmark task. Yi Jin 0001, Xu Wang 0053, Mengxia Hu, Hui Yu 0001, Yidong Li, Tao Wang 0011, Songhe Feng, Congyan Lang |
SMC | 2 |
| 2023 | RS-TNet: point cloud transformer with relation-shape awareness for fine-grained 3D visual processing
Xu Wang 0053, Yuqiao Zeng, Yi Jin 0001, Yi-Gang Cen, Baifu Liu, Shaohua Wan 0001 |
Soft Comput. | 1 |
| 2022 | Boundary Corrected Multi-Scale Fusion Network for Real-Time Semantic SegmentationabstractImage semantic segmentation aims at the pixel-level classification of images, which has requirements for both accuracy and speed in practical application. Existing semantic segmentation methods mainly rely on the high-resolution input to achieve high accuracy and do not meet the requirements of inference time. Although some methods focus on high-speed scene parsing with lightweight architectures, they can not fully mine semantic features under low computation with relatively low performance. To realize the real-time and high-precision segmentation, we propose a new method named Boundary Corrected Multi-scale Fusion Network, which uses the designed Low-resolution Multi-scale Fusion Module to extract semantic information. Moreover, to deal with boundary errors caused by low-resolution feature map fusion, we further design an additional Boundary Corrected Loss to constrain overly smooth features. Extensive experiments show that our method achieves a state-of-the-art balance of accuracy and speed for the real-time semantic segmentation. Tianjiao Jiang, Yi Jin 0001, Tengfei Liang, Xu Wang 0053, Yidong Li |
ICIP | 4 |
| 2022 | VSLN: View-aware sphere learning network for cross-view vehicle re-identificationabstractCross-view vehicle Reidentification (ReID) has attracted widespread attention as an increasingly important vision task in intelligent transportation and urban surveillance. Benefiting from Convolutional Neural Network (CNN), recent studies have promoted the development of vehicle ReID by extracting discriminative local features. However, two fundamental challenges of small interclass discrepancy caused by different views and large intraclass distance caused by similar appearance still hinder the performance of cross-view vehicle ReID. In this paper, a novel View-aware Sphere Learning Network (VSLN) is proposed to alleviate the above issues while maintaining the merits of CNN-based approaches to generate view-aware sphere-based features. First, a Sphere Feature Embedding Network (SFEN) is proposed to constrain the images into hypersphere for extracting sphere features. On the other hand, this study presents a sphere similarity triple loss to help SFEN concentrate more on robust and discriminative vehicle parts. Second, since the vehicle images are usually captured from different viewpoints, this study further extends SFEN by introducing a Vehicle Viewpoint Predictor (VVP) combined with global attention mechanism to enlarge the discrepancy of interclass and shorten the distance of intraclass. Moreover, a city-scale data set, named Vehicle from Different Viewpoints, containing image-level viewpoint labels, is collected for training VVP. As a result, the proposed VLSN can achieve 96.31% Top-1 accuracy and 79.46% Top-1 accuracy on VeRi-776 and VRIC data sets, respectively. Overall, extensive experimental results on two benchmark data sets show that the proposed VSLN outperforms state-of-the-art methods. Xu Wang 0053, Yi Jin 0001, Chenning Li, Yi-Gang Cen, Yidong Li |
Int. J. Intell. Syst. | 1 |
| 2021 | PST-NET: Point Cloud Sampling via Point-Based Transformer
Xu Wang 0053, Yi Jin 0001, Yi-Gang Cen, Congyan Lang, Yidong Li |
ICIG (3) | 1 |