VLDB 2026 Research / reviewers in the wild / expert
Xueyang Zhang
dblp:210/5124
· DBLP profile ↗
17ranked-venue papers
1as first author
15since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 1 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DriveLiDAR4D: Sequential and Controllable LiDAR Scene Generation for Autonomous DrivingabstractThe generation of realistic LiDAR point clouds plays a crucial role in the development and evaluation of autonomous driving systems. Although recent methods for 3D LiDAR point cloud generation have shown significant improvements, they still face notable limitations, including the lack of sequential generation capabilities and the inability to produce accurately positioned foreground objects and realistic backgrounds. These shortcomings hinder their practical applicability. In this paper, we introduce DriveLiDAR4D, a novel LiDAR generation pipeline consisting of multimodal conditions and a novel sequential noise prediction model LiDAR4DNet, capable of producing temporally consistent LiDAR scenes with highly controllable foreground objects and realistic backgrounds. To the best of our knowledge, this is the first work to address the sequential generation of LiDAR scenes with full scene manipulation capability in an end-to-end manner. We evaluated DriveLiDAR4D on the nuScenes and KITTI datasets, where we achieved an FRD score of 743.13 and an FVD score of 16.96 on the nuScenes dataset, surpassing the current state-of-the-art (SOTA) method, UniScene, with an performance boost of 37.2% in FRD and 24.1% in FVD, respectively. Kaiwen Cai, Hengtong Hu, Xueyang Zhang, Kun Zhan, Yifei Zhan, Xianpeng Lang |
AAAI | 7 |
| 2026 | CorrectAD: A Self-Correcting Agentic System to Improve End-to-end Planning in Autonomous DrivingabstractEnd-to-end planning methods are the de-facto standard of the current autonomous driving system, while the robustness of the data-driven approaches suffers due to the notorious long-tail problem (i.e., rare but safety-critical failure cases). In this work, we explore whether recent diffusion-based video generation methods (a.k.a. world models), paired with structured 3D layouts, can enable a fully automated pipeline to self-correct such failure cases. We first introduce an agent to simulate the role of product manager, dubbed PM-Agent, which formulates data requirements to collect data similar to the failure cases. Then, we use a generative model that can simulate both data collection and annotation. However, existing generative models struggle to generate high-fidelity data conditioned on 3D layouts. To address this, we propose DriveSora, which can generate spatiotemporally consistent videos aligned with the 3D annotations requested by PM-Agent. We integrate these components into our self-correcting agentic system, CorrectAD. Importantly, our pipeline is end-to-end model agnostic and can be applied to improve any end-to-end planner. Evaluated on both nuScenes and a more challenging in-house dataset across multiple end-to-end planners, CorrectAD corrects 62.5% and 49.8% of failure cases, reducing collision rates by 39% and 27%, respectively. Enhui Ma, Junpeng Jiang, Kun Zhan, Xueyang Zhang, Xianpeng Lang, Di Lin 0002, Kaicheng Yu |
AAAI | 9 |
| 2025 | ReconDreamer: Crafting World Models for Driving Scene Reconstruction via Online RestorationabstractClosed-loop simulation is crucial for end-to-end autonomous driving. Existing sensor simulation methods (e.g., NeRF and 3DGS) reconstruct driving scenes based on conditions that closely mirror training data distributions. However, these methods struggle with rendering novel trajectory, such as lane changes. Recent works have demonstrated that integrating world model knowledge alleviates these issues. Despite their efficiency, these approaches still encounter difficulties in the accurate representation of more complex maneuvers, with multi-lane shifts being a notable example. Therefore, we introduce ReconDreamer, which enhances driving scene reconstruction through incremental integration of world model knowledge. Specifically, DriveRestorer is proposed to mitigate artifacts via online restoration. This is complemented by a progressive data update strategy designed to ensure high-quality rendering for more complex maneuvers. To the best of our knowledge, ReconDreamer is the first method to effectively render in large maneuvers. Experimental results demonstrate that ReconDreamer outperforms Street Gaussians in the NTA-IoU, NTL-IoU, and FID, with relative improvements by 24.87%, 6.72%, and 29.97%. Furthermore, ReconDreamer surpasses DriveDreamer4D with PVG during large maneuver rendering, as verified by a relative improvement of 195.87% in the NTA-IoU metric and a user study. Chaojun Ni, Guosheng Zhao, Wenkang Qin, Guan Huang 0003, Yuyin Chen, Xueyang Zhang, Yifei Zhan, Kun Zhan, Peng Jia 0007, Xianpeng Lang, Xingang Wang 0003, Wenjun Mei |
CVPR | 10 |
| 2025 | DriveDreamer4D: World Models Are Effective Data Machines for 4D Driving Scene RepresentationabstractClosed-loop simulation is essential for advancing end-to-end autonomous driving systems. Contemporary sensor simulation methods, such as NeRF and 3DGS, rely predominantly on conditions closely aligned with training data distributions, which are largely confined to forward-driving scenarios. Consequently, these methods face limitations when rendering complex maneuvers (e.g., lane change, acceleration, deceleration). Recent advancements in autonomous-driving world models have demonstrated the potential to generate diverse driving videos. However, these approaches remain constrained to 2D video generation, inherently lacking the spatiotemporal coherence required to capture intricacies of dynamic driving environments. In this paper, we introduce DriveDreamer4D, which enhances 4D driving scene representation leveraging world model priors. Specifically, we utilize the world model as a data machine to synthesize novel trajectory videos, where structured conditions are explicitly leveraged to control the spatial-temporal consistency of traffic elements. Besides, the cousin data training strategy is proposed to facilitate merging real and synthetic data for optimizing 4DGS. To our knowledge, DriveDreamer4D is the first to utilize video generation models for improving 4D reconstruction in driving scenarios. Experimental results reveal that DriveDreamer4D significantly enhances generation quality under novel trajectory views, achieving a relative improvement in FID by 32.1%, 46.4%, and 16.3% compared to PVG, S3Gaussian, and Deformable-GS. Moreover, DriveDreamer4D markedly enhances the spatiotemporal coherence of driving agents, which is verified by a comprehensive user study and the relative increases of 22.6%, 43.5%, and 15.6% in the NTA-IoU metric. Guosheng Zhao, Chaojun Ni, Xueyang Zhang, Guan Huang 0003, Xinze Chen, Youyi Zhang, Wenjun Mei, Xingang Wang 0003 |
CVPR | 5 |
| 2025 | HiNeuS: High-Fidelity Neural Surface Mitigating Low-Texture and Reflective AmbiguityabstractNeural surface reconstruction faces persistent challenges in reconciling geometric fidelity with photometric consistency under complex scene conditions. We present HiNeuS, a unified framework that holistically addresses three core limitations in existing approaches: multi-view radiance inconsistency, missing keypoints in textureless regions, and structural degradation from over-enforced Eikonal constraints during joint optimization. To resolve these issues through a unified pipeline, we introduce: 1) Differential visibility verification through SDF-guided ray tracing, resolving reflection ambiguities via continuous occlusion modeling; 2) Planar-conformal regularization via ray-aligned geometry patches that enforce local surface coherence while preserving sharp edges through adaptive appearance weighting; and 3) Physically-grounded Eikonal relaxation that dynamically modulates geometric constraints based on local radiance gradients, enabling detail preservation without sacrificing global regularity. Unlike prior methods that handle these aspects through sequential optimizations or isolated modules, our approach achieves cohesive integration where appearance-geometry constraints evolve synergistically throughout training. Comprehensive evaluations across synthetic and real-world datasets demonstrate state-of-the-art performance, including a 21.4% reduction in Chamfer distance over reflection-aware baselines and 2.32 dB PSNR improvement against neural rendering counterparts. Qualitative analyses reveal superior capability in recovering specular instruments, urban layouts with centimeter-scale infrastructure, and low-textured surfaces without local patch collapse. The method's generalizability is further validated through successful application to inverse rendering tasks, including material decomposition and view-consistent relighting. Xueyang Zhang, Kun Zhan, Peng Jia 0007, Xianpeng Lang |
ICCV | 2 |
| 2025 | A Dental Periapical X-ray Images Segmentation Network Based on Pixel-Wise Contrastive Learning with Dual Attention MechanismsabstractPeriapical radiographs tend to have poor quality due to factors like acquisition techniques, equipment limitations, and patient differences. These factors lead to discrepancies in images, making precise segmentation challenging. However, accurate segmentation is crucial in dental practices. To address this issue, deep learning can be employed to improve segmentation accuracy and efficiency, thereby providing more reliable support for clinical diagnosis. In this study, we propose a deep learning network based on an encoder-decoder architecture, which integrates a dual attention mechanism and pixel-wise contrastive learning to address the tooth segmentation problem. We design the Dual Attention Contrast (DAC) module, which enhances feature maps through joint spatial and channel attention before utilizing optimized multi-scale features for both segmentation prediction and pixel-wise contrastive learning. This module, implemented with a multi-level deployment strategy, strengthens the network's ability to extract discriminative features from the multi-scale anatomical structures in periapical radiographs while constructing a globally structured feature space across datasets. The dual attention mechanism improves the recognition of key local features through spatial-channel collaborative calibration, while pixel-wise contrastive learning explicitly constrains the topological structure of the feature space, effectively mitigating the impact of inter-image variations on segmentation accuracy. Experimental results show that the proposed model outperforms current mainstream models across various evaluation metrics in periapical radiograph segmentation tasks. Yibin Tian, Zhiyuan Zhang 0004, Xueyang Zhang, Shan LianLei |
IJCNN | 5 |
| 2025 | PosePilot: Steering Camera Pose for Generative World Models with Self-supervised DepthabstractRecent advancements in autonomous driving (AD) systems have highlighted the potential of world models in achieving robust and generalizable performance across both ordinary and challenging driving conditions. However, a key challenge remains: precise and flexible camera pose control, which is crucial for accurate viewpoint transformation and realistic simulation of scene dynamics. In this paper, we introduce PosePilot, a lightweight yet powerful framework that significantly enhances camera pose controllability in generative world models. Drawing inspiration from self-supervised depth estimation, PosePilot leverages structure-from-motion principles to establish a tight coupling between camera pose and video generation. Specifically, we incorporate self-supervised depth and pose readouts, allowing the model to infer depth and relative camera motion directly from video sequences. These outputs drive pose-aware frame warping, guided by a photometric warping loss that enforces geometric consistency across synthesized frames. To further refine camera pose estimation, we introduce a reverse warping step and a pose regression loss, improving viewpoint precision and adaptability. Extensive experiments on autonomous driving and general-domain video datasets demonstrate that PosePilot significantly enhances structural understanding and motion reasoning in both diffusion-based and auto-regressive world models. By steering camera pose with self-supervised depth, PosePilot sets a new benchmark for pose controllability, enabling physically consistent, reliable viewpoint synthesis in generative world models. Bu Jin, Weize Li 0001, Baihan Yang, Zhenxin Zhu, Junpeng Jiang, Huan-ang Gao, Kun Zhan, Hengtong Hu, Xueyang Zhang, Peng Jia 0007, Hao Zhao 0002 |
IROS | 10 |
| 2025 | OmniGen: Unified Multimodal Sensor Generation for Autonomous DrivingabstractAutonomous driving has seen remarkable advancements, largely driven by extensive real-world data collection. However, acquiring diverse and corner-case data remains costly and inefficient. Generative models have emerged as a promising solution by synthesizing realistic sensor data. However, existing approaches primarily focus on single-modality generation, leading to inefficiencies and misalignment in multimodal sensor data. To address these challenges, we propose OminiGen, which generates aligned multimodal sensor data in a unified framework. Our approach leverages a shared Bird's Eye View (BEV) space to unify multimodal features and designs a novel generalizable multimodal reconstruction method, UAE, to jointly decode LiDAR and multi-view camera data. UAE achieves multimodal sensor decoding through volume rendering, enabling accurate and flexible reconstruction. Furthermore, we incorporate a Diffusion Transformer (DiT) with a ControlNet branch to enable controllable multimodal sensor generation. Our comprehensive experiments demonstrate that OminiGen achieves desired performances in unified multimodal sensor data generation with multimodal consistency and flexible sensor adjustments. Enhui Ma, Tianyi Yan, Xueyang Zhang, Kun Zhan, Peng Jia 0007, Xianpeng Lang, Jiawang Bian, Kaicheng Yu, Xiaodan Liang |
ACM Multimedia | 6 |
| 2025 | RLGF: Reinforcement Learning with Geometric Feedback for Autonomous Driving Video GenerationabstractSynthetic data is crucial for advancing autonomous driving (AD) systems, yet current state-of-the-art video generation models, despite their visual realism, suffer from subtle geometric distortions that limit their utility for downstream perception tasks.
We identify and quantify this critical issue, demonstrating a significant performance gap in 3D object detection when using synthetic versus real data.
To address this, we introduce Reinforcement Learning with Geometric Feedback (RLGF), RLGF uniquely refines video diffusion models by incorporating rewards from specialized latent-space AD perception models.
Its core components include an efficient Latent-Space Windowing Optimization technique for targeted feedback during diffusion, and a Hierarchical Geometric Reward (HGR) system providing multi-level rewards for point-line-plane alignment, and scene occupancy coherence.
To quantify these distortions, we propose GeoScores. Applied to models like DiVE on nuScenes, RLGF substantially reduces geometric errors (e.g., VP error by 21\%, Depth error by 57\%) and dramatically improves 3D object detection mAP by 12.7\%, narrowing the gap to real-data performance. RLGF offers a plug-and-play solution for generating geometrically sound and reliable synthetic videos for AD development. Tianyi Yan, Wencheng Han, Xueyang Zhang, Kun Zhan, Cheng-Zhong Xu 0001, Jianbing Shen |
NeurIPS | 4 |
| 2025 | Detection of Incomplete Root Canal Obturations in Dental X-ray Images via Spatial-Semantic Attention and Dynamic Feature CalibrationabstractTo address the challenges of low resolution, loss of small target features, and interference from complex anatomical structures in detecting incomplete root canal obturations in dental periapical radiographs, this article proposes an improved YOLOv8 model. First, we design a Convolution module with Space-to-Depth Transformation (SDT-Conv) that preserves feature map resolution through spatial depth-wise separable convolutions, effectively mitigating loss of small targets caused by downsampling operations. Second, we construct a Dynamic Iterative Token Aggregator (DITA) architecture that enhances global feature representation through hyper-token spatial aggregation and semantic correlation, while employing a spatial-semantic dual-stream attention mechanism to strengthen multiscale feature fusion capabilities, thereby providing richer feature information for the entire network. Finally, we embed an Efficient Multiscale Attention (EMA) dynamic calibration mechanism in the detection head, which optimizes feature responses through cross-channel weight adaptation, enabling the model to precisely localize small object boundaries. The experimental results demonstrate that the improved model achieves 81.5% mAP@50 on the validation set, representing a 12.8% improvement over YOLOv8n. It effectively overcomes the challenges posed by variations in obturation materials, dental structure occlusions, and low-contrast interference. Zhiqi Ren, Shanglei Chai, Zhiyuan Zhang 0004, Xueyang Zhang, Yibin Tian |
SMC | 4 |
| 2024 | Teeth Segmentation from Bite-Wing X-Ray Images by Integrating Nested Dual UNet with Swin TransformersabstractIn medical practice, the precision of image segmentation is crucial for diagnosis and treatment evaluations. Specifically, in dentistry, accurate teeth segmentation from bite-wing images is important for automatic and objective evaluations of root canal treatments. This study introduces$\text{Swin}-\mathrm{U}^{2} \text{Net}$, a model merging the nested dual UNet with residual U-block and Swin Transformers. It combines the local feature extraction capability of the former and the global attention and context understanding of the latter. It has been evaluated for tooth root segmentation using 500 bite-wing dental x-ray images obtained from a root canal treatment clinic. It achieved the best segmentation outcome in terms of Intersection over Union (IOU) and the third best result in terms of Dice Similarity Coefficient (DSC) with the second least amount of network parameters among six UNet-like models, thus it is effective and efficient. Yibin Tian, Zhiyuan Zhang 0004, Xueyang Zhang, Bingran Du |
SMC | 4 |
| 2024 | Research of Federated Learning Application Methods and Social ResponsibilityabstractFederated learning is a multi-party distributed machine learning system that allows each participant to complete the training task without their data out of the locality. The real meaning of the realization of data availability is invisible to protect personal privacy and data security. Although some institutions need to obtain their data's benefit by sharing their own data sets, the potential trust risks lead to the data not being better utilized cooperatively. In addition,regulatory requirements on data computing and sharing are different in various countries. These contribute to enormous challenges in the actual application of federated learning. In this research, we make a comprehensive review of the application patterns of federated learning in different fields. Then, we illustrate the social responsibilities of federated learning from four dimensions, including compliance application, system security mechanism, the trust mechanism, and ethical security. Finally, based on the current characteristics and regulatory requirements of federated learning, we discuss the future research directions for federated learning. Meiyun Xie, Xueyang Zhang |
IEEE Trans. Big Data | 4 |
| 2023 | Frame-Level Embedding Learning for Few-shot Bioacoustic Event DetectionabstractWe propose an effective frame-level embedding learning framework for few-shot bioacoustic event detection (FSBED). First, the duration of different animal calls varies greatly, so we innovatively propose a frame-level embedding learning scheme, which can obtain adaptive event receptive fields with more accurate frame-level units. Next, we develop a transfer learning-based approach to deal with the mismatch between training and testing data. Finally, we use the idea of semi-supervised learning to solve the problem of too little labeled data in few-shot learning. By incorporating these several sets of techniques, our overall system ranked first place in the FSBED task of Detection and Classification of Acoustic Scenes and Events (DCASE) Challenge 2022. Xueyang Zhang, Jun Du 0002, Genwei Yan, Jigang Tang, Tian Gao 0005, Jianqing Gao |
ICME | 1 |
| 2023 | Nonlocal Correntropy Matrix Representation for Hyperspectral Image ClassificationabstractHyperspectral image (HSI) classification is a hot topic in the remote sensing community. However, it is challenging to fully use spatial–spectral information for HSI classification due to the high dimensionality of the data, high intraclass variability, and the limited availability of training samples. To deal with these issues, we propose a novel feature extraction method called nonlocal correntropy matrix (NLCM) representation in this letter. NLCM can characterize the spectral correlation and effectively extract discriminative features for HSI classification. We verify the effectiveness of the proposed method on two widely used datasets. The results show that NLCM performs better than the state-of-the-art methods, especially when the training set size is small. Furthermore, the experimental results also demonstrate that the proposed method outperforms compared methods significantly when the land covers are complex and with irregular distributions. Guochao Zhang, Xueting Hu, Yantao Wei, Weijia Cao, Huang Yao, Xueyang Zhang, Keyi Song |
IEEE Geosci. Remote. Sens. Lett. | 6 |
| 2021 | Supplier selection with different risk preferences and attribute sets: An innovative study based on generalized linguistic term sets
Jianxin You, Xueyang Zhang |
Adv. Eng. Informatics | 3 |
| 2020 | Progressive Multi-Target Network Based Speech Enhancement with Snr-Preselection for Robust Speaker DiarizationabstractIn this paper, we design a novel front-end processing system for speaker diarization under realistic conditions with challenging background noises. To cope with diversified environments, we first extend our perviously proposed progressive learning based speech enhancement model by adding multi-task learning in each intermediate layer. The corresponding progressive multi-target (PMT) in various layers includes both progressive ratio mask (PRM) and progressively enhanced log-power spectra (PELPS) with specified signal-to-noise ratios (SNRs). Speech distortions are commonly introduced during the front-end processing, which often deteriorate the back-end performance. However, the proposed speech enhancement model can be regarded as a bagging of models with multiple learning objectives, which provides flexibility for selecting the most appropriate output for robust speaker diarzation. In addition, a global SNR estimation is performed using the results of deep neural network (DNN) based speech activity detection (SAD) to decide whether the audio should be enhanced. We evaluate the speaker diarzation performance on the second DIHARD dataset which includes several different realistic conditions. Compared with the original data, experiments demonstrate that the enhanced data processed by our proposed method can effectively avoid the performance loss of every single domain, and achieve consistent improvements in most domains. Lei Sun 0010, Jun Du 0002, Xueyang Zhang, Tian Gao 0005, Chin-Hui Lee 0001 |
ICASSP | 3 |
| 2018 | Speaker Diarization with Enhancing Speech for the First DIHARD Challenge
Lei Sun 0010, Jun Du 0002, Xueyang Zhang, Chin-Hui Lee 0001 |
INTERSPEECH | 4 |