VLDB 2026 Research / reviewers in the wild / expert
Yikang Ding
dblp:307/5268
· DBLP profile ↗
19ranked-venue papers
5as first author
19since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 4 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 4 first-author · 14 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | UniScene: Unified Occupancy-centric Driving Scene GenerationabstractGenerating high-fidelity, controllable, and annotated training data is critical for autonomous driving. Existing methods typically generate a single data form directly from a coarse scene layout, which not only fails to output rich data forms required for diverse downstream tasks but also struggles to model the direct layout-to-data distribution. In this paper, we introduce UniScene, the first unified framework for generating three key data forms — semantic occupancy, video, and LiDAR — in driving scenes. UniScene employs a progressive generation process that decomposes the complex task of scene generation into two hierarchical steps: (a) first generating semantic occupancy from a customized scene layout as a meta scene representation rich in both semantic and geometric information, and then (b) conditioned on occupancy, generating video and LiDAR data, respectively, with two novel transfer strategies of Gaussian-based Joint Rendering and Prior-guided Sparse Modeling. This occupancy-centric approach reduces the generation burden, especially for intricate scenes, while providing detailed intermediate representations for the subsequent generation stages. Extensive experiments demonstrate that UniScene outperforms previous SOTAs in the occupancy, video, and LiDAR generation, which also indeed benefits downstream driving tasks. The Project is available at https://arlo0o.github.io/uniscene/. Bohan Li 0015, Jiazhe Guo, Hongsi Liu, Yingshuang Zou, Yikang Ding, Xiwu Chen, Hu Zhu, Feiyang Tan, Tiancai Wang, Shuchang Zhou 0001, Li Zhang 0040, Xiaojuan Qi 0001, Hao Zhao 0002, Mu Yang, Wenjun Zeng 0001, Xin Jin 0014 |
CVPR | 5 |
| 2025 | DiST-4D: Disentangled Spatiotemporal Diffusion with Metric Depth for 4D Driving Scene GenerationabstractCurrent generative models struggle to synthesize dynamic 4D driving scenes that simultaneously support temporal extrapolation and spatial novel view synthesis (NVS) without per-scene optimization. A key challenge lies in finding an efficient and generalizable geometric representation that seamlessly connects temporal and spatial synthesis. To address this, we propose DiST-4D, the first disentangled spatiotemporal diffusion framework for 4D driving scene generation, which leverages metric depth as the core geometric representation. DiST-4D decomposes the problem into two diffusion processes: DiST-T, which predicts future metric depth and multi-view RGB sequences directly from past observations, and DiST-S, which enables spatial NVS by training only on existing viewpoints while enforcing cycle consistency. This cycle consistency mechanism introduces a forward-backward rendering constraint, reducing the generalization gap between observed and unseen viewpoints. Metric depth is essential for both accurate reliable forecasting and accurate spatial NVS, as it provides a view-consistent geometric representation that generalizes well to unseen perspectives. Experiments demonstrate that DiST-4D achieves state-of-the-art performance in both temporal prediction and NVS tasks, while also delivering competitive performance in planning-related evaluations. Jiazhe Guo, Yikang Ding, Xiwu Chen, Bohan Li 0015, Yingshuang Zou, Xiaoyang Lyu, Feiyang Tan, Xiaojuan Qi 0001, Hao Zhao 0002 |
ICCV | 2 |
| 2025 | HERMES: A Unified Self-Driving World Model for Simultaneous 3D Scene Understanding and GenerationabstractDriving World Models (DWMs) have become essential for autonomous driving by enabling future scene prediction. However, existing DWMs are limited to scene generation and fail to incorporate scene understanding, which involves interpreting and reasoning about the driving environment. In this paper, we present a unified Driving World Model named HERMES. We seamlessly integrate 3D scene understanding and future scene evolution (generation) through a unified framework in driving scenarios. Specifically, HERMES leverages a Bird's-Eye View (BEV) representation to consolidate multi-view spatial information while preserving geometric relationships and interactions. We also introduce world queries, which incorporate world knowledge into BEV features via causal attention in the Large Language Model, enabling contextual enrichment for understanding and generation tasks. We conduct comprehensive studies on nuScenes and OmniDrive-nuScenes datasets to validate the effectiveness of our method. HERMES achieves state-of-the-art performance, reducing generation error by 32.4% and improving understanding metrics such as CIDEr by 8.0%. The model and code will be publicly released at https://github.com/LMD0311/HERMES. Xin Zhou 0013, Dingkang Liang, Sifan Tu, Xiwu Chen, Yikang Ding, Dingyuan Zhang, Feiyang Tan, Hengshuang Zhao, Xiang Bai |
ICCV | 5 |
| 2025 | CLAIM: Camera-LiDAR Alignment with Intensity and MonodepthabstractIn this paper, we unleash the potential of the powerful monodepth model in camera-LiDAR calibration and propose CLAIM, a novel method of aligning data from the camera and LiDAR. Given the initial guess and pairs of images and LiDAR point clouds, CLAIM utilizes a coarse-to-fine searching method to find the optimal transformation minimizing a patched Pearson correlation-based structure loss and a mutual information-based texture loss. These two losses serve as good metrics for camera-LiDAR alignment results and require no complicated steps of data processing, feature extraction, or feature matching like most methods, rendering our method simple and adaptive to most scenes. We validate CLAIM on public KITTI, Waymo, and MIAS-LCEC datasets, and the experimental results demonstrate its superior performance compared with the state-of-the-art methods. The code is available at https://github.com/Tompson11/claim. Meijie Zhang, Feiyang Tan, Yikang Ding |
IROS | 5 |
| 2025 | Joint multi-layer network and coupling redundancy minimization for semi-supervised EEG-based emotion recognition
Liangliang Hu, Daowen Xiong, Congming Tan, Yikang Ding, Jiahao Jin, Yin Tian |
Knowl. Based Syst. | 5 |
| 2025 | Interpretable Cross-Modal Alignment Network for EEG Visual Decoding With Algorithm UnrollingabstractAccurate decoding in electroencephalography (EEG) technology, particularly for rapid visual stimuli, remains challenging due to the low signal-to-noise ratio (SNR). Additionally, existing neural networks struggle with issues related to generalization and interpretability. This article proposes a cross-modal aligned network, E2IVAE, which leverages shared information from multiple modalities for self-supervised alignment of EEG to images for extracting visual perceptual information and features a novel EEG encoder, ISTANet, based on algorithm unrolling. This network framework significantly enhances the accuracy and stability of EEG decoding for object recognition in novel classes while reducing the extensive neural data typically required for training neural decoders. The proposed ISTANet employs algorithm unrolling to transform the multilayer sparse coding algorithm into an end-to-end format, extracting features from noisy EEG signals while incorporating the interpretability of traditional machine learning. The experimental results demonstrate that our method achieves SOTA top-1 accuracy of 62.39% and top-5 accuracy of 88.98% on a comprehensive rapid serial visual presentation (RSVP) dataset for public comparison in a 200-class zero-shot neural decoding task. Additionally, ISTANet enables visualization and analysis of multiscale atom features and overall reconstruction features, exploring biological plausibility across temporal, spatial, and spectral dimensions. On another more challenging RSVP large-scale dataset, the proposed framework also achieves significantly above chance-level performance, proving its robustness and generalization. This research provides critical insights into neural decoding and brain-computer interfaces (BCIs) within the fields of cognitive science and artificial intelligence. Daowen Xiong, Liangliang Hu, Jiahao Jin, Yikang Ding, Congming Tan, Yin Tian |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | M2Depth: Self-supervised Two-Frame Multi-camera Metric Depth Estimation
Yingshuang Zou, Yikang Ding, Xi Qiu, Haoqian Wang |
ECCV (46) | 2 |
| 2023 | Adaptive Assignment for Geometry Aware Local Feature MatchingabstractThe detector-free feature matching approaches are currently attracting great attention thanks to their excellent performance. However, these methods still struggle at large-scale and viewpoint variations, due to the geometric inconsistency resulting from the application of the mutual nearest neighbour criterion (i.e., one-to-one assignment) in patch-level matching. Accordingly, we in-troduce AdaMatcher, which first accomplishes the feature correlation and co-visible area estimation through an elaborate feature interaction module, then performs adaptive assignment on patch-level matching while es-timating the scales between images, and finally refines the co-visible matches through scale alignment and sub-pixel regression module. Extensive experiments show that AdaMatcher outperforms solid baselines and achieves state-of-the-art results on many downstream tasks. Ad-ditionally, the adaptive assignment and sub-pixel refinement module can be used as a refinement network for other matching methods, such as SuperGlue, to boost their performance further. The code will be publicly available at https://github.com/AbyssGaze/AdaMatcher. Dihe Huang, Yong Liu 0032, Shang Xu, Yikang Ding, Fan Tang, Chengjie Wang 0001 |
CVPR | 7 |
| 2023 | Rethinking Feature Context in Learning Image-Guided Depth Completion
Yikang Ding, Pengzhi Li, Dihe Huang, Zhiheng Li 0001 |
ICANN (3) | 1 |
| 2023 | Volumetric 3D Reconstruction with Window-Wise Global Feature AggregationabstractVolumetric 3D reconstruction methods have shown great performance in reconstructing indoor scenarios from monocular videos. However, as such approaches utilize discrete feature voxels to encode the observed scenes, the global feature interaction within and across different voxels is ignored, leading to imperfect reconstructions. To solve this problem, we propose a novel volumetric 3D reconstruction method named VolGARecon. The core portion of VolGARecon includes two parts: first, we use an MLP-based weighted fusion module (WFM) to unproject the extracted features to each voxel, which considers the visibility and is capable to reduce the noise caused by occlusion; second, a 3D transformer module (3DTR) is used to perform window-wise global feature interaction in a local sliding window, which strengthens the feature expression in 3D space and benefits estimating more complete and spatially coherent 3D models. In addition, we propose a multi-dimensional hybrid loss (MHL) that incorporates the 3D supervision in classical volumetric methods and the 2D supervision in novel view synthesis works. Extensive experiments show our method achieves superior performance on multiple datasets. Shihao Ren, Yikang Ding, Jinli Liao, Xinghui Li, Wensen Feng, Xueqian Wang 0001 |
ICASSP | 2 |
| 2023 | Edge-aware Neural Implicit Surface ReconstructionabstractRecently, neural implicit 3D reconstruction in indoor scenarios has achieved impressive performance. Utilizing the volume rendering method and neural implicit representation to learn 3D scenes, such per-scene optimization methods could reconstruct pretty complete models but also suffer from missing details and overly-smoothed reconstructions. In this paper, we propose a novel edge-aware neural implicit surface reconstruction method, named Ea-NeuS, to learn high-quality 3D models with fine details. Specifically, we use the edge of objects to locate the important areas, and propose a simple yet effective edge-guided ray-sampling strategy to learn the 3D models. The aforementioned edge information further guides the normal prior supervision, which helps reduce inaccurate optimization in detailed regions. We additionally use the visibility-aware sparse points to pilot the 3D points sampling along the rays and perform explicit supervision. As a result, our method achieves superior performance compared with existing methods on various scenes. Xinghui Li, Yikang Ding, Xiansong Lai, Shihao Ren, Wensen Feng, Long Zeng 0001 |
ICME | 2 |
| 2023 | Towards Practical Consistent Video Depth EstimationabstractMonocular depth estimation algorithms aim to explore the possible links between 2D and 3D data, but challenges remain for existing methods to predict consistent depth from a casual video. Relying on camera poses and the optical flow in the time-consuming test-time training phases makes these methods fail in many scenarios and cannot be used for practical applications. In this work, we present a data-driven post-processing method to overcome these challenges and achieve online processing. Based on a deep recurrent network, our method takes the adjacent original and optimized depth map as inputs to learn temporal consistency from the dataset and achieves higher depth accuracy. Our approach can be applied to multiple single-frame depth estimation models and used for various real-world scenes in real-time. In addition, to tackle the lack of a temporally consistent video depth training dataset of dynamic scenes, we propose an approach to generate the training video sequences dataset from a single image based on inferring motion field. To the best of our knowledge, this is the first data-driven plug-and-play method to improve the temporal consistency of depth estimation for casual videos. Extensive experiments on three datasets and three depth estimation models show that our method outperforms the state-of-the-art methods. Pengzhi Li, Yikang Ding, Linge Li, Jingwei Guan, Zhiheng Li 0001 |
ICMR | 2 |
| 2023 | Sem-Avatar: Semantic Controlled Neural Field for High-Fidelity Audio Driven Avatar
Yikang Ding |
PRCV (2) | 3 |
| 2022 | Adaptive Range Guided Multi-view Depth Estimation with Normal Ranking Loss
Yikang Ding, Dihe Huang, Kai Zhang 0012, Zhiheng Li 0001, Wensen Feng |
ACCV (1) | 1 |
| 2022 | TransMVSNet: Global Context-aware Multi-view Stereo Network with TransformersabstractIn this paper, we present TransMVSNet, based on our exploration of feature matching in multi-view stereo (MVS). We analogize MVS back to its nature of a feature matching task and therefore propose a powerful Feature Matching Transformer (FMT) to leverage intra- (self-) and inter-(cross-) attention to aggregate long-range context information within and across images. To facilitate a better adaptation of the FMT, we leverage an Adaptive Receptive Field (ARF) module to ensure a smooth transit in scopes of features and bridge different stages with a feature pathway to pass transformed features and gradients across different scales. In addition, we apply pair-wise feature correlation to measure similarity between features, and adopt ambiguity-reducing focal loss to strengthen the supervision. To the best of our knowledge, TransMVSNet is the first attempt to leverage Transformer into the task of MVS. As a result, our method achieves state-of-the-art performance on DTU dataset, Tanks and Temples benchmark, and BlendedMVS dataset. Code is available at https://github.com/MegviiRobot/TransMVSNet. Yikang Ding, Qingtian Zhu, Xiangyue Liu 0004, Yuanjiang Wang |
CVPR | 1 |
| 2022 | KD-MVS: Knowledge Distillation Based Self-supervised Learning for Multi-view Stereo
Yikang Ding, Qingtian Zhu, Xiangyue Liu 0004 |
ECCV (31) | 1 |
| 2022 | Sobolev Training for Implicit Neural Representations with Approximated Image Derivatives
Qingtian Zhu, Xiangyue Liu 0004, Yikang Ding |
ECCV (15) | 4 |
| 2022 | Enhancing Multi-View Stereo with Contrastive Matching and Weighted Focal LossabstractLearning-based multi-view stereo (MVS) methods have made impressive progress and surpassed traditional methods in recent years. However, their accuracy and completeness are still struggling. In this paper, we propose a new method to enhance the performance of existing networks inspired by contrastive learning and feature matching. First, we propose a Contrast Matching Loss (CML), which treats the correct matching points in depth-dimension as positive sample and other points as negative samples, and computes the contrastive loss based on the similarity of features. We further propose a Weighted Focal Loss (WFL) for better classification capability, which weakens the contribution of low-confidence pixels in unimportant areas to the loss according to predicted confidence. Extensive experiments performed on DTU, Tanks and Temples and BlendedMVS datasets show our method achieves state-of-the-art performance and significant improvement over baseline network. Yikang Ding, Dihe Huang, Zhiheng Li 0001, Kai Zhang 0012 |
ICIP | 1 |
| 2022 | WT-MVSNet: Window-based Transformers for Multi-view StereoabstractRecently, Transformers have been shown to enhance the performance of multi-view stereo by enabling long-range feature interaction. In this work, we propose Window-based Transformers (WT) for local feature matching and global feature aggregation in multi-view stereo. We introduce a Window-based Epipolar Transformer (WET) which reduces matching redundancy by using epipolar constraints. Since point-to-line matching is sensitive to erroneous camera pose and calibration, we match windows near the epipolar lines. A second Shifted WT is employed for aggregating global information within cost volume. We present a novel Cost Transformer (CT) to replace 3D convolutions for cost volume regularization. In order to better constrain the estimated depth maps from multiple views, we further design a novel geometric consistency loss (Geo Loss) which punishes unreliable areas where multi-view consistency is not satisfied. Our WT multi-view stereo method (WT-MVSNet) achieves state-of-the-art performance across multiple datasets and ranks $1^{st}$ on Tanks and Temples benchmark. Code will be available upon acceptance. Jinli Liao, Yikang Ding, Yoli Shavit, Dihe Huang, Shihao Ren, Wensen Feng, Kai Zhang 0012 |
NeurIPS | 2 |