EDBT 2026 Demo / reviewers in the wild / expert
Yuhang Ming 0001
dblp:299/1425-1
· DBLP profile ↗
12ranked-venue papers
4as first author
12since 2021 · last 2026
0000-0002-4548-6388ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 3 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 6 since 2021Systems, architecture and hardware · 4 · 2 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | D3FlowSLAM: Self-supervised dynamic SLAM with flow motion decomposition and DINO guidance
Xingyuan Yu, Weicai Ye, Xiyue Guo, Yuhang Ming 0001, Jinyu Li 0002, Hujun Bao, Zhaopeng Cui, Guofeng Zhang 0001 |
Neurocomputing | 4 |
| 2025 | Multi-Modal Synergistic Implicit Image Enhancement for Efficient Optical Flow EstimationabstractAs a fundamental visual task, optical flow estimation has widespread applications in computer vision. However, it faces significant challenges under adverse lighting conditions, where low texture and noise make accurate optical flow estimation particularly difficult. In this paper, we propose an optical flow method that employs implicit image enhancement through multi-modal synergistic training. To supplement the scene information missing in the original low-quality image, we utilize a high-low frequency feature enhancement network. The enhancement network is implicitly guided by multi-modal data and the specific subsequent tasks, enabling the model to learn multi-modal knowledge that enhances feature information suitable for optical flow estimation during inference. By using RGBD multi-modal data, the proposed method avoids the reliance on the images captured from the same view, a common limitation in traditional image enhancement methods. During training, the encoded features extracted from the enhanced images are synergistically supervised by features from the RGBD fusion as well as by the optical flow task. Experiments conducted on both synthetic and real datasets demonstrate that the proposed method significantly improves performance on public datasets. Weichen Dai 0001, Hexing Wu, Xiaoyang Weng, Yuhang Ming 0001, Wanzeng Kong |
CVPR | 5 |
| 2025 | Benchmarking neural radiance fields for autonomous robots: An overview
Yuhang Ming 0001, Xingrui Yang 0001, Zheng Chen 0016, Jinglun Feng, Yifan Xing, Guofeng Zhang 0001 |
Eng. Appl. Artif. Intell. | 1 |
| 2025 | Independent Components Time-Frequency Purification With Channel Consensus Against Adversarial Attack in SSVEP-Based BCIsabstractThe Steady State Visual Evoked Potential (SSVEP) paradigm has been widely employed in various Brain-Computer Interface (BCI) systems. However, recent studies indicate that SSVEP is vulnerable to adversarial attacks, resulting in manipulated results and drastic degradation in recognition performance, which pose inconveniences and even risks to users. Noticing the fact that the adversarial attack on SSVEP is done by adding subtle waveform perturbations into random EEG channels, we propose Independent Components Time-Frequency Purification with Channel Consensus (ICTFP-CC) as a defensive strategy. In particular, we first detect and remove suspicious perturbations with independent component analysis from the time and frequency domain, and then reconstruct the purified EEG signals. Additionally, we introduce a voting mechanism to achieve channel consensus and enhance overall robustness. We conducted experiments on two public datasets and three SSVEP recognition algorithms. The results demonstrate that our method can significantly improve the classification accuracy and information transfer rate of attacked SSVEP signals by a maximum of 46.79 (%) and 62.87 (bits/min). Hangjie Yi, Jingsheng Qian, Yuhang Ming 0001, Wanzeng Kong |
IEEE Signal Process. Lett. | 3 |
| 2024 | AEGIS-Net: Attention-Guided Multi-Level Feature Aggregation for Indoor Place RecognitionabstractWe present AEGIS-Net, a novel indoor place recognition model that takes in RGB point clouds and generates global place descriptors by aggregating lower-level color, geometry features and higher-level implicit semantic features. However, rather than simple feature concatenation, self-attention modules are employed to select the most important local features that best describe an indoor place. Our AEGIS-Net is made of a semantic encoder, a semantic decoder and an attention-guided feature embedding. The model is trained in a 2-stage process with the first stage focusing on an auxiliary semantic segmentation task and the second one on the place recognition task. We evaluate our AEGIS-Net on the ScanNetPR dataset and compare its performance with a pre-deep-learning feature-based method and five state-of-the-art deep-learning-based methods. Our AEGIS-Net achieves exceptional performance and outperforms all six methods. Yuhang Ming 0001, Jian Ma 0001, Xingrui Yang 0001, Weichen Dai 0001, Yong Peng 0001, Wanzeng Kong |
ICASSP | 1 |
| 2024 | Time-Frequency Jointed Imperceptible Adversarial Attack to Brainprint Recognition with Deep Learning ModelsabstractEEG-based brainprint recognition with deep learning models has garnered much attention in biometric identification. Yet, studies have indicated vulnerability to adversarial attacks in deep learning models with EEG inputs. In this paper, we introduce a novel adversarial attack method that jointly attacks time-domain and frequency-domain EEG signals by employing wavelet transform. Different from most existing methods which only target time-domain EEG signals, our method not only takes advantage of the time-domain attack’s potent adversarial strength but also benefits from the imperceptibility inherent in frequency-domain attack, achieving a better balance between attack performance and imperceptibility. Extensive experiments are conducted in both white- and grey-box scenarios and the results demonstrate that our attack method achieves state-of-the-art attack performance on three datasets and three deep-learning models. In the meanwhile, the perturbations in the signals attacked by our method are barely perceptible to the human visual system. Hangjie Yi, Yuhang Ming 0001, Dongjun Liu, Wanzeng Kong |
ICME | 2 |
| 2023 | PVO: Panoptic Visual OdometryabstractWe present PVO, a novel panoptic visual odometry framework to achieve more comprehensive modeling of the scene motion, geometry, and panoptic segmentation information. Our PVO models visual odometry (VO) and video panoptic segmentation (VPS) in a unified view, which makes the two tasks mutually beneficial. Specifically, we introduce a panoptic update module into the VO Module with the guidance of image panoptic segmentation. This Panoptic-Enhanced VO Module can alleviate the impact of dynamic objects in the camera pose estimation with a panoptic-aware dynamic mask. On the other hand, the VO-Enhanced VPS Module also improves the segmentation accuracy by fusing the panoptic segmentation result of the current frame on the fly to the adjacent frames, using geometric information such as camera pose, depth, and optical flow obtained from the VO Module. These two modules contribute to each other through recurrent iterative optimization. Extensive experiments demonstrate that PVO outperforms state-of-the-art methods in both visual odometry and video panoptic segmentation tasks. Weicai Ye, Xinyue Lan, Yuhang Ming 0001, Xingyuan Yu, Hujun Bao, Zhaopeng Cui, Guofeng Zhang 0001 |
CVPR | 4 |
| 2023 | EDI: ESKF-based Disjoint Initialization for Visual-Inertial SLAM SystemsabstractVisual-inertial initialization can be classified into joint and disjoint approaches. Joint approaches tackle both the visual and the inertial parameters together by aligning observations from feature-bearing points based on IMU integration then use a closed-form solution with visual and acceleration observations to find initial velocity and gravity. In contrast, disjoint approaches independently solve the Structure from Motion (SFM) problem and determine inertial parameters from up-to-scale camera poses obtained from pure monocular SLAM. However, previous disjoint methods have limitations, like assuming negligible acceleration bias impact or accurate rotation estimation by pure monocular SLAM. To address these issues, we propose EDI, a novel approach for fast, accurate, and robust visual-inertial initialization. Our method incorporates an Error-state Kalman Filter (ESKF) to estimate gyroscope bias and correct rotation estimates from monocular SLAM, overcoming dependence on pure monocular SLAM for rotation estimation. To estimate the scale factor without prior information, we offer a closed-form solution for initial velocity, scale, gravity, and acceleration bias estimation. To address gravity and acceleration bias coupling, we introduce weights in the linear least-squares equations, ensuring acceleration bias observability and handling outliers. Extensive evaluation on the EuRoC dataset shows that our method achieves an average scale error of 5.8% in less than 3 seconds, outperforming other state-of-the-art disjoint visual-inertial initialization approaches, even in challenging environments and with artificial noise corruption. Yuhang Ming 0001, Philippos Mordohai |
IROS | 3 |
| 2022 | FD-SLAM: 3-D Reconstruction Using Features and Dense MatchingabstractIt is well known that visual SLAM systems based on dense matching are locally accurate but are also susceptible to long-term drift and map corruption. In contrast, feature matching methods can achieve greater long-term consistency but can suffer from inaccurate local pose estimation when feature information is sparse. Based on these observations, we propose an RGB-D SLAM system that leverages the advantages of both approaches: using dense frame-to-model odometry to build accurate sub-maps and on-the-fly feature-based matching across sub-maps for global map optimisation. In addition, we incorporate a learning-based loop closure component based on 3-D features which further stabilises map building. We have evaluated the approach on indoor sequences from public datasets, and the results show that it performs on par or better than state-of-the-art systems in terms of map reconstruction quality and pose estimation. The approach can also scale to large scenes where other systems often fail. Xingrui Yang 0001, Yuhang Ming 0001, Zhaopeng Cui, Andrew Calway |
ICRA | 2 |
| 2022 | CGiS-Net: Aggregating Colour, Geometry and Implicit Semantic Features for Indoor Place RecognitionabstractWe describe a novel approach to indoor place recognition from RGB point clouds based on aggregating low-level colour and geometry features with high-level implicit semantic features. It uses a 2-stage deep learning framework, in which the first stage is trained for the auxiliary task of semantic segmentation and the second stage uses features from layers in the first stage to generate discriminate descriptors for place recognition. The auxiliary task encourages the features to be semantically meaningful, hence aggregating the geometry and colour in the RGB point cloud data with implicit semantic information. We use an indoor place recognition dataset derived from the ScanNet dataset for training and evaluation, with a test set comprising 3,608 point clouds generated from 100 different rooms. Comparison with a traditional feature-based method and four state-of-the-art deep learning methods demonstrate that our approach significantly outperforms all five methods, achieving, for example, a top-3 average recall rate of 75% compared with 41% for the closest rival method. Our code is available at: https://github.com/YuhangMing/Semantic-Indoor-Place-Recognition Yuhang Ming 0001, Xingrui Yang 0001, Guofeng Zhang 0001, Andrew Calway |
IROS | 1 |
| 2022 | Vox-Fusion: Dense Tracking and Mapping with Voxel-based Neural Implicit RepresentationabstractIn this work, we present a dense tracking and mapping system named Vox-Fusion, which seamlessly fuses neural implicit representations with traditional volumetric fusion methods. Our approach is inspired by the recently developed implicit mapping and positioning system and further extends the idea so that it can be freely applied to practical scenarios. Specifically, we leverage a voxel-based neural implicit surface representation to encode and optimize the scene inside each voxel. Furthermore, we adopt an octree-based structure to divide the scene and support dynamic expansion, enabling our system to track and map arbitrary scenes without knowing the environment like in previous works. Moreover, we proposed a high-performance multi-process framework to speed up the method, thus supporting some applications that require real-time performance. The evaluation results show that our methods can achieve better accuracy and completeness than previous methods. We also show that our Vox-Fusion can be used in augmented reality and virtual reality applications. Our source code is publicly available at https://github.com/zju3dv/Vox-Fusion. Xingrui Yang 0001, Hongjia Zhai, Yuhang Ming 0001, Yuqian Liu, Guofeng Zhang 0001 |
ISMAR | 4 |
| 2021 | Object-Augmented RGB-D SLAM for Wide-Disparity RelocalisationabstractWe propose a novel object-augmented RGB-D SLAM system that is capable of constructing a consistent object map and performing relocalisation based on centroids of objects in the map. The approach aims to overcome the view dependence of appearance-based relocalisation methods using point features or images. During the map construction, we use a pre-trained neural network to detect objects and estimate 6D poses from RGB-D data. An incremental probabilistic model is used to aggregate estimates over time to create the object map. Then in relocalisation, we use the same network to extract objects-of-interest in the ‘lost’ frames. Pairwise geometric matching finds correspondences between map and frame objects, and probabilistic absolute orientation followed by application of iterative closest point to dense depth maps and object centroids gives relocalisation. Results of experiments in desktop environments demonstrate very high success rates even for frames with widely different viewpoints from those used to construct the map, significantly outperforming two appearance- based methods. Yuhang Ming 0001, Xingrui Yang 0001, Andrew Calway |
IROS | 1 |