EDBT 2026 Demo / reviewers in the wild / expert
Zhimin Yuan
dblp:72/1060
· DBLP profile ↗
19ranked-venue papers
4as first author
16since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 8 · 2 first-author · 7 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 5 since 2021Computer networks · 2 · 1 since 2021Systems, architecture and hardware · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Adaptive Co-Operative Prompting and Uncertainty-Aware Implicit Knowledge Enhancement for Cross-Modal RetrievalabstractWith the rapid growth of Internet multimedia data, cross-modal retrieval techniques have garnered significant attention. Given the inherent complexity and non-intuitive nature of cross-modal relationships, tuning pre-trained Large Multimodal Models (LMMs) with cross-modal data has become a mainstream approach. However, cross-modal data commonly exhibit inter-modal information asymmetry and intra-modal distribution diversity. Faced with these challenges, existing paradigms tend to learn ambiguous and asymmetric cross-modal associations, which introduce semantic noise. In addition, their limited adaptability to the high diversity of real-world content further hinders optimal retrieval performance. To address these challenges, this article proposes the A daptive C o-operative K nowledge E nhancement (ACKE) method, which comprises the Uncertainty-Aware Inspire Potential (UAIP) and Adaptive Co-Operative Prompt (ACP) strategies. UAIP utilizes generative LMMs to generate multi-perspective descriptions that enrich semantic information, while employing Dempster-Shafer Theory (DST) to quantify their semantic uncertainty and adjust contribution weights, reducing inaccurate relational mappings and balancing information asymmetry. ACP constructs a prompt pool where instance-specific visual prompts are dynamically selected and projected into text prompts, which collaborate to guide modal encoders toward deep semantic consensus, thus mitigating alignment bias from intra-modal distribution diversity and improving accuracy. Extensive experiments are conducted on two widely used datasets, Flickr30K and MS-COCO, demonstrating the effectiveness of our proposed method. The code is available at https://github.com/nynu-BDAI/ACKE . Zhimin Yuan, Runsheng Zhang |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2025 | SSRFlow: Semantic-Aware Fusion with Spatial Temporal Re-Embedding for Real-World Scene FlowabstractScene flow, which provides the 3D motion field of the first frame from two consecutive point clouds, is vital for dynamic scene perception. However, contemporary scene flow methods face three major challenges. Firstly, they only consider the context of individual point clouds before flow embedding, leading to embedded points struggling to perceive the consistent semantic relationship of another frame. To address this issue, we propose a novel approach called Dual Cross Attentive (DCA) for the latent fusion and alignment between two frames based on semantic contexts. This is then integrated into Global Fusion Flow Embedding (GF) to initialize flow embedding based on global correlations in both contextual and Euclidean spaces. Secondly, deformations exist in non-rigid objects after the warping layer, which distorts the spatiotemporal relation between the consecutive frames. For a more precise estimation of residual flow at next-level, the Spatial Temporal Re-embedding (STR) module is devised to update the point sequence features at current-level. Lastly, poor generalization is often observed due to the significant domain gap between synthetic and LiDAR-scanned datasets. We leverage novel domain adaptive losses to effectively bridge the gap of motion inference from synthetic to real-world. Experiments demonstrate that our approach achieves state-of-the-art (SOTA) performance across various datasets, with particularly outstanding results in real-world LiDAR-scanned situations. Zhiyang Lu, Qinghan Chen, Zhimin Yuan, Chenglu Wen, Ming Cheng 0002, Cheng Wang 0003 |
3DV | 3 |
| 2025 | STGC-NeRF: Spatial-Temporal Geometric Consistency for LiDAR Neural Radiance Fields in Dynamic ScenesabstractWhile Neural Radiance Fields (NeRFs) have advanced the frontiers of novel view synthesis (NVS) using LiDAR data, they still struggle in dynamic scenes. Due to the low frequency and sparsity characteristics of LiDAR point clouds, it is challenging to spontaneously learn a dynamic and consistent scene representation from posed scans. In this paper, we propose STGC-NeRF, a novel LiDAR NeRF method that combines spatial-temporal geometry consistency to enhance the reconstruction of dynamic scenes. First, we propose a temporal geometry consistency regularization to enhance the regression of time-varying scene geometries from low-frequency LiDAR sequences. By estimating the pointwise correspondences between synthetic (or real) and real frames at different times, we convert them into various forms of temporal supervision. This alleviates the inconsistency caused by moving objects in dynamic scenes. Second, to improve the reconstruction of sparse LiDAR data, we propose spatial geometric consistency constraints. By computing multiple neighborhood feature descriptors incorporating geometric and contextual information, we capture structural geometry information from sparse LiDAR data. This helps encourage consistent direction, smoothness, and detail of the local surface. Extensive experiments on the KITTI-360 and nuScenes datasets demonstrate that STGC-NeRF outperforms state-of-the-art methods in both geometry and intensity accuracy for dynamic LiDAR scene reconstruction. Shangshu Yu, Xiaotian Sun 0005, Wen Li 0005, Qingshan Xu 0001, Zhimin Yuan, Rui She 0001, Cheng Wang 0003 |
AAAI | 5 |
| 2025 | GTR-Loc: Geospatial Text Regularization Assisted Outdoor LiDAR LocalizationabstractPrevailing scene coordinate regression methods for LiDAR localization suffer from localization ambiguities, as distinct locations can exhibit similar geometric signatures — a challenge that current geometry-based regression approaches have yet to solve. Recent vision–language models show that textual descriptions can enrich scene understanding, supplying potential localization cues missing from point cloud geometries. In this paper, we propose GTR-Loc, a novel text-assisted LiDAR localization framework that effectively generates and integrates geospatial text regularization to enhance localization accuracy. We propose two novel designs: a Geospatial Text Generator that produces discrete pose-aware text descriptions, and a LiDAR-Anchored Text Embedding Refinement module that dynamically constructs view-specific embeddings conditioned on current LiDAR features. The geospatial text embeddings act as regularization to effectively reduce localization ambiguities. Furthermore, we introduce a Modality Reduction Distillation strategy to transfer textual knowledge. It enables high-performance LiDAR-only localization during inference, without requiring runtime text generation. Extensive experiments on challenging large-scale outdoor datasets, including QEOxford, Oxford Radar RobotCar, and NCLT, demonstrate the effectiveness of GTR-Loc. Our method significantly outperforms state-of-the-art approaches, notably achieving a 9.64%/8.04% improvement in position/orientation accuracy on QEOxford. Our code is available at https://github.com/PSYZ1234/GTR-Loc. Shangshu Yu, Wen Li 0005, Xiaotian Sun 0005, Zhimin Yuan, Rui She 0001, Cheng Wang 0003 |
NeurIPS | 4 |
| 2024 | Density-guided Translator Boosts Synthetic-to-Real Unsupervised Domain Adaptive Segmentation of 3D Point Cloudsabstract3D synthetic-to-real unsupervised domain adaptive seg-mentation is crucial to annotating new domains. Self-training is a competitive approach for this task, but its performance is limited by different sensor sampling patterns (i.e., variations in point density) and incomplete training strate-gies. In this work, we propose a density-guided translator (DGT), which translates point density between domains, and integrates it into a two-stage self-training pipeline named DGT-ST. First, in contrast to existing works that simulta-neously conduct data generation and feature/output align-ment within unstable adversarial training, we employ the non-learnable DGT to bridge the domain gap at the in-put level. Second, to provide a well-initialized model for self-training, we propose a category-level adversarial net-work in stage one that utilizes the prototype to prevent neg-ative transfer. Finally, by leveraging the designs above, a domain-mixed self-training method with source-aware consistency loss is proposed in stage two to narrow the domain gap further. Experiments on two synthetic-to-real segmentation tasks (SynLiDAR → semanticKITTI and SynL- iDAR → semanticPOSS) demonstrate that DGT-ST outper-forms state-of-the-art methods, achieving 9.4% and 4.3% mIoU improvements, respectively. Code is available at https://github.com/yuan-zm/DGT-ST. Zhimin Yuan, Wankang Zeng, Yanfei Su, Weiquan Liu, Ming Cheng 0002, Yulan Guo, Cheng Wang 0003 |
CVPR | 1 |
| 2024 | Domain adaptive remote sensing image semantic segmentation with prototype guidance
Wankang Zeng, Ming Cheng 0002, Zhimin Yuan, Youming Wu, Weiquan Liu, Cheng Wang 0003 |
Neurocomputing | 3 |
| 2023 | Multistage Scene-Level Constraints for Large-Scale Point Cloud Weakly Supervised Semantic SegmentationabstractCompared to fully supervised 3D large-scale point cloud segmentation methods, which necessitate extensive manual point-wise annotations, weakly supervised segmentation has emerged as a popular approach for significantly reducing labeling costs while maintaining effectiveness. However, the existing methods have exhibited inferior segmentation performance and unsatisfactory generalization capabilities in some scenarios with unique structures (e.g., building facades). In this paper, we propose an effective and generalized weakly supervised semantic segmentation framework, called multi-stage scene-level constraints (MSC), to solve the above problem. To address the issue regarding inadequate labeled data, we use pseudo-labels for unlabeled data and propose an uncertainty-guided adaptive reweighting strategy to reduce the negative impact of erroneous pseudo-labeled data on the model learning process. To address the class imbalance issue, we employ multi-stage scene-level constraints (i.e., encoder, decoder, and classifier stages) to treat each class equally and improve perception ability of the model for each class. Evaluations conducted on multiple large-scale point cloud datasets collected in different scenarios, including building facades, indoor scenes, outdoor scenes, and UAV scenes, show that our MSC achieves a large gain over the existing weakly supervised methods and even surpasses some fully supervised methods. Yanfei Su, Ming Cheng 0002, Zhimin Yuan, Weiquan Liu, Wankang Zeng, Cheng Wang 0003 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | Spatial Adaptive Fusion Consistency Contrastive Constraint: Weakly Supervised Building Facade Point Cloud Semantic SegmentationabstractSemantic segmentation of building facade point clouds has diverse applications. The development of semantic segmentation methods is inextricably linked to datasets. The available building facade datasets suffer from a lack of abundant semantic categories and data completeness. To compensate for these shortcomings, we propose a new building facade dataset characterized by various categories and relatively complete 3-D building facades. In addition, most existing methods focus on fully supervised learning, which relies on manually labeling large-scale point cloud data and results in high time and labor costs. In this article, we propose an effective weakly supervised building facade segmentation approach, called spatial adaptive fusion consistency contrastive constraint (SAF-C3), to solve the above problem. We first design a multirandom point cloud augmentor as an auxiliary supervision branch to enhance the learning ability of the original network branch. Then, we present a spatial adaptive fusion (SAF) module to extract discriminative features for building facade point clouds. Finally, we propose a spatial consistency contrastive constraint to explore the contrastive property in feature space and to ensure the predictive consistency among the augmentation and original branches. The proposed method achieves a significant performance improvement against the state-of-the-art methods on two building facade point cloud datasets through extensive experiments. In particular, the performance of SAF-C3 with 1% labels significantly surpasses the baseline network with 100% labels. Yanfei Su, Ming Cheng 0002, Zhimin Yuan, Weiquan Liu, Wankang Zeng, Zhihong Zhang 0001, Cheng Wang 0003 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | Prototype-Guided Multitask Adversarial Network for Cross-Domain LiDAR Point Clouds Semantic SegmentationabstractUnsupervised domain adaptation (UDA) segmentation aims to leverage labeled source data to make accurate predictions on unlabeled target data. The key is to make the segmentation network learn domain-invariant representations. In this work, we propose a prototype-guided multitask adversarial network (PMAN) to achieve this. First, we propose an intensity-aware segmentation network (IAS-Net) that leverages the private intensity information of target data to substantially facilitate feature learning of the target domain. Second, the category-level cross-domain feature alignment strategy is introduced to flee the side effects of global feature alignment. It employs the prototype (class centroid) and includes two essential operations: 1) build an auxiliary nonparametric classifier to evaluate the semantic alignment degree of each point based on the prediction consistency between the main and auxiliary classifiers and 2) introduce two class-conditional point-to-prototype learning objectives for better alignment. One is to explicitly perform category-level feature alignment in a progressive manner, and the other aims to shape the source feature representation to be discriminative. Extensive experiments reveal that our PMAN outperforms state-of-the-art results on two benchmark datasets. Zhimin Yuan, Ming Cheng 0002, Wankang Zeng, Yanfei Su, Weiquan Liu, Shangshu Yu, Cheng Wang 0003 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2023 | Category-Level Adversaries for Outdoor LiDAR Point Clouds Cross-Domain Semantic SegmentationabstractUnsupervised domain adaptation (UDA) is a low-cost way to deal with the lack of annotations in a new domain. For outdoor point clouds in urban transportation scenes, the mismatch of sampling patterns and the transferability difference between classes make cross-domain segmentation extremely difficult. To overcome these challenges, we propose a category-level adversarial framework. Firstly, we propose a multi-scale domain conditioned block that facilitates to extract the critical low-level domain-dependent knowledge and reduce the domain gap caused by distinct LiDAR sampling patterns. Secondly, we make full use of multiple representation forms (i.e., point-based sets and voxel-based cells) and utilize the prediction consistency between the two forms to measure how well each point is semantically aligned. The model then focuses on the poorly-aligned points without affecting the well-aligned points. Experimental results on three autonomous driving point cloud datasets show that the proposed method outperforms existing methods by a large margin, especially on the low-beam to high-beam cross-domain segmentation task. Zhimin Yuan, Chenglu Wen, Ming Cheng 0002, Yanfei Su, Weiquan Liu, Shangshu Yu, Cheng Wang 0003 |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2022 | Local Fusion Attention Network for Semantic Segmentation of Building Facade Point CloudsabstractAutomatic building facade point cloud semantic segmentation is an important step in 3-D urban building reconstruction. How to correctly segment the components (e.g., windows, walls, and columns) from the building facade is still a challenging task. According to the characteristics of building facade point clouds, we introduce local fusion attention network (LFA-Net), an efficient neural network that learns LFA features from building facade point clouds, for better capturing the local neighborhood structure information of each point. The core of LFA-Net is the LFA module, which consists of three neural units: local graph attention (LGA), local aggregation attention (LAA), and fusion attention (FA). The LFA-Net is the standard encoder-decoder architecture. Experiments demonstrate that our LFA-Net outperforms the state-of-the-art methods on the large-scale building facade point cloud dataset. Yanfei Su, Weiquan Liu, Ming Cheng 0002, Zhimin Yuan, Cheng Wang 0003 |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2022 | DLA-Net: Learning dual local attention features for semantic segmentation of large-scale building facade point clouds
Yanfei Su, Weiquan Liu, Zhimin Yuan, Ming Cheng 0002, Zhihong Zhang 0001, Xuelun Shen, Cheng Wang 0003 |
Pattern Recognit. | 3 |
| 2022 | DFAN: Dual-Branch Feature Alignment Network for Domain Adaptation on Point CloudsabstractUnsupervised domain adaptation (UDA) significantly reduces the gap between the source domain and the target domain in machine learning and computer vision tasks. Most UDA approaches are applied to images and videos, and only a few methods implement domain adaptation on 3-D computer vision problems. The existing UDA approaches operating on point clouds try to extract domain-invariant features in different domains for feature alignment. However, higher commonality brings less diversity and results a loss of detailed information. In this article, we propose a novel dual-branch feature alignment network (DFAN) architecture for domain adaptation on point cloud visual tasks to better exploit the respective characteristics of local and global features. Our approach specializes in the extraction and alignment of global and local features with different strategies in each branch to complement each other. We also introduce a hierarchical alignment strategy for local feature alignment and a distribution alignment strategy for global feature alignment. Experiments on the PointDA-10 and PointSegDA datasets show that our approach achieves state-of-the-art performance on the UDA of point cloud classification and segmentation tasks. The ablation study demonstrates the effectiveness of the dual-branch design and the feature alignment strategies. Liangwei Shi, Zhimin Yuan, Ming Cheng 0002, Yiping Chen 0002, Cheng Wang 0003 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2021 | Learning Cross-Domain Descriptors for 2D-3D Matching with Hard Triplet Loss and Spatial Transformer Network
Baiqi Lai, Weiquan Liu, Cheng Wang 0003, Xuesheng Bian, Yanfei Su, Xiuhong Lin, Zhimin Yuan, Ming Cheng 0002 |
ICIG (3) | 7 |
| 2021 | PP-PG: Combining Parameter Perturbation with Policy Gradient Methods for Effective and Efficient Explorations in Deep Reinforcement LearningabstractEfficient and stable exploration remains a key challenge for deep reinforcement learning (DRL) operating in high-dimensional action and state spaces. Recently, a more promising approach by combining the exploration in the action space with the exploration in the parameters space has been proposed to get the best of both methods. In this article, we propose a new iterative and close-loop framework by combining the evolutionary algorithm (EA), which does explorations in a gradient-free manner directly in the parameters space with an actor-critic, and the deep deterministic policy gradient (DDPG) reinforcement learning algorithm, which does explorations in a gradient-based manner in the action space to make these two methods cooperate in a more balanced and efficient way. In our framework, the policies represented by the EA population (the parametric perturbation part) can evolve in a guided manner by utilizing the gradient information provided by the DDPG and the policy gradient part (DDPG) is used only as a fine-tuning tool for the best individual in the EA population to improve the sample efficiency. In particular, we propose a criterion to determine the training steps required for the DDPG to ensure that useful gradient information can be generated from the EA generated samples and the DDPG and EA part can work together in a more balanced way during each generation. Furthermore, within the DDPG part, our algorithm can flexibly switch between fine-tuning the same previous RL-Actor and fine-tuning a new one generated by the EA according to different situations to further improve the efficiency. Experiments on a range of challenging continuous control benchmarks demonstrate that our algorithm outperforms related works and offers a satisfactory trade-off between stability and sample efficiency. Shilei Li, Jiongming Su, Shaofei Chen, Zhimin Yuan |
ACM Trans. Intell. Syst. Technol. | 5 |
| 2021 | HDC-Net: Hierarchical Decoupled Convolution Network for Brain Tumor SegmentationabstractAccurate segmentation of brain tumor from magnetic resonance images (MRIs) is crucial for clinical treatment decision and surgical planning. Due to the large diversity of the tumors and complex boundary interactions between sub-regions, it is of a great challenge. Besides accuracy, resource constraint is another important consideration. Recently, impressive improvement has been achieved for this task by using deep convolutional networks. However, most of state-of-the-art models rely on expensive 3D convolutions as well as model cascade/ensemble strategies, which result in high computational overheads and undesired system complexity. For clinical usage, the challenge is how to pursue the best accuracy within very limited computational budgets. In this study, we segment 3D volumetric image in one-pass with a hierarchical decoupled convolution network (HDC-Net), which is a light-weight but efficient pseudo-3D model. Specifically, we replace 3D convolutions with a novel hierarchical decoupled convolution (HDC) module, which can explore multi-scale multi-view spatial contexts with high efficiency. Extensive experiments on the BraTS 2018 and 2017 challenge datasets show that our method performs favorably against state of the art in accuracy yet with greatly reduced computational complexity. Zhengrong Luo, Zhongdao Jia, Zhimin Yuan, Jialin Peng |
IEEE J. Biomed. Health Informatics | 3 |
| 2020 | Mitochondria Segmentation From EM Images via Hierarchical Structured Contextual ForestabstractDelineation of mitochondria from electron microscopy (EM) images is crucial to investigate its morphology and distribution, which are directly linked to neural dysfunction. However, it is a challenging task due to the varied appearances, sizes and shapes of mitochondria, and complicated surrounding structures. Exploiting sufficient contextual information about interactions in extended neighborhood is crucial to address the challenges. To this end, we introduce a novel class of contextual features, namely local patch pattern (LPP), to eliminate the ambiguity of local appearance and texture features. To achieve accurate segmentation, we propose an automatic method by iterative learning of hierarchical structured contextual forest. With a novel median fusion strategy, the probability predictions from long history iterations are augmented to encode spatial and temporal contexts and suppress false detections. Moreover, the LPP features are extracted on both images and history predictions, resulting in a hierarchy of contextual features with increasing receptive fields. Other than using computationally demanding graph based methods, we perform joint label prediction using structured random forest. In addition to direct 3D segmentation of EM volumes, we introduce a 2D variant without sacrificing accuracy using a novel hierarchical multi-view fusion strategy. We evaluated our proposed methods on public EPFL Hippocampus benchmark, achieving state-of-the-art performance of 90.9% in Dice. Quantitative comparison showed the effectiveness of the proposed features and strategies. Jialin Peng, Zhimin Yuan |
IEEE J. Biomed. Health Informatics | 2 |
| 2019 | Spatial Compression for Fronthaul-Constrained Uplink Receiver in 5G SystemsabstractCloud-radio access network (C-RAN) based architecture is widely considered as a fundamental part of 5G networks, where the disaggregation of RAN functionality between a Central Unit (CU) and multiple Distributed Unites (DU) is under hot discussion. With massive multiple-input-multiple-output (MIMO) implemented in the DU part to guarantee the high throughput performance, the limited capacity of the fronthaul between CU and DU becomes the bottleneck for the system performance improvement in the network deployment. Therefore, the spatial compression technique is essential in the uplink receiver structure. In this work, the fronthaul-constrained uplink CU-DU structure is fully discussed and the performance of two alternatives, closed-loop and open-loop spatial compression, are evaluated and compared with 2D cross-polarized antenna arrays. A 2D Kronecker product codebook to dynamically adapt with the antenna array dimension is proposed to achieve better compression performance and it is investigated in the context of the 3rdGeneration Partnership Project (3GPP) 5G New Radio (NR) Release 15 standard. The link level simulation results show that the performance with the spatial compression is sensitive with the oversampling size of codebook and reference signal transmission period and the proposed adaptive codebook outperforms the traditional ones. Zhimin Yuan |
WCNC | 4 |
| 2007 | Asynchronous Spiking Neural P System with Promoters
Zhimin Yuan |
APPT | 1 |