EDBT 2026 Demo / reviewers in the wild / expert
Yanmin Wu
dblp:151/5893
· DBLP profile ↗
26ranked-venue papers
9as first author
25since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 16 · 4 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 3 first-author · 9 since 2021Systems, architecture and hardware · 3 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 3 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Improved WTOD Protocol-Based H∞ Fuzzy PID Control for Networked Nonlinear Systems by an Adaptive Momentum Estimation AlgorithmabstractThis paper is concerned with the problem of the optimalH∞fuzzy proportional-integral-derivative (PID) controller design for Takagi-Sugeno fuzzy systems subjected to infinite-distributed time delays and weighted try-once-discard (WTOD) scheduling effects. To avoid data collisions in communication networks, an improved WTOD communication protocol is proposed, which offers higher transmission precedence to the most needed data node. According to non-parallel distributed compensation scheme, a novel yet easy-to-implement fuzzy PID controller is devised, in which the integral term of the fuzzy PID controller equipped with a limited time window can be utilized for mitigating computational burden. Subsequently, considering the feasible regions of membership functions (MFs) of the fuzzy PID controller, in the case of guaranteeing stability of explored systems, a novel MFs online iterative learning strategy employing an adaptive momentum estimation algorithm is first provided for networked fuzzy systems. The proposed algorithm addresses the limitation of designing fuzzy controller MFs based on designers’ experience. On the basis of the provided MFs online iterative learning approach, MFs of the fuzzy PID controller are renewed in real-time for achieving betterH∞performance, namely, disturbance attenuation capacity for the explored systems. Ultimately, the feasibility of the developed control strategy is illustrated by means of some simulation results. Yanmin Wu, Wei Qian 0002, Bo Shen 0001 |
IEEE Internet Things J. | 1 |
| 2026 | Adaptive Dynamic Event-Triggered H∞ Control for Fuzzy Systems Through an Online Iterative Asynchronous Premise Reconstruction StrategyabstractThis paper is concerned with the adaptive dynamic event-triggered (DET)H∞control for a class of networked nonlinear systems with non-uniform sampling under the Takagi-Sugeno (T-S) fuzzy model. To address the problems of high energy consumption in sensors and the high occupancy rate of limited bandwidth, an improved adaptive DET communication mechanism with non-uniform sampling is proposed, in which the interval dynamic variable is constructed with an adaptive law. Unlike existing works, an online iterative asynchronous premise reconstruction (APR) technique is devised to tackle the challenges caused by the mismatch of premise variables. Based on the adaptive DET mechanism and the online iterative APR method, switched-like fuzzy controllers are designed, which can be switched at different triggering instances. Furthermore, to ensure stability and achieve the desired control performance, new triggering instants-dependent Lyapunov functions are construct- ed in accordance with the idea of event interval partitioning. Fi- nally, the practicability of the proposed strategy is demonstrated using two practical examples. Yanmin Wu, Wei Qian 0002, Zidong Wang 0001 |
IEEE Trans Autom. Sci. Eng. | 1 |
| 2025 | InstanceGaussian: Appearance-Semantic Joint Gaussian Representation for 3D Instance-Level Perceptionabstract3D scene understanding is vital for applications in autonomous driving, robotics, and augmented reality. However, scene understanding based on 3D Gaussian Splatting faces three key challenges: (i) an imbalance between appearance and semantics, (ii) inconsistencies in object boundaries, and (iii) difficulties with top-down instance segmentation. To address these challenges, we propose InstanceGaussian, a method that jointly learns appearance and semantic features while adaptively aggregating instances. Our contributions are as follows: (i) a new Semantic-Scaffold-GS representation to improve feature representation and boundary delineation, (ii) a progressive training strategy for enhanced stability and segmentation, and (iii) a category-agnostic, bottom-up instance aggregation approach for better segmentation. Experimental results demonstrate that our approach achieves state-of-the-art performance in category-agnostic, open-vocabulary 3D point-level segmentation, validating the effectiveness of our proposed method. Project page: https://lhj-git.github.io/InstanceGaussian/ Haijie Li, Yanmin Wu, Jiarui Meng, Qiankun Gao, Ronggang Wang |
CVPR | 2 |
| 2025 | MMSearch: Unveiling the Potential of Large Models as Multi-modal Search EnginesabstractThe advent of Large Language Models (LLMs) has paved the way for AI search engines, e.g., SearchGPT, showcasing a new paradigm in human-internet interaction. However, most current AI search engines are limited to text-only settings, neglecting the multimodal user queries and the text-image interleaved nature of website information. Recently, Large Multimodal Models (LMMs) have made impressive strides. Yet, whether they can function as AI search engines remains under-explored, leaving the potential of LMMs in multimodal search an open question. To this end, we first design a delicate pipeline, MMSearch-Engine, to empower any LMMs with multimodal search capabilities. On top of this, we introduce MMSearch, a comprehensive evaluation benchmark to assess the multimodal search performance of LMMs. The curated dataset contains 300 manually collected instances spanning 14 subfields, which involves no overlap with the current LMMs' training data, ensuring the correct answer can only be obtained within searching. By using MMSearch-Engine, the LMMs are evaluated by performing three individual tasks (requery, rerank, and summarization), and one challenging end-to-end task with a complete searching process. We conduct extensive experiments on closed-source and open-source LMMs. Among all tested models, GPT-4o with MMSearch-Engine achieves the best results, which surpasses the commercial product, Perplexity Pro, in the end-to-end task, demonstrating the effectiveness of our proposed pipeline. We further present error analysis to unveil current LMMs still struggle to fully grasp the multimodal search tasks, and conduct ablation study to indicate the potential of scaling test-time computation for AI search engine. We hope MMSearch may provide unique insights to guide the future development of multimodal AI search engine. Dongzhi Jiang, Renrui Zhang, Yanmin Wu, Jiayi Lei, Pengshuo Qiu, Pan Lu, Guanglu Song, Peng Gao 0007, Yu Liu 0015, Chunyuan Li, Hongsheng Li 0001 |
ICLR | 4 |
| 2025 | SecureGS: Boosting the Security and Fidelity of 3D Gaussian Splatting Steganographyabstract3D Gaussian Splatting (3DGS) has emerged as a premier method for 3D representation due to its real-time rendering and high-quality outputs, underscoring the critical need to protect the privacy of 3D assets. Traditional NeRF steganography methods fail to address the explicit nature of 3DGS since its point cloud files are publicly accessible. Existing GS steganography solutions mitigate some issues but still struggle with reduced rendering fidelity, increased computational demands, and security flaws, especially in the security of the geometric structure of the visualized point cloud. To address these demands, we propose a \textbf{SecureGS}, a secure and efficient 3DGS steganography framework inspired by Scaffold-GS's anchor point design and neural decoding. SecureGS uses a hybrid decoupled Gaussian encryption mechanism to embed offsets, scales, rotations, and RGB attributes of the hidden 3D Gaussian points in anchor point features, retrievable only by authorized users through privacy-preserving neural networks. To further enhance security, we propose a density region-aware anchor growing and pruning strategy that adaptively locates optimal hiding regions without exposing hidden information. Extensive experiments show that SecureGS significantly surpasses existing GS steganography methods in rendering fidelity, speed, and security. Xuanyu Zhang 0003, Jiarui Meng, Zhipei Xu, Shuzhou Yang, Yanmin Wu, Ronggang Wang, Jian Zhang 0018 |
ICLR | 5 |
| 2025 | ReCon-GS: Continuum-Preserved Gaussian Streaming for Fast and Compact Reconstruction of Dynamic ScenesabstractTo address these challenges, we propose the Reconfigurable Continuum Gaussian Stream, dubbed ReCon-GS, a novel storage-aware framework that enables high-fidelity online dynamic scene reconstruction and real-time rendering. Specifically, we dynamically allocate multi-level Anchor Gaussians in a density-adaptive fashion to capture inter-frame geometric deformations, thereby decomposing scene motion into compact coarse-to-fine representations. Then, we design a dynamic hierarchy reconfiguration strategy that preserves localized motion expressiveness through on-demand anchor re-hierarchization, while ensuring temporal consistency through intra-hierarchical deformation inheritance that confines transformation priors to their respective hierarchy levels. Furthermore, we introduce a storage-aware optimization mechanism that flexibly adjusts the density of Anchor Gaussians at different hierarchy levels, enabling a controllable trade-off between reconstruction fidelity and memory usage. Extensive experiments on three widely used datasets demonstrate that, compared to state‐of‐the‐art methods, ReCon-GS improves training efficiency by approximately 15% and achieves superior FVV synthesis quality with enhanced robustness and stability. Moreover, at equivalent rendering quality, ReCon-GS slashes memory requirements by over 50% compared to leading state‑of‑the‑art methods. Code is avaliable at: https://github.com/jyfu-vcl/ReCon-GS/. Jiaye Fu, Qiankun Gao, Chengxiang Wen, Yanmin Wu, Siwei Ma 0001, Jiaqi Zhang 0007 |
NeurIPS | 4 |
| 2025 | Adaptive memory-event-triggered-based double asynchronous fuzzy control for nonlinear semi-Markov jump systems
Yanmin Wu, Wei Qian 0002 |
Fuzzy Sets Syst. | 1 |
| 2025 | Consensus of second-order multi-agent systems based on PIDD-like control protocol with time delay
Wei Qian 0002, Yanmin Wu |
Neurocomputing | 3 |
| 2025 | DYO-SLAM: Visual Localization and Object Mapping in Dynamic ScenesabstractAddressing the impact of dynamic factors on localization accuracy and constructing a long-term consistent map containing only static elements are two crucial tasks in visual simultaneous localization and mapping (SLAM) for dynamic scenes. The introduction of dynamic elements can compromise the geometric constraints essential for visual SLAM, leading to a decrease in localization accuracy. Existing related research faces challenges in simultaneously ensuring localization accuracy in both low-dynamic and high-dynamic scenarios, while also maintaining the system’s real-time performance. To address this issue, we propose a two-stage, coarse-to-fine static-probability-based localization scheme. The construction of object-level maps offers strong support for tasks involving higher-level intelligent agent manipulation as well as augmented reality (AR). However, current research is inadequate for dynamic scenes where the objects to be modeled are frequently and irregularly obscured by dynamic objects, and where there are significant challenges such as severe image and point cloud noise, semantic noise, and lack of observational perspectives. To overcome these challenges, we first propose an object parameter estimation algorithm that combines clustering, weighted Principal Component Analysis (PCA) based on an energy function, and a minimum bounding rectangle. Then, we design a multi-modal object data association strategy based on appearance, semantic, and spatial features. The proposed object parameter estimation algorithm and data association strategy demonstrate improved accuracy and robustness in dynamic scenes with the aforementioned challenges. Finally, based on the entire system, we further develop a dynamic object tracking algorithm and construct an AR system to demonstrate the system’s application prospects. A series of public datasets and real-world scene results have been used to evaluate the effectiveness of the proposed system. Xinggang Hu, Yanmin Wu, Zhenzhong Cao, Xiangkui Zhang, Xiangyang Ji |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | PAS-SLAM: A Visual SLAM System for Planar-Ambiguous ScenesabstractVisual SLAM (Simultaneous Localization and Mapping) systems based on planar features have been widely applied in fields such as environmental structure perception and augmented reality (AR). However, current research still faces challenges in accurate localization and map construction in planar ambiguous scenes, primarily due to the insufficient accuracy of the planar features and data association methods employed. In this paper, we propose a visual SLAM system based on planar features designed for ambiguous planar scenes, including planar analysis and processing, data association, and multi-constraint factor graph optimization. Initially, we introduce a planar analysis and processing strategy that integrates semantic information to analyze the structure of planes and further refine the selection of planes, providing accurate planar information for subsequent association and optimization processes. Then, we integrate various planar data to propose a multimodal fusion data association strategy, achieving accurate and robust planar data association in ambiguous planar scenes. Finally, based on accurate and rich planar information along with related constraints, we design a set of multi-constraint factor graphs for camera pose optimization. Public datasets and real-world experiments demonstrate that, compared to state-of-the-art related research, our proposed system shows significant competitive advantages in terms of accuracy and robustness for both map construction and camera localization. Regarding quantifiable localization accuracy, our system achieves an average improvement in Absolute Trajectory Error (ATE) of approximately 57% in planar ambiguous scenes and about 25% in non-planar ambiguous scenes. Additionally, the system exhibits great application potential in fields such as augmented reality. Xinggang Hu, Yanmin Wu, Linghao Yang, Xiangkui Zhang, Xiangyang Ji |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | UniQuadric: A SLAM Backend for Unknown Rigid Object 3-D Tracking and Light-Weight ModelingabstractTracking and modeling unknown rigid objects in the environment play a crucial role in autonomous uncrewed systems and virtual-real interactive applications. However, many existing Simultaneous Localization, Mapping and Moving Object Tracking (SLAMMOT) methods focus solely on estimating specific object poses and lack estimation of object scales and are unable to effectively track unknown objects. In this paper, we propose a novel SLAM backend that unifies ego-motion tracking, rigid object motion tracking, and modeling within a joint optimization framework. In the perception part, we designed a pixel-level asynchronous object tracker (AOT) based on the Segment Anything Model (SAM) and DeAOT, enabling the tracker to effectively track target unknown objects guided by various predefined tasks and prompts. In the modeling part, we present a novel object-centric quadric parameterization to unify both static and dynamic object initialization and optimization. Subsequently, in the part of object state estimation, we propose a tightly coupled optimization model for object pose and scale estimation, incorporating hybrids constraints into a novel dual sliding window optimization framework for joint estimation. To our knowledge, we are the first to tightly couple object pose tracking with light-weight modeling of dynamic and static objects using quadric. We conduct qualitative and quantitative experiments on simulation datasets and real-world datasets, demonstrating the state-of-the-art robustness and accuracy in motion estimation and modeling. This showcases the significant potential of our method for object perception in complex dynamic environments. Linghao Yang, Yanmin Wu, Rui Tian 0002, Xinggang Hu |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2025 | Language-Assisted 3D Scene UnderstandingabstractThe scale and quality of point cloud datasets constrain the advancement of point cloud learning. Recently, with the development of multi-modal learning, the incorporation of domain-agnostic prior knowledge from other modalities, such as images and text, to assist in point cloud feature learning has been considered a promising avenue. Existing methods have demonstrated the effectiveness of multi-modal contrastive training and feature distillation on point clouds. However, challenges remain, including the requirement for paired triplet data, redundancy and ambiguity in supervised features, and the disruption of the original priors. In this paper, we propose alanguage-assisted approach topointcloud featurelearning (LAST-PCL), enriching semantic concepts through large language model-based text enrichment. We achieve de-redundancy and feature dimensionality reduction without compromising textual priors by statistical-based and training-free significant feature selection. Furthermore, we also delve into an in-depth analysis of the impact of text contrastive training on the point cloud. Extensive experiments validate that the proposed method learns semantically meaningful point cloud features and achieves state-of-the-art or comparable performance in 3D semantic segmentation, 3D object detection, and 3D scene classification tasks. Yanmin Wu, Qiankun Gao, Renrui Zhang, Haijie Li, Jian Zhang 0018 |
IEEE Trans. Multim. | 1 |
| 2025 | New WTOD Protocol-Based Fault Detection Filter Design for Interval Type-2 Fuzzy Systems via an Adaptive Differential Evolution AlgorithmabstractThis article is concerned with the design problem of an $H_{\infty }$ optimal fault detection (FD) filter for networked interval type-2 (IT2) fuzzy systems that are subjected to stochastic cyberattacks. To effectively reduce the utilization of constrained network resources, a new dynamically adjusted event-triggered weighted try-once-discard (DAET-WTOD) protocol is developed, in which two adaptive rules are constructed based on the measured output and the probability of denial-of-service (DoS) attacks. Furthermore, a fuzzy switched-like FD filter is designed with the purpose of detecting system fault signals, while simultaneously considering the DAET-WTOD protocol and stochastic cyberattacks. Subsequently, by utilizing an imperfect premise matching (IPM) scheme, an opposition-based learning adaptive differential evolution algorithm is proposed to deal with the networked IT2 fuzzy systems. This algorithm is capable of iteratively searching the membership function values of the fuzzy filter in real time, thereby achieving improved $H_{\infty }$ performance. Finally, some simulation results are provided to verify the feasibility and advantages of the proposed $H_{\infty }$ optimal FD technique. Wei Qian 0002, Yanmin Wu, Zidong Wang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2025 | A Novel Finite-Frequency Optimization Fault Detection for Fuzzy Systems by Membership Functions Iterative LearningabstractThis article presents the finite-frequency optimization fault detection (FD) strategy for Takagi-Sugeno (T-S) fuzzy systems. Under the imperfect premise matching (IPM) policy, a weighted fuzzy FD observer (WFFDO) with the$L_{\infty }/L_{2}$robustness performance and the finite-frequency$ H_{-}$fault sensitivity performance is first proposed, which signifies the residual signal is robust to the external interference and sensitive to potential faults. Some parameters and slack matrices are introduced to obtain more relaxed conditions of designing the WFFDO with mixed performance. Afterward, a new online membership functions (MFs) iterative learning algorithm with the exponential decay learning rate is proposed for the sake of updating the observer MFs in real-time such that optimal$L_{\infty }/L_{2}$performance can be achieved in this article. In addition, sufficient criterion is established so as to ensure the convergence of the structured mean squared error cost function by means of Lyapunov stability theory. Eventually, two simulation examples are given for illustrating the feasibility and superiority of the developed optimization FD technique. Wei Qian 0002, Yanmin Wu, Junqi Yang |
IEEE Trans. Syst. Man Cybern. Syst. | 2 |
| 2024 | OpenGaussian: Towards Point-Level 3D Gaussian-based Open Vocabulary UnderstandingabstractThis paper introduces OpenGaussian, a method based on 3D Gaussian Splatting (3DGS) that possesses the capability for 3D point-level open vocabulary understanding. Our primary motivation stems from observing that existing 3DGS-based open vocabulary methods mainly focus on 2D pixel-level parsing. These methods struggle with 3D point-level tasks due to weak feature expressiveness and inaccurate 2D-3D feature associations. To ensure robust feature presentation and 3D point-level understanding, we first employ SAM masks without cross-frame associations to train instance features with 3D consistency. These features exhibit both intra-object consistency and inter-object distinction. Then, we propose a two-stage codebook to discretize these features from coarse to fine levels. At the coarse level, we consider the positional information of 3D points to achieve location-based clustering, which is then refined at the fine level.
Finally, we introduce an instance-level 3D-2D feature association method that links 3D points to 2D masks, which are further associated with 2D CLIP features. Extensive experiments, including open vocabulary-based 3D object selection, 3D point cloud understanding, click-based 3D object selection, and ablation studies, demonstrate the effectiveness of our proposed method. The source code is available at our project page https://3d-aigc.github.io/OpenGaussian. Yanmin Wu, Jiarui Meng, Haijie Li, Chenming Wu, Yahao Shi, Xinhua Cheng, Chen Zhao 0011, Haocheng Feng, Errui Ding, Jingdong Wang 0001, Jian Zhang 0018 |
NeurIPS | 1 |
| 2024 | Mirror-3DGS: Incorporating Mirror Reflections into 3D Gaussian Splattingabstract3D Gaussian Splatting (3DGS) has significantly advanced 3D scene reconstruction and novel view synthesis. However, like Neural Radiance Fields (NeRF), 3DGS struggles with accurately modeling physical reflections, particularly in mirrors, leading to incorrect reconstructions and inconsistent reflective properties. To address this challenge, we introduce Mirror-3DGS, a novel framework designed to accurately handle mirror geometries and reflections, thereby generating realistic mirror reflections. By incorporating mirror attributes into 3DGS and leveraging plane mirror imaging principles, Mirror-3DGS simulates a mirrored viewpoint from behind the mirror, enhancing the realism of scene renderings. Extensive evaluations on both synthetic and real-world scenes demonstrate that our method can render novel views with improved fidelity in real-time, surpassing the state-of-the-art Mirror-NeRF, especially in mirror regions. Jiarui Meng, Haijie Li, Yanmin Wu, Qiankun Gao, Shuzhou Yang, Jian Zhang 0018, Siwei Ma 0001 |
VCIP | 3 |
| 2024 | Event-Driven Reduced-Order Fault Detection Filter Design for Nonlinear Systems With Complex Communication ChannelabstractThis article studies the problem of adaptive event-triggered-based reduced-order fault detection (FD) filter design for a category of nonlinear systems with complex communication channel under interval type-2 (IT2) fuzzy-approximation technique. For the purpose of further saving communication resources, a new adaptive event-triggered mechanism equipped with more comprehensive system information is proposed, which can also ensure that the data transmission rate is not lower than the minimum limit predesigned in real time. Considering the complexity of communication network environment, a new data transmission mathematical model with fading measurements and network induced delays is constructed. In addition, to decrease the design complexity of FD systems, a reduced-order FD filter subject to the adaptive event-triggered mechanism as well as mismatched membership functions is devised. By means of the Lyapunov theory, the sufficient criteria are derived to ensure stochastic stability with$H_{\infty }$index for the FD systems. Finally, two practical examples are provided for the sake of verifying the effectiveness of the presented FD technique. Wei Qian 0002, Yanmin Wu, Junqi Yang |
IEEE Trans. Fuzzy Syst. | 2 |
| 2023 | Panoptic Compositional Feature Field for Editable Scene Rendering with Network-Inferred Labels via Metric LearningabstractDespite neural implicit representations demonstrating impressive high-quality view synthesis capacity, decom-posing such representations into objects for instance-level editing is still challenging. Recent works learn object-compositional representations supervised by ground truth instance annotations and produce promising scene editing results. However, ground truth annotations are manually labeled and expensive in practice, which limits their usage in real-world scenes. In this work, we attempt to learn an object-compositional neural implicit representation for editable scene rendering by leveraging labels inferred from the off-the-shelf 2D panoptic segmentation networks instead of the ground truth annotations. We propose a novel framework named Panoptic Compositional Feature Field (PCFF), which introduces an instance quadruplet metric learning to build a discriminating panoptic feature space for reliable scene editing. In addition, we propose semantic-related strategies to further exploit the correlations between semantic and appearance attributes for achieving better rendering results. Experiments on multiple scene datasets including ScanNet, Replica, and ToyDesk demonstrate that our proposed method achieves superior performance for novel view synthesis and produces convincing real-world scene editing results. Xinhua Cheng, Yanmin Wu, Mengxi Jia, Jian Zhang 0018 |
CVPR | 2 |
| 2023 | EDA: Explicit Text-Decoupling and Dense Alignment for 3D Visual Groundingabstract3D visual grounding aims to find the object within point clouds mentioned by free-form natural language descriptions with rich semantic cues. However, existing methods either extract the sentence-level features coupling all words or focus more on object names, which would lose the word-level information or neglect other attributes. To alleviate these issues, we present EDA that Explicitly Decouples the textual attributes in a sentence and conducts Dense Alignment between such fine-grained language and point cloud objects. Specifically, we first propose a text decoupling module to produce textual features for every semantic component. Then, we design two losses to supervise the dense matching between two modalities: position alignment loss and semantic alignment loss. On top of that, we further introduce a new visual grounding task, locating objects without object names, which can thoroughly evaluate the model's dense alignment capacity. Through experiments, we achieve state-of-the-art performance on two widely-adopted 3D visual grounding datasets, ScanRefer and SR3D/NR3D, and obtain absolute leadership on our newly-proposed task. The source code is available at https://github.com/yanmin-wu/EDA. Yanmin Wu, Xinhua Cheng, Renrui Zhang, Zesen Cheng, Jian Zhang 0018 |
CVPR | 1 |
| 2023 | Implicit Neural Representation for Cooperative Low-light Image EnhancementabstractThe following three factors restrict the application of existing low-light image enhancement methods: unpredictable brightness degradation and noise, inherent gap between metric-favorable and visual-friendly versions, and the limited paired training data. To address these limitations, we propose an implicit Neural Representation method for Cooperative low-light image enhancement, dubbed NeRCo. It robustly recovers perceptual-friendly results in an unsupervised manner Concretely, NeRCo unifies the diverse degradation factors of real-world scenes with a controllable fitting function, leading to better robustness. In addition, for the output results, we introduce semantic-oriented supervision with priors from the pretrained vision-language model. Instead of merely following reference images, it encourages results to meet subjective expectations, finding more visual-friendly solutions. Further, to ease the reliance on paired data and reduce solution space, we develop a dual-closed-loop constrained enhancement module. It is trained cooperatively with other affiliated modules in a self-supervised manner. Finally, extensive experiments demonstrate the robustness and superior effectiveness of our proposed NeRCo. Our code is available at https://glthub.com/Ysz2022/NeRCo. Shuzhou Yang, Moxuan Ding, Yanmin Wu, Jian Zhang 0018 |
ICCV | 3 |
| 2023 | BSH-Det3D: Improving 3D Object Detection with BEV Shape HeatmapabstractThe progress of LiDAR-based 3D object detection has significantly enhanced developments in autonomous driving and robotics. However, due to the limitations of LiDAR sensors, object shapes suffer from deterioration in occluded and distant areas, which creates a fundamental challenge to 3D perception. Existing methods estimate specific 3D shapes and achieve remarkable performance. However, these methods rely on extensive computation and memory, causing imbalances between accuracy and real-time performance. To tackle this challenge, we propose a novel LiDAR-based 3D object detection model named BSH-Det3D, which applies an effective way to enhance spatial features by estimating complete shapes from a bird's eye view (BEV). Specifically, we design the Pillar-based Shape Completion (PSC) module to predict the probability of occupancy whether a pillar contains object shapes. The PSC module generates a BEV shape heatmap for each scene. After integrating with heatmaps, BSH-Det3D can provide additional information in shape deterioration areas and generate high-quality 3D proposals. We also design an attention-based densification fusion module (ADF) to adaptively associate the sparse features with heatmaps and raw points. The ADF module integrates the advantages of points and shapes knowledge with negligible overheads. Extensive experiments on the KITTI benchmark achieve state-of-the-art (SOTA) performance in terms of accuracy and speed, demonstrating the efficiency and flexibility of BSH-Det3D. The source code is available on https://github.com/mystorm16/BSH-Det3D. You Shen, Yunzhou Zhang, Yanmin Wu, Zhenyu Wang 0010, Linghao Yang, Sonya A. Coleman, Dermot Kerr |
IROS | 3 |
| 2023 | An Object SLAM Framework for Association, Mapping, and High-Level TasksabstractObject SLAM is considered increasingly significant for robot high-level perception and decision-making. Existing studies fall short in terms of data association, object representation, and semantic mapping and frequently rely on additional assumptions, limiting their performance. In this article, we present a comprehensive object SLAM framework that focuses on object-based perception and object-oriented robot tasks. First, we propose an ensemble data association approach for associating objects in complicated conditions by incorporating parametric and nonparametric statistic testing. In addition, we suggest an outlier-robust centroid and scale estimation algorithm for modeling objects based on the iForest and line alignment. Then a lightweight and object-oriented map is represented by estimated general object models. Taking into consideration the semantic invariance of objects, we convert the object map to a topological map to provide semantic descriptors to enable multimap matching. Finally, we suggest an object-driven active exploration strategy to achieve autonomous mapping in the grasping scenario. A range of public datasets and real-world results in mapping, augmented reality, scene matching, relocalization, and robotic manipulation have been used to evaluate the proposed object SLAM framework for its efficient performance. Yanmin Wu, Yunzhou Zhang, Delong Zhu 0001, Zhiqiang Deng, Wenkai Sun, Jian Zhang 0018 |
IEEE Trans. Robotics | 1 |
| 2022 | CFP-SLAM: A Real-time Visual SLAM Based on Coarse-to-Fine Probability in Dynamic EnvironmentsabstractThe dynamic factors in the environment will lead to the decline of camera localization accuracy due to the violation of the static environment assumption of SLAM algorithm. Recently, some related works generally use the combination of semantic constraints and geometric constraints to deal with dynamic objects, but problems can still be raised, such as poor real-time performance, easy to treat people as rigid bodies, and poor performance in low dynamic scenes. In this paper, a dynamic scene-oriented visual SLAM algorithm based on object detection and coarse-to-fine static probability named CFP-SLAM is proposed. The algorithm combines semantic constraints and geometric constraints to calculate the static probability of objects, keypoints and map points, and takes them as weights to participate in camera pose estimation. Extensive evaluations show that our approach can achieve almost the best results in high dynamic and low dynamic scenarios compared to the state-of-the-art dynamic SLAM methods, and shows quite high real-time ability. Xinggang Hu, Yunzhou Zhang, Zhenzhong Cao, Yanmin Wu, Zhiqiang Deng, Wenkai Sun |
IROS | 5 |
| 2022 | Security-Based Fuzzy Control for Nonlinear Networked Control Systems With DoS Attacks via a Resilient Event-Triggered SchemeabstractThis article studies the issue of resilient event-triggered (RET)-based security controller design for nonlinear networked control systems (NCSs) described by interval type-2 (IT2) fuzzy models subject to nonperiodic denial of service (DoS) attacks. Under the nonperiodic DoS attacks, the state error caused by the packets loss phenomenon is transformed into an uncertain variable in the designed event-triggered condition. Then, an RET strategy based on the uncertain event-triggered variable is firstly proposed for the nonlinear NCSs. The existing results that utilized the hybrid triggered scheme have the defect of complex control structure, and most of the security compensation methods for handling the impacts caused by DoS attacks need to transmit some compensation data when the DoS attacks disappear, which may lead to large performance loss of the systems. Different from these existing results, the proposed RET strategy can transmit the necessary packets to the controller under nonperiodic DoS attacks to reduce the performance loss of the systems and a new security controller subject to the RET scheme and mismatched membership functions is designed to simplify the network control structure under DoS attacks. Finally, some simulation results are utilized to testify the advantages of the presented approach. Yingnan Pan, Yanmin Wu, Hak-Keung Lam |
IEEE Trans. Fuzzy Syst. | 2 |
| 2021 | Object SLAM-Based Active Mapping and Robotic GraspingabstractThis paper presents the first active object mapping framework for complex robotic manipulation and autonomous perception tasks. The framework is built on an object SLAM system integrated with a simultaneous multi-object pose estimation process that is optimized for robotic grasping. Aiming to reduce the observation uncertainty on target objects and increase their pose estimation accuracy, we also design an object-driven exploration strategy to guide the object mapping process, enabling autonomous mapping and high-level perception. Combining the mapping module and the exploration strategy, an accurate object map that is compatible with robotic grasping can be generated. Additionally, quantitative evaluations also indicate that the proposed framework has a very high mapping accuracy. Experiments with manipulation (including object grasping and placement) and augmented reality significantly demonstrate the effectiveness and advantages of our proposed framework. Yanmin Wu, Yunzhou Zhang, Delong Zhu 0001, Sonya A. Coleman, Wenkai Sun, Xinggang Hu, Zhiqiang Deng |
3DV | 1 |
| 2020 | EAO-SLAM: Monocular Semi-Dense Object SLAM Based on Ensemble Data AssociationabstractObject-level data association and pose estimation play a fundamental role in semantic SLAM, which remain unsolved due to the lack of robust and accurate algorithms. In this work, we propose an ensemble data associate strategy for integrating the parametric and nonparametric statistic tests. By exploiting the nature of different statistics, our method can effectively aggregate the information of different measurements, and thus significantly improve the robustness and accuracy of data association. We then present an accurate object pose estimation framework, in which an outliers-robust centroid and scale estimation algorithm and an object pose initialization algorithm are developed to help improve the optimality of pose estimation results. Furthermore, we build a SLAM system that can generate semi-dense or lightweight object-oriented maps with a monocular camera. Extensive experiments are conducted on three publicly available datasets and a real scenario. The results show that our approach significantly outperforms state-of-the-art techniques in accuracy and robustness. The source code is available on https://github.com/yanmin-wu/EAO-SLAM. Yanmin Wu, Yunzhou Zhang, Delong Zhu 0001, Yonghui Feng, Sonya A. Coleman, Dermot Kerr |
IROS | 1 |