Ruyu Liu

dblp:216/8363 · DBLP profile ↗
← Back
47ranked-venue papers
10as first author
41since 2021 · last 2026
0000-0003-2130-9122ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 17 · 4 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 15 · 3 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 3 first-author · 9 since 2021Human-computer interaction and ubiquitous computing · 6 · 1 first-author · 5 since 2021Systems, architecture and hardware · 4 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Computer networks · 1 · 1 since 2021Theory of computation · 1 · 1 since 2021
YearPublicationVenuePosition
2026 WassDRO-AD: Distributionally Robust Anomaly Detection for Multivariate IoT Data Streams
Xiufeng Liu 0001, Ruyu Liu, Yanyan Yang 0002
DEXA (2)2
2026 TwinDB: Interactive What-If Analysis for Digital Twins
Xiufeng Liu 0001, Ruyu Liu, Per Sieverts Nielsen, Hua Lu 0001
EDBT2
2026 NavLLM: Interactive LLM-Assisted Navigation over Multidimensional Data Cubes
Xiufeng Liu 0001, Ruyu Liu, Yanyan Yang 0002
EDBT2
2026 BTKD++: Beyond Teachers by Critically Distilling Knowledge from Teacher's Bias
abstract
Abstract Existing knowledge distillation methods indiscriminately transfer knowledge from teacher networks, including output-level decisional biases, i.e., incorrect final predictions that can mislead student learning and limit student performance. We challenge this paradigm by proposing BTKD++, a framework that systematically filters and rectifies teacher’s output-level biased knowledge into corrective signals. Our approach partitions training data into Easy Tasks (correct teacher predictions) and Hard Tasks (incorrect predictions), then applies bias elimination and rectification modules orchestrated by dynamic learning curriculum. We provide an interpretive information-theoretic abstraction to explain the observed competence-threshold phenomenon, under which bias rectification becomes more effective when teacher errors contain sufficiently structured corrective information. BTKD++ demonstrates broad applicability across classification, detection, and segmentation tasks when task outputs are equipped with suitable probabilistic interfaces, and shows consistent effectiveness across CNNs, Transformers, and State-Space Models. Extensive experiments show consistent student-teacher transcendence, establishing new state-of-the-art results. This work redefines knowledge distillation from blind mimicry to critical learning, proving that students can surpass teachers through principled bias correction. The source code is available at https://github.com/smartyige/BTKD .
Jianhua Zhang 0002, Yu He 0001, Xu Cheng 0003, Xiufeng Liu 0001, Shengyong Chen, Houxiang Zhang, Ruyu Liu
Int. J. Comput. Vis.9
2026 SEMDI-Net: Deep learning techniques for denoising scanning electron microscope images of fiber masterbatches
Ruyu Liu, Lusi Li
Neurocomputing4
2026 Object shape differentiation and texture rendering for neural implicit SLAM
Jiaming Lu, Ruyu Liu, Jianhua Zhang 0002, Xu Cheng 0003
Mach. Vis. Appl.2
2026 Multi-branch perturbation learning with constraint simulation for semi-supervised semantic segmentation
abstract
Current semi-supervised semantic segmentation (SSS) methods improve generalization via weak-to-strong pseudo-supervision with image perturbations. However, many methods are limited by employing a single perturbation mode and a specific weak-to-strong learning strategy, restricting exploration of the perturbation space and hindering performance in fine-grained segmentation. While diverse perturbations are intuitively beneficial, simply combining them can lead to inefficient optimization and instability. In this paper, we propose a multi-branch strong perturbation constraint learning framework for SSS. Our framework introduces a novel multi-branch perturbation learning (MSPL) strategy, employing multiple parallel branches with diverse strong augmentations to expand the perturbation space and capture complex semantic variations. We further design a novel constraint simulation loss (CSSL), based on a hierarchical consistency learning structure (weak-to-strong and strong-to-strong), which enforces strong-to-strong consistency between different perturbation branches. CSSL mitigates instability and enhances robustness to perturbation-induced noise, enabling the network to better generalize and achieve more accurate segmentation, especially for fine object boundaries. Extensive evaluations on benchmark datasets (PASCAL VOC 2012, Cityscapes, COCO) demonstrate that our method achieves state-of-the-art performance. Ablation studies further validate the effectiveness of our proposed MSPL and CSSL components.
Ruyu Liu, Feng Xiao 0005, Jianhua Zhang 0002, Xiufeng Liu 0001, Xu Cheng 0003, Shengyong Chen, Houxiang Zhang
Pattern Recognit.1
2026 PIVOT: A Framework for High-Fidelity Facade PV Layouts via Generative Seeding and Differentiable Optimization
abstract
The automation of Building-Integrated Photovoltaic (BIPV) design is essential for the digital transformation of urban renewable energy, yet current methods struggle to bridge the gap between semantic facade assessment and engineering-grade installation plans. Existing AI-based approaches typically yield raw segmentation masks that lack the geometric precision and physical validity required for real-world deployment. This paper presents PIVOT, a novel end-to-end framework for the automated synthesis of high-fidelity, collision-free PV layouts from single 2D facade images. We formulate the layout generation as a two-stage neuro-symbolic and differentiable optimization process. First, an LLM-based programmatic seeding stage uses structured facade information and elementary spatial-relation cues to generate initial layout hypotheses. Second, a differentiable layout optimization engine refines these proposals by treating panel placement as a continuous constrained optimization problem, effectively resolving spatial overlaps while maximizing geometric regularity. Validated on a diverse dataset of 80 building facades, experimental results demonstrate that PIVOT significantly outperforms heuristic (MaxRects) and evolutionary (MOGA) baselines. Our framework reduces the layout collision rate from 9.7% to 0.7% and improves geometric regularity by over 10%, while requiring only approximately 80 seconds of processing time per building. This work advances BIPV design automation by converting semantically parsed facade information into geometrically constrained, constructibility-oriented PV layouts for downstream performance evaluation.
Dongxu Zhuang, Ruyu Liu, Xiufeng Liu 0001, Per Sieverts Nielsen, Tangao Hu, Jianhua Zhang 0002
IEEE Trans Autom. Sci. Eng.2
2026 Adaptive Kernel Selection Module Combined With Feature Enhanced Perception Network for Camouflaged Object Detection
abstract
Camouflaged object detection plays a crucial role in applications such as automatic sorting and defect inspection in industrial production, yet existing methods often struggle to flexibly capture features of diverse shapes, orientations, and scales due to their reliance on fixed receptive fields and rigid windowing schemes. To address these limitations, we propose a dual-branch joint network comprising a reference branch and a segmentation branch. The reference branch learns supplementary cues from salient objects that co-occur with camouflaged targets, guiding the segmentation branch toward more accurate delineation. Within the segmentation branch, we introduce three novel modules: 1) a deformable window interaction mechanism that replaces fixed-size transformer windows with learnable quadrilateral windows to adaptively extract features of arbitrary shape and orientation; 2) a feature enhancement perception module that fuses rich multiscale representations through parallel dilated convolutions at varying rates and channel-/spatial-attention mechanisms; and 3) a receptive field adjustment adaptive module that dynamically adjusts its receptive field size to balance sensitivity to fine details and global context. Comprehensive experiments on COD10 K, NC4K, CAMO, and R2C7K benchmarks demonstrate that our model outperforms the majority of current state-of-the-art approaches, while ablation studies and sensitivity analyses confirm the individual and combined effectiveness of our proposed components.
Ruyu Liu, Feng Xiao 0005, Jianhua Zhang 0002, Shengyong Chen
IEEE Trans. Ind. Informatics1
2026 Frequency-Aware B-Line and Pleural Line Analysis in Lung Ultrasound Videos
abstract
Accurately identifying B-lines and pleural line (P-line) in lung ultrasound (LUS) videos is valuable for evaluating certain lung conditions. However, manual interpretation remains subjective and highly dependenton operator expertise. Existing deep learning methods often suffer from performance degradation due to speckle noise and motion artifacts. Moreover, the limited availability of LUS video data annotated for multiple diagnostic features such as B-lines and the P-line limits model development. Therefore, this paper introduces ILD-LUS, a new clinical LUS database designed based on interstitial lung disease (ILD) analysis by category labeling, comprising 2,149 ultrasound videos (193,410 frames). Also, we construct an external test set based on the public Covid-BLUES dataset for the evaluation of B-lines and P-line recognition in different pulmonary pathologies. Then, we propose a novel video analysis framework that integrates wavelet enhancement with temporal attention modeling. Specifically, we employ a dual-component frequency feature enhancement method using the Discrete Wavelet Transform (DWT), which effectively suppresses noise while preserving important landmarks. Subsequently, an adaptive attention module is introduced to model long-range temporal dependencies and improve dynamic feature representation across consecutive frames. Experimental results show that the proposed method achieves over 94% AUC and 82% ACC for both B-lines and P-line classification on both the ILD-LUS and Covid-BLUES datasets, outperforming existing methods. These findings demonstrate the robustness and generalizability of our approach across different pathological conditions. Overall, the proposed framework shows strong potential for supporting clinical decision-making in LUS analysis.
Kaihui Yang, Guangyu Guo 0001, Linxuan Pang, Zhaohui Zheng 0004, Ruyu Liu, Jin Ding, Dingwen Zhang, Junwei Han 0001
IEEE J. Biomed. Health Informatics6
2026 InterMamba: Efficient Human-Human Interaction Generation With Adaptive Spatio-Temporal Mamba
abstract
Human-human interaction generation has garnered significant attention in motion synthesis due to its vital role in understanding humans as social beings. However, existing methods typically rely on transformer-based architectures, which often face challenges related to scalability and efficiency. To address these challenges, we propose InterMamba, a novel and efficient human-human interaction generation method built on the Mamba framework, designed to capture long-sequence dependencies effectively while enabling real-time feedback. Specifically, we introduce an adaptive spatio-temporal Mamba framework that utilizes two parallel SSM branches with an adaptive mechanism to integrate the spatial and temporal features of motion sequences. To further enhance the model's ability to capture dependencies within individual motion sequences and the interactions between different individual sequences, we develop two key modules: the self adaptive spatio-temporal Mamba module and the cross adaptive spatio-temporal Mamba module, enabling efficient feature learning. Extensive experiments demonstrate that our method achieves the state-of-the-art results on both two interaction datasets with remarkable quality and efficiency. Compared to the baseline method InterGen, our approach not only improves accuracy but also reduces the parameter size to just 66 M (36% of InterGen's), while achieving an average inference speed of 0.57 seconds, which is 46% of InterGen's execution time.
Zizhao Wu, Xiaoling Gu, Ruyu Liu, Jiazhou Chen 0002
IEEE Trans. Vis. Comput. Graph.5
2025 Can Students Beyond the Teacher? Distilling Knowledge from Teacher's Bias
abstract
Knowledge distillation (KD) is a model compression technique that transfers knowledge from a large teacher model to a smaller student model to enhance its performance. Existing methods often assume that the student model is inherently inferior to the teacher model. However, we identify that the fundamental issue affecting student performance is the bias transferred by the teacher. Current KD frameworks transmit both right and wrong knowledge, introducing bias that misleads the student model. To address this issue, we propose a novel strategy to rectify bias and greatly improve the student model's performance. Our strategy involves three steps: First, we differentiate knowledge and design a bias elimination method to filter out biases, retaining only the right knowledge for the student model to learn. Next, we propose a bias rectification method to rectify the teacher model's wrong predictions, fundamentally addressing bias interference. The student model learns from both the right knowledge and the rectified biases, greatly improving its prediction accuracy. Additionally, we introduce a dynamic learning approach with a loss function that updates weights dynamically, allowing the student model to quickly learn right knowledge-based easy tasks initially and tackle hard tasks corresponding to biases later, greatly enhancing the student model's learning efficiency. To the best of our knowledge, this is the first strategy enabling the student model to surpass the teacher model. Experiments demonstrate that our strategy, as a plug-and-play module, is versatile across various mainstream KD frameworks.
Jianhua Zhang 0002, Ruyu Liu, Xu Cheng 0003, Houxiang Zhang, Shengyong Chen
AAAI3
2025 Multi-Stage Knowledge Distillation for Progressive Image Deblurring in Teacher-Student Networks
abstract
Image deblurring is the process of recovering a high-quality image from a degraded image, which can be lost sharpness by blur filters, noise, compression, or other degradation factors. Image deblurring is a challenging task, as it requires dealing with complex and ill-posed inverse problems, and balancing the trade-off between performance and efficiency. This paper introduces a multi-stage knowledge distillation approach using teacher-student network interactions. Our method employs a three-stage teacher network combining Window-based Transformer and Unet models for the first and second stage, followed by a Channel-level Transformer as the third stage to enhance detail extraction. The student network, streamlined for speed, mirrors the teacher's structure with lower complexity, incorporating a standard Unet model and channel attention mechanism. A new Supervised Attention Module is introduced for effective knowledge transfer and feature enhancement. We evaluate our method on various datasets and show significant enhancements in image deblurring quality, while balancing performance and model efficiency. Our findings suggest that our method has great potential for real-world applications and opens up new possibilities for future research in image deblurring techniques.
Ruyu Liu, Xiufeng Liu 0001, Xianchao Zhang 0002
CSCWD4
2025 QCTKD-PU: Quantum Convolutional Transformer with Knowledge Distillation for Efficient and Robust Point Cloud Upsampling
abstract
Point cloud upsampling is crucial for high-fidelity 3D reconstruction in real-time applications such as autonomous systems. Existing methods based on CNNs or Transformers face three limitations: (1) prohibitive computational complexity hindering real-time deployment, (2) insufficient modeling of multi-scale geometric dependencies in sparse data, (3) sensitivity to noise and outliers. To address these challenges, we propose QCTKD-PU, a framework integrating Quantum Convolutional Transformers (QCT) and Knowledge Distillation (KD) for Point cloud Upsampling. The QCT leverages quantum superposition and self-attention to encode high-dimensional features, enabling efficient multi-scale point interaction learning. Simultaneously, KD transfers knowledge from a teacher model to a lightweight student network, reducing computational costs while maintaining accuracy. Experiments on benchmark datasets demonstrate superior performance in geometric accuracy and noise robustness compared to state-of-the-art methods. This work pioneers the synergy of quantum computing and lightweight learning for resource-constrained 3D vision tasks, while the student model achieves real-time and compact deployment, offering a practical solution for collaborative edge systems.
Yunrui Zhu, Ruyu Liu, Xu Cheng 0003, Jianhua Zhang 0002, Shengyong Chen
CSCWD3
2025 SPRGAN: Streamlined Progressive Refinement for Adversarial Point Cloud Video Upsampling
abstract
Getting dense, uniform, time-series point cloud data is critical for effective rendering. However, due to the limited computational power of edge devices, existing methods cannot achieve real-time results, which affects the visual quality of the consumer experience. To effectively address this issue, this paper presents a self-supervised adversarial upsampling method for point cloud video streams called SPR-GAN. In the generator, we design the Temporal Iterative Graph module to learn local features for each frame and captures long-range spatial information using three iterations of graph convolution operations. Then the Contextual Temporal Fusion module is developed to merge information between different frames, synthesizing temporal information and enriching the dynamic feature representation of the point cloud. Meanwhile, in the discriminator, we introduce the Efficient Shape module. Through dynamic graph convolution operations and stacked learning, it significantly improves the resolution efficiency of global shape information in point clouds. The final experiments show that the proposed method exhibits high practicality and superiority. The model achieves a good result on both the D-FAUST and DeformingThing4D-Animals datasets.
Ruyu Liu, Xianchao Zhang 0002, Jianhua Zhang 0002, Xiufeng Liu 0001
ICASSP2
2025 Efficient Large-Scale Scene Point Cloud Upsampling with Implicit Neural Networks and Spatial Hashing
abstract
Point cloud upsampling is a critical challenge in 3D vision, particularly for large-scale, real-world data. We propose ASFNet, a novel implicit neural network-based approach that uniquely combines adaptive spatial feature representation with efficient spatial hashing. This method significantly improves both upsampling quality and computational efficiency. ASFNet first encodes the point cloud as an implicit surface, employing dynamic search and spatial hashing to optimize query point locations rapidly. This approach creates a uniform, continuous field around surfaces, enabling high-fidelity upsampling. Experiments on benchmark datasets, including Oakland 3D dataset and VMR-Oakland-v2, demonstrate ASFNet’s superiority. Our method achieves state-of-the-art performance with a Chamfer Distance of 5.559 × 10–3on Oakland 3D dataset, while reducing processing time by up to 80% compared to existing methods. On the challenging Oakland 3D dataset, ASFNet completes upsampling in just 150 seconds. These results underscore ASFNet’s potential to advance real-time 3D vision applications in areas such as autonomous navigation and augmented reality.
Yunrui Zhu, Ruyu Liu, Xu Cheng 0003, Jianhua Zhang 0002, Xiufeng Liu 0001
ICASSP3
2025 Mamba-SLAM: Enhancing Neural Implicit SLAM with Uncertainty and Mamba
abstract
Current neural implicit SLAM systems struggle with insufficient object shape constraints due to incomplete depth maps and inefficient pixel sampling strategies, leading to inaccuracies in reconstructed scene morphology and texture. To address these limitations, we introduce Mamba-SLAM, a novel framework featuring two key innovations: (1) an uncertainty-based pixel sampling module that enhances object rendering quality in depth-scarce regions by integrating depth map analysis and color image gradients, improving shape and texture consistency by 6.4% on the Replica dataset compared to state-of-the-art methods; and (2) a Mamba-based keyframe selection module that leverages the efficient feature extraction of Mamba to provide rich semantic cues, optimizing pose estimation and enriching reconstruction detail. Experiments on Replica and ScanNet demonstrate that Mamba-SLAM significantly improves scene rendering and object detail, achieving a 13.5% improvement in reconstruction completeness on Replica. The core novelty lies in the synergistic combination of uncertainty-driven pixel selection and Mamba-powered keyframe management for enhanced neural implicit SLAM.
Jiaming Lu, Yunrui Zhu, Ruyu Liu, Xu Cheng 0003, Jianhua Zhang 0002, Xiufeng Liu 0001
ICME3
2025 Efficiency-Optimized Point Cloud Upsampling with Single-Layer Graph Convolution Network
abstract
Point cloud upsampling is a key technology for improving the density and quality of sparse point clouds, with widespread applications in 3D reconstruction, autonomous driving, and environmental perception. However, traditional point cloud upsampling methods, especially those based on multilayer graph convolution networks (GCNs), typically rely on complex feature extraction modules, which increase computational complexity and model parameters, limiting their use in resource-constrained environments. To overcome these challenges, we propose the EO-PU framework, a lightweight and efficient point cloud upsampling method. This framework combines single-layer GCN and rotation-invariant 4D projection encoding (I4DP) technology, significantly reducing computational load and redundant information, thereby improving upsampling efficiency. Specifically, EO-PU first uses I4DP to map the 3D point cloud data to a rotation-robust 4D feature space, ensuring effective capture of geometric information. A single-layer GCN is then employed to aggregate features, reducing network complexity and computational cost. To further enhance upsampling performance, we introduce the EdgeShuffleNet module, which optimizes feature expansion and rearrangement through efficient local feature aggregation. Experimental results show that EO-PU outperforms or matches existing methods across multiple public datasets while significantly reducing model parameters and computation time, making it highly suitable for deployment in resource-constrained environments.
Yunrui Zhu, Feng Xiao 0005, HaoXiao Wang, Ruyu Liu, Xu Cheng 0003, Jianhua Zhang 0002
IJCNN5
2025 PolypSense3D: A Multi-Source Benchmark Dataset for Depth-Aware Polyp Size Measurement in Endoscopy
abstract
Accurate polyp sizing during endoscopy is crucial for cancer risk assessment but is hindered by subjective methods and inadequate datasets lacking integrated 2D appearance, 3D structure, and real-world size information. We introduce PolypSense3D, the first multi-source benchmark dataset specifically targeting depth-aware polyp size measurement. It uniquely integrates over 43,000 frames from virtual simulations, physical phantoms, and clinical sequences, providing synchronized RGB, dense/sparse depth, segmentation masks, camera parameters, and millimeter-scale size labels derived via a novel forceps-assisted in-vivo annotation technique. To establish its value, we benchmark state-of-the-art segmentation and depth estimation models. Results quantify significant domain gaps between simulated/phantom and clinical data and reveal substantial error propagation from perception stages to final size estimation, with the best fully automated pipelines achieving an average Mean Absolute Error (MAE) of 0.95 mm on the clinical data subset. Publicly released under CC BY-SA 4.0 with code and evaluation protocols, PolypSense3D offers a standardized platform to accelerate research in robust, clinically relevant quantitative endoscopic vision. The benchmark dataset and code are available at: https://github.com/HNUicda/PolypSense3D and https://doi.org/10.7910/DVN/K13H89.
Ruyu Liu, Mingming Zhou, Jianhua Zhang 0002, Xiufeng Liu 0001, Xu Cheng 0003, Sixian Chan 0001, Yanbin Shen, Sheng Dai, Yuping Yan, Yaochu Jin, Lingjuan Lyu
NeurIPS1
2025 Watermark Removal via Boundary-Aware Segmentation and Semantic-Guided Diffusion
Zhenjie Jiang, Ruyu Liu, Jianhua Zhang 0002, Mohammed M. Elmogy, Shengyong Chen
PRCV (2)5
2025 Alice-SLAM: Accurate and Lite-Communication Collaborative SLAM for Resource-Constrained Multi-Agent
abstract
Multi-agent collaborative simultaneous localization and mapping (Mac-SLAM) facilitates mutual localization among multi-agent and mapping in unknown environments. However, Mac-SLAM faces two main practical challenges in resource-constrained situations: heavy communication load and conflicts among multi-source maps. To address these issues, we propose Alice-SLAM: an accurate and lite-communication client-server collaborative SLAM system, reducing communication load while accuracy-guaranteed. Specifically, regarding high communication demand, we optimize communication load by compressing keyframe data and sharing only key map information instead of full map information. For inconsistency among multi-maps, we combine specific bundle adjustments (BA) and an adaptive strategy for active map optimization to enhance the consistency of the global map. A set of experiments demonstrates the superior accuracy and reduced communication load of the proposed Alice-SLAM on the EuRoC dataset and in multi-user augmented reality (AR) experiments conducted in our lab, highlighting its effectiveness in resource-constrained cases. We plan to open-source our code1to encourage further research and collaboration in this area.
Kaiqi Chen 0001, Ruyu Liu, Xu Cheng 0003, Jianhua Zhang 0002, Shengyong Chen, Houxiang Zhang, Arash Ajoudani
IEEE J. Sel. Areas Commun.3
2025 SVD-KD: SVD-based hidden layer feature extraction for Knowledge distillation
Jianhua Zhang 0002, Mian Zhou, Ruyu Liu, Xu Cheng 0003, Sasa Nikolic 0002, Shengyong Chen
Pattern Recognit.4
2025 Distortion-Aware Outdoor Panoramic Depth Estimation via Local-Global Fusion
abstract
Outdoor panoramic depth estimation faces significant challenges due to the wide field of view (FoV), complex scene structures, and severe distortion encountered in such environments. Traditional methods, which often use distortion convolution, fall short in capturing global distortion information and extracting rich contextual details from panoramic images. To overcome these limitations, this article introduces a novel dual-branch framework that synergistically merges the advantages of equirectangular projection (ERP) and tangent projection (TP). First, we design a unique dual-branch framework specifically tailored for panoramic depth estimation. In this framework, the convolutional neural networks branch processes ERP images to extract rich local information, enhancing the detail accuracy of depth estimation, while the vision transformers branch processes TPs to capture comprehensive global information, improving the smoothness of depth estimation. Then, we further enhance our method with a distortion-aware weight map module that adapts the influence of different image regions according to their distortion level, thus prioritizing features from areas with less distortion. In addition, we implement a dual attention fusion module to seamlessly integrate features from both branches at corresponding layers. Comprehensive experiments across various outdoor datasets reveal that our method significantly outperforms state-of-the-art techniques in terms of depth estimation accuracy, adeptly balancing the capture of both overarching scene depth and intricate details, potentially revolutionizing applications in industrial informatics, such as autonomous navigation and environmental mapping.
Ruyu Liu, Yihao Ying, Xiufeng Liu 0001, Weiguo Sheng 0001, Jianhua Zhang 0002, Shengyong Chen
IEEE Trans. Ind. Informatics1
2025 Semantic Visual Simultaneous Localization and Mapping: A Survey
abstract
Visual Simultaneous Localization and Mapping (vSLAM) is a cornerstone technology in computer vision and robotics, underpinning applications such as autonomous vehicles and robot navigation. While traditional vSLAM systems have shown significant progress in indoor or outdoor environments, their performance often degrades in complex scenes, limiting their adaptability and robustness. Semantic vSLAM, which integrates high-level semantic information into vSLAM systems, has emerged as a promising solution to address these limitations by enabling a richer understanding of the environment. In this paper, we provide a comprehensive review of semantic vSLAM, offering a critical analysis of its evolution, methods, and challenges. We begin by revisiting the development of traditional vSLAM, emphasizing its limitations and the motivation for incorporating semantic information. Subsequently, we delve into the core modules of semantic vSLAM, including semantic extraction, object association, semantic loop closing, back-end optimization, and semantic mapping. Then, we present a performance comparison of semantic vSLAM systems under two different datasets, indoor and outdoor, respectively. Furthermore, we also provide a comparative analysis of widely used SLAM datasets to provide guidance for performance testing and validation. To further enrich the discussion, we identify unresolved challenges in semantic vSLAM, such as long-term semantic perception and association, open and unstructured environments. We propose future research directions, including balancing computational resources and quantifying system risk, large model-based navigation and mapping, and embodied AI SLAM. By providing key insights and forward-looking perspectives, this work aims to stimulate future research and improve the capabilities of semantic vSLAM in real-world applications.
Kaiqi Chen 0001, Junhao Xiao 0001, Qiyi Tong, Heng Zhang 0023, Ruyu Liu, Jianhua Zhang 0002, Arash Ajoudani, Shengyong Chen
IEEE Trans. Intell. Transp. Syst.6
2025 Covariance Propagation-Based Accurate Loop Detection for High Confusion Environment
Kaiqi Chen 0001, Ruyu Liu, Shengyong Chen, Arash Ajoudani, Jianhua Zhang 0002
IEEE Trans. Intell. Transp. Syst.3
2024 360ORB-SLAM: A Visual SLAM System for Panoramic Images with Depth Completion Network
abstract
With the advent of the Industry 4.0 era and the increasing performance requirements for AR/VR applications and vision assistance and inspection systems in recent years, visual simultaneous localization and mapping (vSLAM) is a fundamental task in computer vision and robotics. However, traditional vSLAM systems are limited by the camera’s narrow field-of-view, resulting in challenges such as sparse feature distribution and lack of dense depth information. To overcome these limitations, this paper proposes a 360ORB-SLAM system for panoramic images that combines with a depth completion network. The system extracts feature points from the panoramic image, utilizes a panoramic triangulation module to generate sparse depth information, and employs a depth completion network to obtain a dense panoramic depth map. Experimental results on our novel panoramic dataset constructed based on Carla demonstrate that the proposed method achieves superior scale accuracy compared to existing monocular SLAM methods and effectively addresses the challenges of feature association and scale ambiguity. The integration of the depth completion network enhances system stability and mitigates the impact of dynamic elements on SLAM performance.
Yuqi Pan, Ruyu Liu, Guodao Zhang, Jianhua Zhang 0002
CSCWD3
2024 ColVO: Colonoscopic Visual Odometry Considering Geometric and Photometric Consistency
abstract
Locating lesions is the primary goal of colonoscopy examinations.3D perception techniques can enhance the accuracy of lesion localization by restoring 3D spatial information of the colon. However, existing methods focus on the local depth estimation of a single frame and neglect the precise global positioning of the colonoscope, thus failing to provide the accurate 3D location of lesions. The root causes of this shortfall is twofold: Firstly, existing methods treat colon depth and colonoscope pose estimation as independent tasks or design them as parallel sub-task branches. Secondly, the light source in the colon environment moves with the colonoscope, leading to brightness fluctuations among continuous frame images. To address these two issues, we propose ColVO, a novel deep learning-based Visual Odometry framework, which can continuously estimate colon depth and colonoscopic pose using two key components: a deep couple strategy for depth and pose estimation (DCDP) and a light consistent calibration mechanism (LCC). DCDP utilization of multimodal fusion and loss function constraints to couple depth and pose estimation modes ensure seamless alignment of geometric projections between consecutive frames. Meanwhile, LCC accounts for brightness variations by recalibrating the luminosity values of adjacent frames, enhancing ColVO's robustness. A comprehensive evaluation of ColVO on colon odometry benchmarks reveals its superiority over state-of-the-art methods in depth and pose estimation. We also demonstrate two valuable applications: immediate polyp localization and complete 3D reconstruction of the intestine. The code for ColVO is available at https://github.com/HNUicda/CoIVO.
Ruyu Liu, Zhengzhe Liu, Guodao Zhang, Jianhua Zhang 0002, Weiguo Sheng 0001, Xiufeng Liu 0001, Yaochu Jin
ACM Multimedia1
2024 Semi-Supervised Camouflaged Object Detection: Multi Information Fusion Combined with Adaptive Receptive Field Selection Network
Feng Xiao 0005, Ruyu Liu, Jianhua Zhang 0002, Shengyong Chen
PRCV (12)3
2024 FeatureB2SENet: point cloud classification of large scenes
Hangli Weng, Guodao Zhang, Ruyu Liu, Ping-Kuo Chen, Liping Wang 0016
Vis. Comput.4
2024 Correction: FeatureB2SENet: point cloud classification of large scenes
Hangli Weng, Guodao Zhang, Ruyu Liu, Ping-Kuo Chen, Liping Wang 0016
Vis. Comput.4
2024 KSRB-Net: a continuous sign language recognition deep learning strategy based on motion perception mechanism
Feng Xiao 0005, Yunrui Zhu, Ruyu Liu, Jianhua Zhang 0002, Shengyong Chen
Vis. Comput.3
2023 LiDAR Point Cloud Classification with Coordinate Attention Blueprint Separation Involution Neural Network
abstract
With the advent of the era of Industry 4.0 and the continuous development of point cloud data acquisition technology, point cloud data has been widely used in unmanned distribution of intelligent logistics. This paper designs a 3D point cloud classification model with coordinate attention, blueprint separation involution neural network (BICANet). Firstly, the combination of 2D features and 3D features is adopted to maintain the spatial structure of the point cloud. Secondly, the Involution network is introduced to reduce the amount of redundant data for neural network computation and improve the whole network computation efficiency. At the same time, to further enhance the network feature learning capability, the blueprint separation convolution is combined with coordinate attention. The experimental results prove that the overall accuracy of BICANet in Vaihingen and GML B datasets reaches 86.0% and 98.8%, respectively. It is highly competitive with the currently available methods.
Guodao Zhang, Liting Dai, Guangjie Zhou, Ruyu Liu
CSCWD5
2023 Dense Depth Completion Based on Multi-Scale Confidence and Self-Attention Mechanism for Intestinal Endoscopy
abstract
Doctors perform limited one-way intestine endoscopy, in which advanced surgical robots with depth sensors, such as stereo and ToF endoscopes, can only provide sparse and incomplete depth information. However, dense, accurate and instant depth estimation during endoscopy is vital for doctors to judge the 3D location and shape of intestinal tissues, which affects the human-robot interaction between doctors and surgical robots, such as the operation on the subsequent moving of the probe. In this paper, we present a deep learning-based dense depth completion method for intestine endoscopy. We utilize the scattered depth information from depth sensors to make up for the deficiency of features in the intestine and design a multi-scale confidence prediction network to extract dense geometric depth features. Then, we introduce the structure awareness module based on the self-attention mechanism in the depth completion network to enhance the geometry and texture features of the intestine. We also present a virtual multi-modal RGBD intestine dataset and conduct comprehensive experiments on a total of three intestine datasets. The experimental results clearly demonstrate that our method achieves better results in all metrics in all intestinal environments compared to state-of-the-art methods.
Ruyu Liu, Zhengzhe Liu, Guodao Zhang, Zhigui Zuo, Weiguo Sheng 0001
ICRA1
2023 DSP-Based Traffic Target Detection for Intelligent Transportation
abstract
Internet of Things (IoT)-based intelligent transportation is attracting more and more attention. As a key component of intelligent transportation, traffic video monitoring is very important, in which vehicle and pedestrian detection on the road is a crucial task. Although vehicle and pedestrian detection through deep learning (DL) may achieve high accuracy, it tends to require high computing resources, which hinders its use on IoT devices. As an important class of IoT devices, digital signal processor (DSP) has the characteristics of low energy consumption, small size, and strong performance, which has been widely used in intelligent transportation. In order to use DL on DSP for accurate vehicle and pedestrian detection, we first propose a series of general tactics to optimize the object detection convolutional neural network (CNN) model, including convolution layer optimization, cache optimization, compiler optimization, intrinsics optimization and direct memory access (DMA) acceleration, and then a parallel scheme to extend the model to run on multicore, and further quantize the implementation of the model. We evaluate it on UA-DETRAC and KITTI datasets. Experimental results show that our method achieves a faster speed than running the same CNN model on a mainstream desktop CPU, with only 0.06% accuracy loss.
Jianhua Zhang 0002, Rucen Wang, Ruyu Liu, Dongyan Guo, Bo Li 0090, Shengyong Chen
IEEE Trans. Intell. Transp. Syst.3
2022 Point Clouds Classification of Large Scenes based on Blueprint Separation Convolutional Neural Network
abstract
In industry 4.0-related applications such as UAVs, autonomous driving, remote sensing, and navigation, environmental perception based on large-scene point clouds plays a crucial role. Accurate point clouds classification is the key and premise of environment perception. In this paper, we propose a new point clouds classification method, FeatureB2SE. First, we design a feature extraction method for point clouds by projecting features in different directions in 2D and 3D to form feature maps. Then, we present a B2SE convolution that can more adequately leverage the advantages from both blueprints separable convolution and Squeeze-and-Excitation networks. To effectively evaluate the performance of FeatureB2SE, extensive experiments have been conducted on two public datasets, GML_B and Vaihingen. The outcome demonstrates that our strategy has achieved state-of-the-art baselines. Specifically, the classification accuracy achieves 98.91% on the GML_B dataset and 85.11% on the Vaihingen dataset, respectively.
Guodao Zhang, Hangli Weng, Ruyu Liu, Menghui Zhang
CSCWD3
2022 Large Scale Point Cloud Classification Base on Graph-MLP++
abstract
Deep learning has made remarkable achievements in the object classification of 2D images. However, the 3D point cloud classification task is still an open challenge due to the point cloud being irregular and with a mass of noise. This work proposed a large-scale point cloud processing framework that can improve the accuracy and efficiency of point cloud classification. The proposed method calculates the feature values of the point cloud to construct the point cloud feature images, then inputs them into the Graph-MLP++ network to get the point cloud classification result. GraphMLP++ can achieve 97.8% accuracy in the Oakland dataset. Compared with other methods, the efficiency and accuracy of the result are competitive
Ruyu Liu, En Xie, Guodao Zhang
DSAA2
2022 Robust and Accurate Multi-Agent SLAM with Efficient Communication for Smart Mobiles
abstract
In a long-term large-scenario application, the multi-agent collaborative SLAM is expected to improve the robustness and efficiency of executing tasks for mobile agents. In this paper, a multi-agent collaborative visual-inertial SLAM system is proposed based on a centralized client-server (CS) architecture, where the clients run on smart mobiles. In general, multi-agent collaborative SLAM relies on robust and precise experience sharing and efficient communication among agents. The experience sharing requires the place recognition with a high recall and accuracy, the precise estimation of transformation between looping frames, and the map fusion with globally consistency. To this end, we devise an enhanced geometric verification, a re-projection optimization based on the error-aware weighting strategy, and a strategy of flexible fusion to meet these requirements. In addition, the multi-agent collaborative SLAM needs to exchange abundant information, which requires the efficient communication. Therefore, we design a CS collaborative loop detection mechanism which is more robust to network transmission. We perform extensive experiments on the EuRoc dataset and in real environments. Experimental results show that the proposed system achieves better results than state-of-the-art methods. Furthermore, we demonstrate the stability of the proposed collaborative SLAM in real environments with a bandwidth of 7.55Mbps.
Kaiqi Chen 0001, Ruyu Liu, Yanhong Yang, Zhenhua Wang 0003, Jianhua Zhang 0002
ICRA3
2022 Dual attention granularity network for vehicle re-identification
Jianhua Zhang 0002, Jingbo Chen, Jiewei Cao, Ruyu Liu, Linjie Bian, Shengyong Chen
Neural Comput. Appl.4
2022 Cross-Modal 360° Depth Completion and Reconstruction for Large-Scale Indoor Environment
abstract
In a large-scale epidemic, reducing direct contact among medical personnel, attendants and patients has become a necessary means of epidemic prevention and control. Intelligent vehicles and mobile robots in the hospital environment, such as disinfection vehicles, logistics vehicles, nursing robots, and guiding robots, play an important role in improving the operational efficiency of the medical system and promoting epidemic prevention and governance. Powerful capabilities of environmental spatial perception and reconstruction are the keys to accurate localization, navigation, and obstacle avoidance for intelligent vehicles and autonomous robots in such operations. Omnidirectional perception is becoming increasingly important and proliferative in autonomous vehicles and robots since its wide field of view significantly enhances the perception ability. However, the lack of dense and accurate 360° depth datasets has brought the challenge to the omnidirectional perception. In this paper, we propose a depth-sensing and reconstruction system to address this challenge in the large-scale indoor environment. First, we design an omnidirectional depth completion convolutional neural network model, in which a spherical normalized convolutional and the unit sphere area-based loss are introduced to extract features from cross-modal omnidirectional input with unequal sparsity and deal with the imbalanced data distribution and distortion in the panoramic input. In addition, we present a 3D reconstruction system by integrating our depth completion into omnidirectional localization and dense mapping. We evaluate our method on 360D large-scale indoor datasets and real-world sequences of a challenging hospital scene. Extensive experiments show that the proposed method outperforms the other state-of-the-art (SoTA) approaches in terms of depth completion and 3D reconstruction.
Ruyu Liu, Guodao Zhang, Jiangming Wang, Shuwen Zhao
IEEE Trans. Intell. Transp. Syst.1
2021 Collaborative Visual Inertial SLAM for Multiple Smart Phones
abstract
The efficiency and accuracy of mapping are crucial in a large scene and long-term AR applications. Multi-agent cooperative SLAM is the precondition of multi-user AR interaction. The cooperation of multiple smart phones has the potential to improve efficiency and robustness of task completion and can complete tasks that a single agent cannot do. However, it depends on robust communication, efficient location detection, robust mapping, and efficient information sharing among agents. We propose a multi-intelligence collaborative monocular visual-inertial SLAM deployed on multiple ios mobile devices with a centralized architecture. Each agent can independently explore the environment, run a visual-inertial odometry module online, and then send all the measurement information to a central server with higher computing resources. The server manages all the information received, detects overlapping areas, merges and optimizes the map, and shares information with the agents when needed. We have verified the performance of the system in public datasets and real environments. The accuracy of mapping and fusion of the proposed system is comparable to VINS-Mono which requires higher computing resources.
Ruyu Liu, Kaiqi Chen 0001, Jianhua Zhang 0002, Dongyan Guo
ICRA2
2021 Map Recovery and Fusion for Collaborative Augment Reality of Multiple Mobile Devices
abstract
The map recovery and fusion is a key issue in the application of large scale and long-term augmented reality (AR) scenarios. However, they are still not addressed well in an efficient and precise way, especially for complex industrial environments. In this article, we propose a map recovery and fusion strategy based on vision-inertial simultaneous localization and mapping. We first develop a heuristic strategy that can fast search and match map points among multiple maps, and can be used for efficient map fusion. For map recovery, we leverage the inertial sensors for short time motion estimation, and transform the previous lost map to the current map. Based on this strategy, a novel framework for collaborative AR is implemented and can parallelly run in multiple mobile devices in real time. Extensive experiments have been carried out on a public data set, and the results show that the proposed method can recovery and fuse multiple maps with high completeness and precision.
Jianhua Zhang 0002, Kaiqi Chen 0001, Zhiying Pan, Ruyu Liu, Thomas Yang 0001, Shengyong Chen
IEEE Trans. Ind. Informatics5
2020 CalibRCNN: Calibrating Camera and LiDAR by Recurrent Convolutional Neural Network and Geometric Constraints
abstract
In this paper, we present Calibration Recurrent Convolutional Neural Network (CalibRCNN) to infer a 6 degrees of freedom (DOF) rigid body transformation between 3D LiDAR and 2D camera. Different from the existing methods, our 3D-2D CalibRCNN not only uses the LSTM network to extract the temporal features between 3D point clouds and RGB images of consecutive frames, but also uses the geometric loss and photometric loss obtained by the interframe constraint to refine the calibration accuracy of the predicted transformation parameters. The CalibRCNN aims at inferring the correspondence between projected depth image and RGB image to learn the underlying geometry of 2D-3D calibration. Thus, the proposed calibration model achieves a good generalization ability to adapt to unknown initial calibration error ranges, and other 3D LiDAR and 2D camera pairs with different intrinsic parameters from the training dataset. Extensive experiments have demonstrated that our CalibRCNN can achieve state-of-the-art accuracy by comparison with other CNN based methods.
Jieying Shi, Ziheng Zhu, Jianhua Zhang 0002, Ruyu Liu, Zhenhua Wang 0003, Shengyong Chen, Honghai Liu 0001
IROS4
2019 Towards SLAM-Based Outdoor Localization using Poor GPS and 2.5D Building Models
abstract
In this paper, we address the topic of outdoor localization and tracking using monocular camera setups with poor GPS priors. We leverage 2.5D building maps, which are freely available from open-source databases such as OpenStreetMap. The main contributions of our work are a fast initialization method and a non-linear optimization scheme. The initialization upgrades a visual SLAM reconstruction with an absolute scale. The non-linear optimization uses the 2.5D building model footprint, which further improves the tracking accuracy and the scale estimation. A pose optimization step relates the vision-based camera pose estimation from SLAM to the position information received through GPS, in order to fix the common problem of drift. We evaluate our approach on a set of challenging scenarios. The experimental results show that our approach achieves improved accuracy and robustness with an advantage in run-time over previous setups.
Ruyu Liu, Jianhua Zhang 0002, Shengyong Chen, Clemens Arth
ISMAR1
2019 Hierarchical Topic Model Based Object Association for Semantic SLAM
abstract
Object-based simultaneous localization and mapping (SLAM) is a more natural and robust way for agents to interact with their surrounding environment. However, it introduces a problem of semantic objects association. Correct object association is the key factor to achieve a successful object SLAM system because object association and SLAM are inherently coupled and have not been well tackled yet. A novel formulation of the object association problem based on a hierarchical Dirichlet process (HDP) is proposed. Through the HDP, we can hierarchically associate the grouped object measurements. This can improve the object association accuracy and computation efficiency. Thanks to the novel formulation, the proposed method is also able to correct failure object associations according to its sampling inference algorithm. Furthermore, we introduce object poses to the processing of pose optimization. The object association and pose optimization are then solved in a tightly coupled way, by which both aspects can promote each other. The proposed method is evaluated on indoor and outdoor datasets and the experimental results show a very impressive improvement with respect to the traditional SLAM.
Jianhua Zhang 0002, Mengping Gui, Ruyu Liu, Junzhe Xu 0001, Shengyong Chen
IEEE Trans. Vis. Comput. Graph.4
2018 Instant SLAM Initialization for Outdoor Omnidirectional Augmented Reality
abstract
The initialization and absolute scale are two critical issues for an Augmented Reality (AR) system. Most existing methods have to resort to some external sensors or some special steps, to initialize an AR system and to obtain a correct scale. In this paper, we introduce an omnidirectional AR system, which can be instantly initialized, recover the absolute scale without other sensor data, and provide a full 360-degrees field-of-view (Fov) to give users the best immersive feelings. Our system shows how to firstly estimate the absolute orientation using the meaningful line and point cues from a single panorama, and how to estimate the camera global position by aligning a panorama after semantic segmentation with a widely available 2.5D map. Based on resulting absolute pose from a single frame, we subsequently render a depth map to initialize a SLAM system. We can then fuse the virtual elements with a real scene according to the continuous camera motion from the SLAM system. We evaluate the SLAM initialization approach on a challenging dataset. The experiments indicate that localization precision from our method is obviously superior to that from consumer GPS devices and we remain unbeatable in time performance compared to previous methods.
Ruyu Liu, Jianhua Zhang 0002, Kejie Yin, Jia-xin Wu, Ruihao Lin, Shengyong Chen
CASA1
2018 Cross Modal Multiscale Fusion Net for Real-time RGB-D Detection
abstract
This paper presents a novel multi-modal CNN architecture for object detection by exploiting complementary input cues in addition to sole color information. Our one-stage architecture fuses the multiscale mid-level features from two individual feature extractor, so that our end-to-end net can accept cross modal streams to obtain high-precision detection results. In comparison to other cross modal fusion neural networks, our solution successfully reduces runtime to meet the real-time requirement with still high-level accuracy. Experimental evaluation on challenging NYUD2 dataset shows that our network achieves 49.1% mAP, and processes images in real-time at 35.3 frames per second on one single Nvidia GTX1080 GPU. Compared to baseline one stage network SSD on RGB images which gets 39.2% mAP, our method has great accuracy improvement.
Kejie Yin, Sheng Liu 0002, Ruyu Liu, Yibin Chen
ICPR3
2018 Absolute Orientation and Localization Estimation from an Omnidirectional Image
Ruyu Liu, Jianhua Zhang 0002, Kejie Yin, Zhiying Pan, Ruihao Lin, Shengyong Chen
PRICAI1