Jianhua Zhang 0002

dblp:55/2363-2 · DBLP profile ↗
← Back
83ranked-venue papers
16as first author
58since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 32 · 8 first-author · 19 since 2021Graphics, computer vision, multimedia, augmented reality and games · 24 · 4 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 23 · 5 first-author · 20 since 2021Human-computer interaction and ubiquitous computing · 11 · 10 since 2021Systems, architecture and hardware · 7 · 2 first-author · 3 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 SMISMP: Stable Scoring for Few-Shot 3D Anomaly Detection via Stochastic Metric Modeling
Jianhua Zhang 0002
ICIC (20)2
2026 BTKD++: Beyond Teachers by Critically Distilling Knowledge from Teacher's Bias
abstract
Abstract Existing knowledge distillation methods indiscriminately transfer knowledge from teacher networks, including output-level decisional biases, i.e., incorrect final predictions that can mislead student learning and limit student performance. We challenge this paradigm by proposing BTKD++, a framework that systematically filters and rectifies teacher’s output-level biased knowledge into corrective signals. Our approach partitions training data into Easy Tasks (correct teacher predictions) and Hard Tasks (incorrect predictions), then applies bias elimination and rectification modules orchestrated by dynamic learning curriculum. We provide an interpretive information-theoretic abstraction to explain the observed competence-threshold phenomenon, under which bias rectification becomes more effective when teacher errors contain sufficiently structured corrective information. BTKD++ demonstrates broad applicability across classification, detection, and segmentation tasks when task outputs are equipped with suitable probabilistic interfaces, and shows consistent effectiveness across CNNs, Transformers, and State-Space Models. Extensive experiments show consistent student-teacher transcendence, establishing new state-of-the-art results. This work redefines knowledge distillation from blind mimicry to critical learning, proving that students can surpass teachers through principled bias correction. The source code is available at https://github.com/smartyige/BTKD .
Jianhua Zhang 0002, Yu He 0001, Xu Cheng 0003, Xiufeng Liu 0001, Shengyong Chen, Houxiang Zhang, Ruyu Liu
Int. J. Comput. Vis.1
2026 Object shape differentiation and texture rendering for neural implicit SLAM
Jiaming Lu, Ruyu Liu, Jianhua Zhang 0002, Xu Cheng 0003
Mach. Vis. Appl.3
2026 Multi-branch perturbation learning with constraint simulation for semi-supervised semantic segmentation
abstract
Current semi-supervised semantic segmentation (SSS) methods improve generalization via weak-to-strong pseudo-supervision with image perturbations. However, many methods are limited by employing a single perturbation mode and a specific weak-to-strong learning strategy, restricting exploration of the perturbation space and hindering performance in fine-grained segmentation. While diverse perturbations are intuitively beneficial, simply combining them can lead to inefficient optimization and instability. In this paper, we propose a multi-branch strong perturbation constraint learning framework for SSS. Our framework introduces a novel multi-branch perturbation learning (MSPL) strategy, employing multiple parallel branches with diverse strong augmentations to expand the perturbation space and capture complex semantic variations. We further design a novel constraint simulation loss (CSSL), based on a hierarchical consistency learning structure (weak-to-strong and strong-to-strong), which enforces strong-to-strong consistency between different perturbation branches. CSSL mitigates instability and enhances robustness to perturbation-induced noise, enabling the network to better generalize and achieve more accurate segmentation, especially for fine object boundaries. Extensive evaluations on benchmark datasets (PASCAL VOC 2012, Cityscapes, COCO) demonstrate that our method achieves state-of-the-art performance. Ablation studies further validate the effectiveness of our proposed MSPL and CSSL components.
Ruyu Liu, Feng Xiao 0005, Jianhua Zhang 0002, Xiufeng Liu 0001, Xu Cheng 0003, Shengyong Chen, Houxiang Zhang
Pattern Recognit.3
2026 PIVOT: A Framework for High-Fidelity Facade PV Layouts via Generative Seeding and Differentiable Optimization
abstract
The automation of Building-Integrated Photovoltaic (BIPV) design is essential for the digital transformation of urban renewable energy, yet current methods struggle to bridge the gap between semantic facade assessment and engineering-grade installation plans. Existing AI-based approaches typically yield raw segmentation masks that lack the geometric precision and physical validity required for real-world deployment. This paper presents PIVOT, a novel end-to-end framework for the automated synthesis of high-fidelity, collision-free PV layouts from single 2D facade images. We formulate the layout generation as a two-stage neuro-symbolic and differentiable optimization process. First, an LLM-based programmatic seeding stage uses structured facade information and elementary spatial-relation cues to generate initial layout hypotheses. Second, a differentiable layout optimization engine refines these proposals by treating panel placement as a continuous constrained optimization problem, effectively resolving spatial overlaps while maximizing geometric regularity. Validated on a diverse dataset of 80 building facades, experimental results demonstrate that PIVOT significantly outperforms heuristic (MaxRects) and evolutionary (MOGA) baselines. Our framework reduces the layout collision rate from 9.7% to 0.7% and improves geometric regularity by over 10%, while requiring only approximately 80 seconds of processing time per building. This work advances BIPV design automation by converting semantically parsed facade information into geometrically constrained, constructibility-oriented PV layouts for downstream performance evaluation.
Dongxu Zhuang, Ruyu Liu, Xiufeng Liu 0001, Per Sieverts Nielsen, Tangao Hu, Jianhua Zhang 0002
IEEE Trans Autom. Sci. Eng.6
2026 Adaptive Kernel Selection Module Combined With Feature Enhanced Perception Network for Camouflaged Object Detection
abstract
Camouflaged object detection plays a crucial role in applications such as automatic sorting and defect inspection in industrial production, yet existing methods often struggle to flexibly capture features of diverse shapes, orientations, and scales due to their reliance on fixed receptive fields and rigid windowing schemes. To address these limitations, we propose a dual-branch joint network comprising a reference branch and a segmentation branch. The reference branch learns supplementary cues from salient objects that co-occur with camouflaged targets, guiding the segmentation branch toward more accurate delineation. Within the segmentation branch, we introduce three novel modules: 1) a deformable window interaction mechanism that replaces fixed-size transformer windows with learnable quadrilateral windows to adaptively extract features of arbitrary shape and orientation; 2) a feature enhancement perception module that fuses rich multiscale representations through parallel dilated convolutions at varying rates and channel-/spatial-attention mechanisms; and 3) a receptive field adjustment adaptive module that dynamically adjusts its receptive field size to balance sensitivity to fine details and global context. Comprehensive experiments on COD10 K, NC4K, CAMO, and R2C7K benchmarks demonstrate that our model outperforms the majority of current state-of-the-art approaches, while ablation studies and sensitivity analyses confirm the individual and combined effectiveness of our proposed components.
Ruyu Liu, Feng Xiao 0005, Jianhua Zhang 0002, Shengyong Chen
IEEE Trans. Ind. Informatics4
2026 WTCLIP: A Wavelet-Aware CLIP Framework for Boundary-Refined Weakly Supervised Semantic Segmentation
abstract
Some advanced methods have leveraged the zero-shot recognition capability of the contrastive language–image pretraining (CLIP) model and adapted it to weakly supervised semantic segmentation (WSSS), achieving promising performance. However, they primarily use CLIP as an auxiliary feature extractor, leaving the fundamental limitations of class activation mapping unresolved, particularly in preserving fine-grained object boundaries and achieving precise pixelwise localization under sparse supervision. To address these challenges, this article proposes a novel end-to-end WSSS framework WTCLIP, which aims to fully exploit the potential of CLIP for weakly supervised segmentation tasks. Different from traditional methods that use CLIP only as a static feature extractor, we innovatively introduce a learnable wavelet transform decoder to enhance the information extraction capability and significantly improve the model's perception of object boundaries. We dynamically adjust the weight distribution ratio of the CLIP feature layer, capture multiscale edge information, and make full use of the time–frequency localization characteristics of the wavelet transform to significantly improve the quality of pseudolabels and achieve more accurate semantic segmentation. Experimental results show that our method significantly improves the performance of the WSSS task on two public benchmark datasets, notably by4.0%over the state-of-the-art methods, especially in capturing weakly annotated object boundary details.
Feng Xiao 0005, Jianhua Zhang 0002, Peihua Han, Shengyong Chen, Houxiang Zhang
IEEE Trans. Ind. Informatics2
2026 TSCFNet: Temporal Spectral Feature Cross Fusion Network for Imbalanced Sea State Estimation in Autonomous Ships
abstract
Sea state estimation (SSE) is critical to the safety of maritime transport and the reliability of autonomous ships. The frequency of different sea states varies significantly, leading to uneven data distribution. Existing deep learning methods for SSE typically focus on feature extraction, often using simple splicing and fusion, which can result in cross-domain incoherence and degrade model performance. Addressing sea state classification imbalance is often done through distance-based classifiers (e.g., prototype classifiers), but these can be less sensitive to minority classes, and using few prototypes for a class limits the expression of intra-class variations. To overcome these challenges, we propose the Temporal Spectral Cross Fusion Network (TSCFNet), which extracts temporal and spectral features. These are integrated via an innovative temporal spectral cross fusion module to maximize their complementary advantages. Additionally, we introduce a multi-fusion loss function, including temporal, spectral, and fusion losses, to optimize features across different dimensions. This approach improves the performance for minority classes and captures intra-class differences more effectively, solving the problem of category imbalance. Experimental results show that TSCFNet significantly outperforms baseline methods on two imbalanced sea state datasets and multiple multivariate spatio-temporal datasets.
Feng Xiao 0005, Xu Cheng 0003, Xia Xie 0003, Jianhua Zhang 0002
IEEE Trans. Intell. Transp. Syst.5
2025 Can Students Beyond the Teacher? Distilling Knowledge from Teacher's Bias
abstract
Knowledge distillation (KD) is a model compression technique that transfers knowledge from a large teacher model to a smaller student model to enhance its performance. Existing methods often assume that the student model is inherently inferior to the teacher model. However, we identify that the fundamental issue affecting student performance is the bias transferred by the teacher. Current KD frameworks transmit both right and wrong knowledge, introducing bias that misleads the student model. To address this issue, we propose a novel strategy to rectify bias and greatly improve the student model's performance. Our strategy involves three steps: First, we differentiate knowledge and design a bias elimination method to filter out biases, retaining only the right knowledge for the student model to learn. Next, we propose a bias rectification method to rectify the teacher model's wrong predictions, fundamentally addressing bias interference. The student model learns from both the right knowledge and the rectified biases, greatly improving its prediction accuracy. Additionally, we introduce a dynamic learning approach with a loss function that updates weights dynamically, allowing the student model to quickly learn right knowledge-based easy tasks initially and tackle hard tasks corresponding to biases later, greatly enhancing the student model's learning efficiency. To the best of our knowledge, this is the first strategy enabling the student model to surpass the teacher model. Experiments demonstrate that our strategy, as a plug-and-play module, is versatile across various mainstream KD frameworks.
Jianhua Zhang 0002, Ruyu Liu, Xu Cheng 0003, Houxiang Zhang, Shengyong Chen
AAAI1
2025 Robust Online Detection of Anomalies in Evolving Data Streams with GAN Imputation
abstract
Online anomaly detection is a critical technique in intelligent industrial systems. Concept drift and missing data that arise during the transmission of real-time data streams can severely impact the performance of anomaly detection. Therefore, ensuring the accuracy of online anomaly detection and adaptability to concept drift in the presence of missing values is a significant challenge. In this paper, we propose an online anomaly detection model that integrates an online imputation module based on Generative Adversarial Networks (GANs) to efficiently impute missing data. Additionally, following the dynamic model pool strategy, the model dynamically selects and adjusts the optimal models in the pool to effectively respond to changes in data distribution, ensuring detection performance under complex conditions. We evaluated the model's performance across multiple datasets with concept drift and under varying missing data rates. The results demonstrate that the proposed model not only adapts flexibly to rapidly changing data streams but also exhibits enhanced robustness in the presence of missing data.
Mengna Liu, Xu Cheng 0003, Jianhua Zhang 0002, Feng Xiao 0005
CSCWD4
2025 Data-Driven Diffusion-Augmented Network for Imbalanced Sea State Estimation with Dynamic Prototypes
abstract
With the advent of Industry 4.0, which emphasizes automation and data-driven decision-making, sea state estimation (SSE) has become a critical component in marine engineering and autonomous vessels. However, traditional SSE methods face significant challenges due to poor real-time performance and high manual costs. Additionally, existing deep learning (DL) approaches struggle to handle imbalanced sea state data effectively, which hinders their generalization ability. To address these issues, this paper proposes a novel DL-based SSE model. The model enhances the expressive capacity of the extracted features through a feature augmentation module and utilizes a diffusion model to extract more robust features from the data. Furthermore, a dynamic prototype update module is designed to effectively address class imbalance and overcome the limitations of traditional prototype classifiers in handling boundary samples. Experimental results show that the proposed method outperforms existing baseline models on two imbalanced sea state datasets, achieving F1-score improvements of 4.5% and 3.5%, respectively, compared to the current state-of-the-art methods. Additionally, the model demonstrates superior performance on several publicly available multivariate time series classification datasets.
Feng Xiao 0005, Xu Cheng 0003, Jianhua Zhang 0002
CSCWD4
2025 Lightweight Self-Supervised Monocular Depth Estimation via Context-Aware Fusion and Separable Depthwise Convolution
abstract
Monocular depth estimation is a critical problem in computer vision, with wide-ranging applications across various domains. However, existing methods often involve high computational costs, making them challenging to deploy efficiently on edge devices. To address the trade-off between computational complexity and inference accuracy, this paper presents an efficient and lightweight model for self-supervised monocular depth estimation. Our model incorporates a Context-Aware Fusion (CAF) module to capture both global and local feature dependencies. In addition, Separable Depthwise Convolution (SDC) module are utilized to reduce computational overhead, and the Multi-Scale Structural Similarity (MS-SSIM) loss function is employed to improve both depth estimation accuracy and visual perception quality. Experimental results show that the proposed model delivers improved accuracy while maintaining a lightweight and efficient architecture, ensuring its compatibility with edge device deployment. A comprehensive analysis of the findings is also presented, along with insights for future optimizations and improvements.
Meina Zhao, Shixin Wang 0014, Feng Xiao 0005, Jianhua Zhang 0002, Xu Cheng 0003, Yunrui Zhu
CSCWD4
2025 QCTKD-PU: Quantum Convolutional Transformer with Knowledge Distillation for Efficient and Robust Point Cloud Upsampling
abstract
Point cloud upsampling is crucial for high-fidelity 3D reconstruction in real-time applications such as autonomous systems. Existing methods based on CNNs or Transformers face three limitations: (1) prohibitive computational complexity hindering real-time deployment, (2) insufficient modeling of multi-scale geometric dependencies in sparse data, (3) sensitivity to noise and outliers. To address these challenges, we propose QCTKD-PU, a framework integrating Quantum Convolutional Transformers (QCT) and Knowledge Distillation (KD) for Point cloud Upsampling. The QCT leverages quantum superposition and self-attention to encode high-dimensional features, enabling efficient multi-scale point interaction learning. Simultaneously, KD transfers knowledge from a teacher model to a lightweight student network, reducing computational costs while maintaining accuracy. Experiments on benchmark datasets demonstrate superior performance in geometric accuracy and noise robustness compared to state-of-the-art methods. This work pioneers the synergy of quantum computing and lightweight learning for resource-constrained 3D vision tasks, while the student model achieves real-time and compact deployment, offering a practical solution for collaborative edge systems.
Yunrui Zhu, Ruyu Liu, Xu Cheng 0003, Jianhua Zhang 0002, Shengyong Chen
CSCWD5
2025 SPRGAN: Streamlined Progressive Refinement for Adversarial Point Cloud Video Upsampling
abstract
Getting dense, uniform, time-series point cloud data is critical for effective rendering. However, due to the limited computational power of edge devices, existing methods cannot achieve real-time results, which affects the visual quality of the consumer experience. To effectively address this issue, this paper presents a self-supervised adversarial upsampling method for point cloud video streams called SPR-GAN. In the generator, we design the Temporal Iterative Graph module to learn local features for each frame and captures long-range spatial information using three iterations of graph convolution operations. Then the Contextual Temporal Fusion module is developed to merge information between different frames, synthesizing temporal information and enriching the dynamic feature representation of the point cloud. Meanwhile, in the discriminator, we introduce the Efficient Shape module. Through dynamic graph convolution operations and stacked learning, it significantly improves the resolution efficiency of global shape information in point clouds. The final experiments show that the proposed method exhibits high practicality and superiority. The model achieves a good result on both the D-FAUST and DeformingThing4D-Animals datasets.
Ruyu Liu, Xianchao Zhang 0002, Jianhua Zhang 0002, Xiufeng Liu 0001
ICASSP5
2025 Efficient Large-Scale Scene Point Cloud Upsampling with Implicit Neural Networks and Spatial Hashing
abstract
Point cloud upsampling is a critical challenge in 3D vision, particularly for large-scale, real-world data. We propose ASFNet, a novel implicit neural network-based approach that uniquely combines adaptive spatial feature representation with efficient spatial hashing. This method significantly improves both upsampling quality and computational efficiency. ASFNet first encodes the point cloud as an implicit surface, employing dynamic search and spatial hashing to optimize query point locations rapidly. This approach creates a uniform, continuous field around surfaces, enabling high-fidelity upsampling. Experiments on benchmark datasets, including Oakland 3D dataset and VMR-Oakland-v2, demonstrate ASFNet’s superiority. Our method achieves state-of-the-art performance with a Chamfer Distance of 5.559 × 10–3on Oakland 3D dataset, while reducing processing time by up to 80% compared to existing methods. On the challenging Oakland 3D dataset, ASFNet completes upsampling in just 150 seconds. These results underscore ASFNet’s potential to advance real-time 3D vision applications in areas such as autonomous navigation and augmented reality.
Yunrui Zhu, Ruyu Liu, Xu Cheng 0003, Jianhua Zhang 0002, Xiufeng Liu 0001
ICASSP5
2025 Mamba-SLAM: Enhancing Neural Implicit SLAM with Uncertainty and Mamba
abstract
Current neural implicit SLAM systems struggle with insufficient object shape constraints due to incomplete depth maps and inefficient pixel sampling strategies, leading to inaccuracies in reconstructed scene morphology and texture. To address these limitations, we introduce Mamba-SLAM, a novel framework featuring two key innovations: (1) an uncertainty-based pixel sampling module that enhances object rendering quality in depth-scarce regions by integrating depth map analysis and color image gradients, improving shape and texture consistency by 6.4% on the Replica dataset compared to state-of-the-art methods; and (2) a Mamba-based keyframe selection module that leverages the efficient feature extraction of Mamba to provide rich semantic cues, optimizing pose estimation and enriching reconstruction detail. Experiments on Replica and ScanNet demonstrate that Mamba-SLAM significantly improves scene rendering and object detail, achieving a 13.5% improvement in reconstruction completeness on Replica. The core novelty lies in the synergistic combination of uncertainty-driven pixel selection and Mamba-powered keyframe management for enhanced neural implicit SLAM.
Jiaming Lu, Yunrui Zhu, Ruyu Liu, Xu Cheng 0003, Jianhua Zhang 0002, Xiufeng Liu 0001
ICME5
2025 Multi-Scale Convolutional Networks with Class-Normalized Logit Clipping for Robust Sea State Estimation from Noisy Ship Motion Data
abstract
Autonomous ships utilize automation systems to achieve unmanned navigation, driving innovation in maritime transportation. However, sea conditions, influenced by dynamic factors such as wave height, wind speed, and ocean currents, present a challenge in accurately assessing these conditions. Traditional classification models often assume accurate labels, but noisy labels are prevalent in real-world applications. Existing methods, such as noise sample filtering or loss function adjustment, have limited applicability and poor generalization when dealing with complex sea condition data. To address this issue, this study proposes an end-to-end neural network model. The model's feature extraction module uses deep representation learning to capture latent patterns in the data, and a loss function is designed to mitigate the impact of outliers. The integration of these components allows the model to perform accurate classification even in the presence of noisy labels. Extensive experiments on public and sea condition datasets validate the effectiveness of this approach, demonstrating that the model exhibits strong generalization capabilities and holds great promise for practical applications.
Mengna Liu, Xu Cheng 0003, Xiufeng Liu 0001, Fan Shi 0001, Jianhua Zhang 0002, Shengyong Chen
ICRA6
2025 Efficiency-Optimized Point Cloud Upsampling with Single-Layer Graph Convolution Network
abstract
Point cloud upsampling is a key technology for improving the density and quality of sparse point clouds, with widespread applications in 3D reconstruction, autonomous driving, and environmental perception. However, traditional point cloud upsampling methods, especially those based on multilayer graph convolution networks (GCNs), typically rely on complex feature extraction modules, which increase computational complexity and model parameters, limiting their use in resource-constrained environments. To overcome these challenges, we propose the EO-PU framework, a lightweight and efficient point cloud upsampling method. This framework combines single-layer GCN and rotation-invariant 4D projection encoding (I4DP) technology, significantly reducing computational load and redundant information, thereby improving upsampling efficiency. Specifically, EO-PU first uses I4DP to map the 3D point cloud data to a rotation-robust 4D feature space, ensuring effective capture of geometric information. A single-layer GCN is then employed to aggregate features, reducing network complexity and computational cost. To further enhance upsampling performance, we introduce the EdgeShuffleNet module, which optimizes feature expansion and rearrangement through efficient local feature aggregation. Experimental results show that EO-PU outperforms or matches existing methods across multiple public datasets while significantly reducing model parameters and computation time, making it highly suitable for deployment in resource-constrained environments.
Yunrui Zhu, Feng Xiao 0005, HaoXiao Wang, Ruyu Liu, Xu Cheng 0003, Jianhua Zhang 0002
IJCNN7
2025 PolypSense3D: A Multi-Source Benchmark Dataset for Depth-Aware Polyp Size Measurement in Endoscopy
abstract
Accurate polyp sizing during endoscopy is crucial for cancer risk assessment but is hindered by subjective methods and inadequate datasets lacking integrated 2D appearance, 3D structure, and real-world size information. We introduce PolypSense3D, the first multi-source benchmark dataset specifically targeting depth-aware polyp size measurement. It uniquely integrates over 43,000 frames from virtual simulations, physical phantoms, and clinical sequences, providing synchronized RGB, dense/sparse depth, segmentation masks, camera parameters, and millimeter-scale size labels derived via a novel forceps-assisted in-vivo annotation technique. To establish its value, we benchmark state-of-the-art segmentation and depth estimation models. Results quantify significant domain gaps between simulated/phantom and clinical data and reveal substantial error propagation from perception stages to final size estimation, with the best fully automated pipelines achieving an average Mean Absolute Error (MAE) of 0.95 mm on the clinical data subset. Publicly released under CC BY-SA 4.0 with code and evaluation protocols, PolypSense3D offers a standardized platform to accelerate research in robust, clinically relevant quantitative endoscopic vision. The benchmark dataset and code are available at: https://github.com/HNUicda/PolypSense3D and https://doi.org/10.7910/DVN/K13H89.
Ruyu Liu, Mingming Zhou, Jianhua Zhang 0002, Xiufeng Liu 0001, Xu Cheng 0003, Sixian Chan 0001, Yanbin Shen, Sheng Dai, Yuping Yan, Yaochu Jin, Lingjuan Lyu
NeurIPS4
2025 Watermark Removal via Boundary-Aware Segmentation and Semantic-Guided Diffusion
Zhenjie Jiang, Ruyu Liu, Jianhua Zhang 0002, Mohammed M. Elmogy, Shengyong Chen
PRCV (2)6
2025 An expert features enhanced temporal and contextual contrasting learning model for detecting wind turbine blade icing
abstract
With global carbon neutrality goals, wind power has rapidly developed, but blade icing remains a major challenge. AI(artificial intelligence) methods show great promise for detecting icing on wind turbine blades. However, early icing data overlap, difficulty obtaining continuous labeled data, and small variations between samples due to short sampling intervals complicate the task. This study proposes an expert feature-enhanced temporal and contextual contrastive learning model for detecting blade icing. This approach efficiently extracts data features and combines self-supervised contrastive learning, maximizing data utilization without requiring extensive labeled data. To validate the effectiveness of this method, extensive experiments were conducted on two public datasets. The results achieved the best performance across multiple metrics, with F1-Score and AUC exceeding 98%, significantly enhancing wind power generation efficiency.
Jiamei Zhou, Feng Xiao 0005, Xu Cheng 0003, Jianhua Zhang 0002
Eng. Appl. Artif. Intell.5
2025 Alice-SLAM: Accurate and Lite-Communication Collaborative SLAM for Resource-Constrained Multi-Agent
abstract
Multi-agent collaborative simultaneous localization and mapping (Mac-SLAM) facilitates mutual localization among multi-agent and mapping in unknown environments. However, Mac-SLAM faces two main practical challenges in resource-constrained situations: heavy communication load and conflicts among multi-source maps. To address these issues, we propose Alice-SLAM: an accurate and lite-communication client-server collaborative SLAM system, reducing communication load while accuracy-guaranteed. Specifically, regarding high communication demand, we optimize communication load by compressing keyframe data and sharing only key map information instead of full map information. For inconsistency among multi-maps, we combine specific bundle adjustments (BA) and an adaptive strategy for active map optimization to enhance the consistency of the global map. A set of experiments demonstrates the superior accuracy and reduced communication load of the proposed Alice-SLAM on the EuRoC dataset and in multi-user augmented reality (AR) experiments conducted in our lab, highlighting its effectiveness in resource-constrained cases. We plan to open-source our code1to encourage further research and collaboration in this area.
Kaiqi Chen 0001, Ruyu Liu, Xu Cheng 0003, Jianhua Zhang 0002, Shengyong Chen, Houxiang Zhang, Arash Ajoudani
IEEE J. Sel. Areas Commun.5
2025 SVD-KD: SVD-based hidden layer feature extraction for Knowledge distillation
Jianhua Zhang 0002, Mian Zhou, Ruyu Liu, Xu Cheng 0003, Sasa Nikolic 0002, Shengyong Chen
Pattern Recognit.1
2025 Wavelet-Discrete Cosine Transform Synergy for Ship Motion-Based Sea State Estimation in Autonomous Ships
abstract
Developing a robust autonomous sea state estimation (SSE) model stands as a pivotal challenge in advancing autonomous ships. Presently, deep learning (DL) methodologies have showcased remarkable efficacy in SSE tasks. Nonetheless, the dynamic nature of ship motion introduces temporal variations alongside frequency domain characteristics like periodic swinging, posing challenges for existing DL approaches. Most prevailing DL techniques, predominantly leveraging Convolutional Neural Networks or Long Short-Term Memory Networks, often fail to effectively harness frequency domain information post feature extraction. To tackle these limitations head-on, this paper introduces a pioneering SSE model. Specifically, in order to solve the frequency-domain feature extraction problem, we design a wavelet transform-based frequency domain encoder to extract relevant frequency-domain features from ship motion data by discriminating the contribution of different frequencies in the signal. Subsequently, in order to better integrate the extracted ship motion features, we designed a Feature Perception module based on discrete cosine transform. This module adeptly merges the extracted feature insights while prioritizing crucial frequency domain features. Following rigorous experimentation, our methodology exhibits superior performance compared to existing baseline techniques in SSE, a capability of profound significance for autonomous ships. Moreover, across diverse public multivariate time series classification datasets, our model outperforms current state-of-the-art approaches, underscoring its scalability across distinct domains.
Feng Xiao 0005, Xu Cheng 0003, Sasa Nikolic 0002, Jianhua Zhang 0002, Shengyong Chen
IEEE Trans Autom. Sci. Eng.5
2025 Distortion-Aware Outdoor Panoramic Depth Estimation via Local-Global Fusion
abstract
Outdoor panoramic depth estimation faces significant challenges due to the wide field of view (FoV), complex scene structures, and severe distortion encountered in such environments. Traditional methods, which often use distortion convolution, fall short in capturing global distortion information and extracting rich contextual details from panoramic images. To overcome these limitations, this article introduces a novel dual-branch framework that synergistically merges the advantages of equirectangular projection (ERP) and tangent projection (TP). First, we design a unique dual-branch framework specifically tailored for panoramic depth estimation. In this framework, the convolutional neural networks branch processes ERP images to extract rich local information, enhancing the detail accuracy of depth estimation, while the vision transformers branch processes TPs to capture comprehensive global information, improving the smoothness of depth estimation. Then, we further enhance our method with a distortion-aware weight map module that adapts the influence of different image regions according to their distortion level, thus prioritizing features from areas with less distortion. In addition, we implement a dual attention fusion module to seamlessly integrate features from both branches at corresponding layers. Comprehensive experiments across various outdoor datasets reveal that our method significantly outperforms state-of-the-art techniques in terms of depth estimation accuracy, adeptly balancing the capture of both overarching scene depth and intricate details, potentially revolutionizing applications in industrial informatics, such as autonomous navigation and environmental mapping.
Ruyu Liu, Yihao Ying, Xiufeng Liu 0001, Weiguo Sheng 0001, Jianhua Zhang 0002, Shengyong Chen
IEEE Trans. Ind. Informatics7
2025 Cascaded State Space and Contrastive Learning for Cross-Domain Few-Shot Segmentation
abstract
Current cross-domain few-shot semantic segmentation (CD-FSS) faces multiple challenges, including inconsistent feature mapping among domains and insufficient utilization of low-level information and background information from the source domain. To address these issues, this article proposes a novel cascade feature enhancement and contrastive learning framework to improve the generalization capability of CD-FSS. Within this framework, we first introduce a cascade feature enhancement module to construct distinctive feature representations, enhancing the model’s transferability across domains. By effectively integrating multilevel feature information from support images, this module strengthens the representation capability of query images. Second, we employ contrastive learning to form positive and negative sample pairs for the foreground and background, capturing rich correlations between them. Finally, the iterative prototype enhancement module we propose gradually refines the correspondence between the support image and the query image through iteration, making full use of the embedded supervisory information in the limited support samples. Experimental results demonstrate that the proposed method outperforms existing approaches on multiple benchmark datasets, achieving up to a 9.7% improvement over state-of-the-art methods.
Feng Xiao 0005, Jianhua Zhang 0002, Peihua Han, Shengyong Chen, Houxiang Zhang
IEEE Trans. Ind. Informatics2
2025 Cross-Scale Denoising Reverse Distillation for Anomaly Detection
abstract
Effective discrepancy representation of anomalies plays a crucial role in visual anomaly detection. Recent advances build upon reverse distillation paradigm that boost the teacher–student model’s discrimination capability on anomalies; however, they are still susceptible to the size variation of unpredictable anomalies. To generalize the anomaly size variation, we propose a new algorithm cross-scale denoising reverse distillation (CDRD), which integrates cross-scale denoising with reverse distillation to exchange multiscale perception and enhance the fine-grained representation of features. Specifically, we introduce a cross-scale anomalous signal suppression procedure in the teacher network to facilitate the interaction of information across different scales, thereby enabling the student network to learn more robust normal data representations. In the knowledge transfer process, a fusion compression module acts as an intermediate transmitter of information, aiming to obtain a compact embedding while abandoning anomaly perturbations. Moreover, we construct a detail supplement module in the student network to prevent the loss of key information in the deconvolution process of the decoder. Experiments on well-known datasets demonstrate that our CDRD brings significant improvements over the next best competitor.
Yanhong Yang, Feng Xiao 0005, Jianhua Zhang 0002, Guodao Zhang, Shengyong Chen
IEEE Trans. Ind. Informatics4
2025 Semantic Visual Simultaneous Localization and Mapping: A Survey
abstract
Visual Simultaneous Localization and Mapping (vSLAM) is a cornerstone technology in computer vision and robotics, underpinning applications such as autonomous vehicles and robot navigation. While traditional vSLAM systems have shown significant progress in indoor or outdoor environments, their performance often degrades in complex scenes, limiting their adaptability and robustness. Semantic vSLAM, which integrates high-level semantic information into vSLAM systems, has emerged as a promising solution to address these limitations by enabling a richer understanding of the environment. In this paper, we provide a comprehensive review of semantic vSLAM, offering a critical analysis of its evolution, methods, and challenges. We begin by revisiting the development of traditional vSLAM, emphasizing its limitations and the motivation for incorporating semantic information. Subsequently, we delve into the core modules of semantic vSLAM, including semantic extraction, object association, semantic loop closing, back-end optimization, and semantic mapping. Then, we present a performance comparison of semantic vSLAM systems under two different datasets, indoor and outdoor, respectively. Furthermore, we also provide a comparative analysis of widely used SLAM datasets to provide guidance for performance testing and validation. To further enrich the discussion, we identify unresolved challenges in semantic vSLAM, such as long-term semantic perception and association, open and unstructured environments. We propose future research directions, including balancing computational resources and quantifying system risk, large model-based navigation and mapping, and embodied AI SLAM. By providing key insights and forward-looking perspectives, this work aims to stimulate future research and improve the capabilities of semantic vSLAM in real-world applications.
Kaiqi Chen 0001, Junhao Xiao 0001, Qiyi Tong, Heng Zhang 0023, Ruyu Liu, Jianhua Zhang 0002, Arash Ajoudani, Shengyong Chen
IEEE Trans. Intell. Transp. Syst.7
2025 Covariance Propagation-Based Accurate Loop Detection for High Confusion Environment
Kaiqi Chen 0001, Ruyu Liu, Shengyong Chen, Arash Ajoudani, Jianhua Zhang 0002
IEEE Trans. Intell. Transp. Syst.6
2025 Graph Attention Network for Context-Aware Visual Tracking
abstract
Siamese-network-based trackers convert the general object tracking as a similarity matching task between a template and a search region. Using convolutional feature cross correlation (Xcorr) for similarity matching, a large number of Siamese trackers are proposed and achieved great success. However, due to the predefined size of the target feature, these trackers suffer from either retaining much background information or losing important foreground information. Moreover, the global matching between the target and search region also largely neglects the part-level structural information and the contextual information of the target. To tackle the aforementioned obstacles, in this article, we propose a simple context-aware Siamese graph attention network, which establishes part-to-part correspondence between the Siamese branches with a complete bipartite graph. The object information from the template is propagated to the search region via a graph attention mechanism. With such a design, a target-aware template input is enabled to replace the prefixed template region, which can adaptively fit the size and aspect ratio variations in different objects. Based on it, we further construct a context-aware feature matching mechanism to embed both the target and the contextual information in the search region. Experiments on challenging benchmarks including GOT-10k, TrackingNet, LaSOT, VOT2020, and OTB-100 demonstrate that the proposed SiamGAT* outperforms many state-of-the-art trackers and achieves leading performance. Code is available at: https://git.io/SiamGAT.
Yanyan Shao, Dongyan Guo, Zhenhua Wang 0003, Liyan Zhang 0001, Jianhua Zhang 0002
IEEE Trans. Neural Networks Learn. Syst.6
2024 Intentional Evolutionary Learning for Untrimmed Videos with Long Tail Distribution
abstract
Human intention understanding in untrimmed videos aims to watch a natural video and predict what the person’s intention is. Currently, exploration of predicting human intentions in untrimmed videos is far from enough. On the one hand, untrimmed videos with mixed actions and backgrounds have a significant long-tail distribution with concept drift characteristics. On the other hand, most methods can only perceive instantaneous intentions, but cannot determine the evolution of intentions. To solve the above challenges, we propose a loss based on Instance Confidence and Class Accuracy (ICCA), which aims to alleviate the prediction bias caused by the long-tail distribution with concept drift characteristics in video streams. In addition, we propose an intention-oriented evolutionary learning method to determine the intention evolution pattern (from what action to what action) and the time of evolution (when the action evolves). We conducted extensive experiments on two untrimmed video datasets (THUMOS14 and ActivityNET v1.3), and our method has achieved excellent results compared to SOTA methods. The code and supplementary materials are available at https://github.com/Jennifer123www/UntrimmedVideo.
Xiujie Wang, Jianhua Zhang 0002, Shengyong Chen
AAAI3
2024 360ORB-SLAM: A Visual SLAM System for Panoramic Images with Depth Completion Network
abstract
With the advent of the Industry 4.0 era and the increasing performance requirements for AR/VR applications and vision assistance and inspection systems in recent years, visual simultaneous localization and mapping (vSLAM) is a fundamental task in computer vision and robotics. However, traditional vSLAM systems are limited by the camera’s narrow field-of-view, resulting in challenges such as sparse feature distribution and lack of dense depth information. To overcome these limitations, this paper proposes a 360ORB-SLAM system for panoramic images that combines with a depth completion network. The system extracts feature points from the panoramic image, utilizes a panoramic triangulation module to generate sparse depth information, and employs a depth completion network to obtain a dense panoramic depth map. Experimental results on our novel panoramic dataset constructed based on Carla demonstrate that the proposed method achieves superior scale accuracy compared to existing monocular SLAM methods and effectively addresses the challenges of feature association and scale ambiguity. The integration of the depth completion network enhances system stability and mitigates the impact of dynamic elements on SLAM performance.
Yuqi Pan, Ruyu Liu, Guodao Zhang, Jianhua Zhang 0002
CSCWD7
2024 CRED: A Corneal Reflection and Environment Dataset
abstract
Existing studies have proved that corneal reflection images can not only visualize the human surroundings, but also accurately reflect the attention information of the eyes to the environment, which promotes the research and application of visual tracking and human posture localization in the field of human-computer interaction. The cornea is a small and transparent reflective surface with weak reflective ability, and its reflected images always have dull colors and low resolution. Some researches try to obtain clearer and brighter corneal reflection images when a person is facing a screen or outdoors, but the reflected images are highly susceptible to the interference of iris color and texture. However, strong corneal reflections are highly susceptible to obscuring the iris and pupil regions, affecting the accuracy of gaze tracking. In addition, a large number of reflected images interfered by the iris must rely on iris features for image enhancement. These two limitations make it difficult to directly apply eye images taken outdoors. We try to propose a corneal reflection and human eye surroundings dataset, CRED, which contains not only segmented images of human eye images and ocular structures (e.g., iris, pupil, and eyelid margins) with significant corneal reflections, but also corneal reflections and ground truth of the human eye surroundings scene separated from the ocular images. We believe that with the help of the CRED dataset, a large number of deep learning-based end-to-end works can be performed for iris and pupil position estimation and localization in the presence of strong corneal reflection interference. Similarly, the clarity and usability of corneal reflection images will be significantly improved.
Mengqi Du, Yue Zhang 0072, Jianhua Zhang 0002, Honghai Liu 0001
CSCWD3
2024 AdaFSNet: Time Series Classification Based on Convolutional Network with a Adaptive and Effective Kernel Size Configuration
abstract
Time series classification is one of the most critical and challenging problems in data mining, existing widely in various fields and holding significant research importance. Despite extensive research and notable achievements with successful real-world applications, addressing the challenge of capturing the appropriate receptive field (RF) size from one-dimensional or multi-dimensional time series of varying lengths remains a persistent issue, which greatly impacts performance and varies considerably across different datasets. In this paper, we propose an Adaptive and Effective Full-Scope Convolutional Neural Network (AdaFSNet) to enhance the accuracy of time series classification. This network includes two Dense Blocks. Particularly, it can dynamically choose a range of kernel sizes that effectively encompass the optimal RF size for various datasets by incorporating multiple prime numbers corresponding to the time series length. We also design a TargetDrop block, which can reduce redundancy while extracting a more effective RF. To assess the effectiveness of the AdaFSNet network, comprehensive experiments were conducted using the UCR and UEA datasets, which include one-dimensional and multi-dimensional time series data, respectively. Our model surpassed baseline models in terms of classification accuracy, underscoring the AdaFSNet network’s efficiency and effectiveness in handling time series classification tasks.
Haoxiao Wang, Jianhua Zhang 0002, Xu Cheng 0003
IJCNN3
2024 ColVO: Colonoscopic Visual Odometry Considering Geometric and Photometric Consistency
abstract
Locating lesions is the primary goal of colonoscopy examinations.3D perception techniques can enhance the accuracy of lesion localization by restoring 3D spatial information of the colon. However, existing methods focus on the local depth estimation of a single frame and neglect the precise global positioning of the colonoscope, thus failing to provide the accurate 3D location of lesions. The root causes of this shortfall is twofold: Firstly, existing methods treat colon depth and colonoscope pose estimation as independent tasks or design them as parallel sub-task branches. Secondly, the light source in the colon environment moves with the colonoscope, leading to brightness fluctuations among continuous frame images. To address these two issues, we propose ColVO, a novel deep learning-based Visual Odometry framework, which can continuously estimate colon depth and colonoscopic pose using two key components: a deep couple strategy for depth and pose estimation (DCDP) and a light consistent calibration mechanism (LCC). DCDP utilization of multimodal fusion and loss function constraints to couple depth and pose estimation modes ensure seamless alignment of geometric projections between consecutive frames. Meanwhile, LCC accounts for brightness variations by recalibrating the luminosity values of adjacent frames, enhancing ColVO's robustness. A comprehensive evaluation of ColVO on colon odometry benchmarks reveals its superiority over state-of-the-art methods in depth and pose estimation. We also demonstrate two valuable applications: immediate polyp localization and complete 3D reconstruction of the intestine. The code for ColVO is available at https://github.com/HNUicda/CoIVO.
Ruyu Liu, Zhengzhe Liu, Guodao Zhang, Jianhua Zhang 0002, Weiguo Sheng 0001, Xiufeng Liu 0001, Yaochu Jin
ACM Multimedia5
2024 Semi-Supervised Camouflaged Object Detection: Multi Information Fusion Combined with Adaptive Receptive Field Selection Network
Feng Xiao 0005, Ruyu Liu, Jianhua Zhang 0002, Shengyong Chen
PRCV (12)5
2024 Online Anomaly Detection for Streaming Data in the Presence of Missing Values
abstract
Online anomaly detection is a critical area in data analysis, particularly for handling dynamic data streams and addressing the challenge of concept drift. While current methods for online anomaly detection have achieved significant breakthroughs, creating a system that can continuously and effectively learn in scenarios with missing data remains a formidable challenge. In this paper, we introduce an autoencoder-based online deep anomaly detection model that addresses both missing data and concept drift. The model features a lightweight module specifically designed for efficient missing value processing. Additionally, it incorporates an adaptive model pool to manage the time-varying concept drift commonly observed in dynamic data streams. This flexible and dynamic management mechanism allows the model to adapt to changes in the data stream, maintaining robust anomaly detection performance across various conditions. Empirical validation of our model through ten comparative experiments on high-dimensional datasets affected by concept drift shows that it outperforms existing state-of-the-art methods. These results underscore the effectiveness and practicality of our approach.
Mengna Liu, Xu Cheng 0003, Lei Song 0011, Jianhua Zhang 0002
SMC5
2024 Object discriminability re-extraction for distractor-aware visual object tracking
Dongyan Guo, Xiangjie Kong 0001, Zhenhua Wang 0003, Jianhua Zhang 0002
Comput. Vis. Image Underst.6
2024 MLKAF-Net: Multiscale Large Kernel Attention Network for Hyperspectral and Multispectral Image Fusion
abstract
The fusion of a low spatial resolution hyperspectral image (LR-HSI) with a high spatial resolution multispectral image (HR-MSI) aims to synthesize a high-resolution hyperspectral image (HR-HSI), enabling a broader range of applications for hyperspectral images (HSIs). However, existing fusion methods struggle to capture both long-range dependencies and fine-grained spatial features, resulting in block artifacts and spatial distortions in the reconstructed HR-HSIs. Therefore, we introduce MLKAF-Net, a multiscale HSI-MSI fusion method, which effectively formulates cross-modality fused features in both spatial and spectral domains. MLKAF-Net mainly consists of three modules: the multiscale large kernel attention module (MLKAM), the spatial information aggregation module (SIAM), and the spectral attention module (SPAM). Specifically, the MLKAM incorporates a multiscale mechanism into the large kernel decomposition, adaptively capturing both long-range dependencies and local granular information. We develop the SIAM to establish the spatial quality of the reconstructed HR-HSIs by aggregating abundant spatial information. The SPAM introduces the channel attention to effectively mitigate spectral distortion through preserving beneficial spectral information. Extensive experiments demonstrate that our MLKAF-Net importantly enhances the fusion performance compared to state-of-the-art methods.
Haozheng Zhang, Yanhong Yang, Jianhua Zhang 0002, Shengyong Chen
IEEE Geosci. Remote. Sens. Lett.3
2024 Bilevel Fusion With Local and Global Cues for Point Cloud Upsampling
abstract
This study focuses on point cloud upsampling, crucial in 3-D data processing but hindered by current 3-D sensor limitations. Point clouds from RGB-D cameras and light detection and ranging (LiDAR) scanners are often sparse, noisy, and irregular, challenging traditional processing methods reliant on prior knowledge and hindering detail preservation. Despite deep learning's transformative impact, issues like hole overfitting and insufficient local-global feature fusion persist. To address these, we introduce the bilevel fusion point cloud upsampling (BiPU) network. It features a parallel extractor for simultaneous local and global feature extraction and a consistency-based feature alignment module employing cross-attention for enhanced multiscale feature transfer. BiPU also incorporates 4-D encoding for rotational invariance and depthwise separable convolutions to reduce complexity and parameters. Tested across multiple datasets, BiPU excels in maintaining hole contours and reducing costs, marking a notable advancement in point cloud processing.
Yunrui Zhu, Xu Cheng 0003, Jianhua Zhang 0002
IEEE Trans. Ind. Informatics4
2024 KSRB-Net: a continuous sign language recognition deep learning strategy based on motion perception mechanism
Feng Xiao 0005, Yunrui Zhu, Ruyu Liu, Jianhua Zhang 0002, Shengyong Chen
Vis. Comput.4
2023 DSP-Based Industrial Defect Detection for Intelligent Manufacturing
abstract
Internet of Things (IoT) based industrial defect detection has attracted more and more attention. As a key component of intelligent manufacturing, defect detection is very important. Although deep learning (DL) can reduce the cost of traditional manual inspection and improve accuracy and efficiency, it requires huge computing resources and cannot be simply deployed on IoT devices. Digital signal processor (DSP) is an important IoT device with the characteristics of small size, strong performance and low energy consumption, and has been widely used in intelligent manufacturing. In order to achieve accurate defect detection on DSP, we proposed a variety of optimization strategies, and then extended the model to run on multi-core using a parallel scheme, and further quantified the implementation of the model. We evaluated it on three datasets, i.e. NEUSDD, MTDD and RSDD. Experimental results show that our method achieves a faster speed than running the same CNN model on a mainstream desktop CPU, with slightly accuracy loss.
Rucen Wang, Ailing Xia, Jianhua Zhang 0002
CSCWD5
2023 Human Interaction Understanding With Consistency-Aware Learning
abstract
Compared with the progress made on human activity classification, much less success has been achieved on human interaction understanding (HIU). Apart from the latter task is much more challenging, the main causation is that recent approaches learn human interactive relations via shallow graphical representations, which are inadequate to model complicated human interactive-relations. This paper proposes a deep consistency-aware framework aiming at tackling the grouping and labelling inconsistencies in HIU. This framework consists of three components, including a backbone CNN to extract image features, a factor graph network to implicitly learn higher-order consistencies among labelling and grouping variables, and a consistency-aware reasoning module to explicitly enforcing consistencies. The last module is inspired by our key observation that the consistency-aware reasoning bias can be embedded into an energy function or a particular loss function, minimizing which delivers consistent predictions. An efficient mean-field inference algorithm is proposed, such that all modules of our network could be trained in an end-to-end fashion. Experimental results demonstrate that the two proposed consistency-learning modules complement each other, and both make considerable contributions in achieving leading performance on three benchmarks of HIU. The effectiveness of the proposed approach is further validated by experiments on detecting human-object interactions.
Jiajun Meng, Zhenhua Wang 0003, Kaining Ying, Jianhua Zhang 0002, Dongyan Guo, Zhen Zhang 0008, Qinfeng Shi, Shengyong Chen
IEEE Trans. Pattern Anal. Mach. Intell.4
2023 DSP-Based Traffic Target Detection for Intelligent Transportation
abstract
Internet of Things (IoT)-based intelligent transportation is attracting more and more attention. As a key component of intelligent transportation, traffic video monitoring is very important, in which vehicle and pedestrian detection on the road is a crucial task. Although vehicle and pedestrian detection through deep learning (DL) may achieve high accuracy, it tends to require high computing resources, which hinders its use on IoT devices. As an important class of IoT devices, digital signal processor (DSP) has the characteristics of low energy consumption, small size, and strong performance, which has been widely used in intelligent transportation. In order to use DL on DSP for accurate vehicle and pedestrian detection, we first propose a series of general tactics to optimize the object detection convolutional neural network (CNN) model, including convolution layer optimization, cache optimization, compiler optimization, intrinsics optimization and direct memory access (DMA) acceleration, and then a parallel scheme to extend the model to run on multicore, and further quantize the implementation of the model. We evaluate it on UA-DETRAC and KITTI datasets. Experimental results show that our method achieves a faster speed than running the same CNN model on a mainstream desktop CPU, with only 0.06% accuracy loss.
Jianhua Zhang 0002, Rucen Wang, Ruyu Liu, Dongyan Guo, Bo Li 0090, Shengyong Chen
IEEE Trans. Intell. Transp. Syst.1
2022 A Swift Gaze Estimate Method Based On The Corneal Image System
abstract
With the development of intelligent manufacturing, the demand of the incoming Human-Machine Interaction such as the augment reality rapidly increasing. However, the existing interaction modes in the augment reality, rely heavily on the hands or head movement. The inflexible modes is inefficient in the busy work flow. In this paper, we propose a gaze estimation work based on the Corneal Image System which can improve the efficiency of the interaction. Several prior works have proved, single Corneal Image contains the subject’s gaze information. However, the quality of Corneal Image is always impacted by the color and texture of the iris or the light from the surrounding, is hard to be applied directly. In order to improve the quality of the Corneal Image, people usually import additional devices into their work, such as infrared camera or eye tracker. These extra devices cause their gaze estimation works to become cumbersome and hard to be re-implemented commonly. Our gaze estimation work requires no additional device, can be seamlessly integrated into the AR domain with the help of the AprilTag mark. An AprilTag mark contained in an eye image, is distinct enough to be recognized, meanwhile, owns the hybrid pose relationship information between the eye, camera, and the focused AprilTag mark. The gaze can be inferred through the rigid body coordinate transformation naturally from this relationship. Many experiments have demonstrated that our approach is much easier to be re-implemented than the previous Corneal Image System based gaze computing works, at the same time, have the near performance to the state of the art.
Mengqi Du, Kaiqi Chen 0001, Jianhua Zhang 0002, Honghai Liu 0001
CSCWD3
2022 TXSLAM: A Monocular Semantic SLAM Tightly Coupled with Planar Text Features
abstract
We propose a new monocular semantic simultaneous localization and mapping (SLAM) system that tightly couples planar text features. The system treats text features as a plane with rich texture information and semantic information, and more accurate camera pose estimation can be obtained by tightly coupling the semantic plane. Unlike previous work, it pioneers the use of words contained in the text to represent the semantic information of the plane, which enables the use of simpler and more efficient data association algorithms to match geometric planes. We evaluate our method in public datasets, and the final experimental results prove that our proposed system improves the accuracy of camera pose estimation. Additionally, the system augments the sparse map with semantic plane information, enhancing the applicability of the system in robotics, unmanned driving, augmented reality (AR), and virtual reality (VR).
Qiyi Tong, Luzhen Ma, Kaiqi Chen 0001, Jianhua Zhang 0002
CSCWD6
2022 Robust and Accurate Multi-Agent SLAM with Efficient Communication for Smart Mobiles
abstract
In a long-term large-scenario application, the multi-agent collaborative SLAM is expected to improve the robustness and efficiency of executing tasks for mobile agents. In this paper, a multi-agent collaborative visual-inertial SLAM system is proposed based on a centralized client-server (CS) architecture, where the clients run on smart mobiles. In general, multi-agent collaborative SLAM relies on robust and precise experience sharing and efficient communication among agents. The experience sharing requires the place recognition with a high recall and accuracy, the precise estimation of transformation between looping frames, and the map fusion with globally consistency. To this end, we devise an enhanced geometric verification, a re-projection optimization based on the error-aware weighting strategy, and a strategy of flexible fusion to meet these requirements. In addition, the multi-agent collaborative SLAM needs to exchange abundant information, which requires the efficient communication. Therefore, we design a CS collaborative loop detection mechanism which is more robust to network transmission. We perform extensive experiments on the EuRoc dataset and in real environments. Experimental results show that the proposed system achieves better results than state-of-the-art methods. Furthermore, we demonstrate the stability of the proposed collaborative SLAM in real environments with a bandwidth of 7.55Mbps.
Kaiqi Chen 0001, Ruyu Liu, Yanhong Yang, Zhenhua Wang 0003, Jianhua Zhang 0002
ICRA6
2022 Human Interaction Recognition with Skeletal Attention and Shift Graph Convolution
abstract
Human interaction recognition has wide applications including intelligent surveillance, intelligent transportation and the analysis of sports videos. In recent years, benefiting from the development of action recognition based on deep learning, the performance of human interaction recognition has been boosted. This paper tackles two vital issues in recognizing human interactions, namely target missing and inadequate feature expression. To this end, we first design a data preprocessing method using skeleton estimation and multi-object tracking, which effectively reduces the chance of missing detection. Second, we propose a two-stream network composing of an appearance branch and a pose branch. The appearance branch extracts features enhanced via part affinity maps and part confidences maps, while the pose branch trains a customized Shift-GCN to extract skeletal features from people-pairs. Appearance and pose features are then fused to generate a more powerful representation of human interactions. Extensive experiments on two existing benchmarks, UT and BIT-Interaction, as well as a new dataset crafted by us, namely Campus-Interaction (CI), demonstrate the superior performance of the proposed approach over the state-of-the-arts.
Zhenhua Wang 0003, Jiajun Meng, Sheng Liu 0002, Jianhua Zhang 0002, Shengyong Chen
IJCNN5
2022 Robust Visual-Lidar Simultaneous Localization and Mapping System for UAV
abstract
Obtaining 3-D data by LIDAR from unmanned aerial vehicles (UAVs) is vital for the field of remote sensing; however, the highly dynamic movement of UAVs and narrow viewpoint of LIDAR pose a great challenge to the self-localization for UAVs based on solely LIDAR sensor. To this end, we propose a robust simultaneous localization and mapping (SLAM) system, which combines the image data obtained by vision sensor and point clouds obtained by LIDAR. In the front-end of the proposed system, the more stable line and plane features are extracted from point clouds through clustering. Then the relative pose between two consecutive frames is computed by the least squares iterative closest point algorithm. Afterward, a novel direct odometry algorithm is developed by combining the image frames and sparse point clouds, where the relative pose is used as a prior. In the back-end, the pose estimation is refined and the 3-D map with texture information is built at a lower frequency. Extensive experiments show that our method can achieve robust and highly precise localization and mapping for UAVs.
Kaiqi Chen 0001, Qinying Chen, Yanhong Yang, Jianhua Zhang 0002, Shengyong Chen
IEEE Geosci. Remote. Sens. Lett.5
2022 Hyperspectral Image Restoration via Subspace-Based Nonlocal Low-Rank Tensor Approximation
abstract
In this letter, we present a subspace-based nonlocal low-rank tensor approximation framework (SNLRTA) for hyperspectral image (HSI) restoration. The proposed method consists of a subspace learning method to achieve an accurate subspace characterization of HSI and a nonlocal low-rank tensor approximation to take spatial nonlocal self-similarity into consideration. Specifically, the HSI first exploits residual statistics on median filtered image to estimate a robust subspace. Laplacian scale mixture (LSM) modeling is then investigated to model tensor coefficients from overlapping cubes in low-rank subspace. Both the hidden scale parameters and the sparse coefficients therein are adaptively shrink, characterizing the sparsity of similar patches. Meanwhile, the$\ell _{1}$data fidelity facilitates the implicit detection of outliers after median filtering. Substantiated by extensive experimental results, the proposed method outperforms several state-of-the-art approaches on mixed noise removal, qualitatively and quantitatively.
Yanhong Yang, Yuan Feng 0002, Jianhua Zhang 0002, Shengyong Chen
IEEE Geosci. Remote. Sens. Lett.3
2022 Dual attention granularity network for vehicle re-identification
Jianhua Zhang 0002, Jingbo Chen, Jiewei Cao, Ruyu Liu, Linjie Bian, Shengyong Chen
Neural Comput. Appl.1
2022 Dual-View 3D Reconstruction via Learning Correspondence and Dependency of Point Cloud Regions
abstract
Multi-view 3D reconstruction generally adopts the feature fusion strategy to guide the generation of 3D shape for objects with different views. Empirically, the correspondence learning of object regions across different views enables better feature fusion. However, such idea has not been fully exploited in existing methods. Furthermore, current methods fail to explore the intrinsic dependency among regions within a 3D shape, leading to a rough reconstruction result. To address the above issues, we propose a Dual-View 3D Point Cloud reconstruction architecture named DVPC, which takes two views images as inputs, and progressively generates a refined 3D point cloud. First, a point cloud generation network is assigned to generate a coarse point cloud for each input view. Second, a dual-view point clouds synthesis network is presented in DVPC. It constructs a regional attention mechanism to learn a high-quality correspondence among regions across two coarse point clouds in different views, so that our DVPC can achieve feature fusion accurately. And then it develops a point cloud deformation module to produce a relatively-precise point cloud via establishing the communication between the coarse point cloud and the fused feature. Lastly, a point-region transformer network is devised to model the dependency among regions within the relatively-precise point cloud. With the dependency, the relatively-precise point cloud is refined into a desirable 3D point cloud with rich details. Qualitative and quantitative experiments on the ShapeNet and Pix3D datasets demonstrate that the proposed DVPC outperforms the state-of-the-art methods in terms of reconstruction quality.
Jianxin Wang 0001, Shourui Yang, Yunbo Wang, Jianhua Zhang 0002, Yuxin Peng 0001, Shengyong Chen
IEEE Trans. Image Process.4
2022 Accurate Object Association and Pose Updating for Semantic SLAM
abstract
Current pandemic has caused the medical system to operate under high load. To relieve it, robots with high autonomy can be used to effectively execute contactless operations in hospitals and reduce cross-infection between medical staff and patients. Although semantic Simultaneous Localization and Mapping (SLAM) technology can improve the autonomy of robots, semantic object association is still a problem that is worthy of being studied. The key to solving this problem is to correctly associate multiple object measurements of one object landmark by using semantic information, and to refine the pose of object landmark in real time. To this end, we propose a hierarchical object association strategy and a pose-refinement approach. The former one consists of two levels, i.e., a short-term object association and a global one. In the first level, we employ the multiple-object-tracking for short-term object association, through which the incorrect association among objects whose locations are close and appearances are similar can be avoided. Moreover, the short-term object association can provide more abundant object appearance and more robust estimation of object pose for the global object association in the second level. To refine the object pose in the map, we develop an approach to choose the optimal object pose from all object measurements associated with an object landmark. The proposed method is comprehensively evaluated on seven simulated hospital sequences, a real hospital environment and the KITTI dataset. Experimental results show that our method has an obviously improvement in terms of robustness and accuracy for the object association and the trajectory estimation in the semantic SLAM.
Kaiqi Chen 0001, Qinying Chen, Zhenhua Wang 0003, Jianhua Zhang 0002
IEEE Trans. Intell. Transp. Syst.5
2021 Consistency-Aware Graph Network for Human Interaction Understanding
abstract
Compared with the progress made on human activity classification, much less success has been achieved on human interaction understanding (HIU). Apart from the latter task is much more challenging, the main cause is that recent approaches learn human interactive relations via shallow graphical models, which is inadequate to model complicated human interactions. In this paper, we propose a consistency-aware graph network, which combines the representative ability of graph network and the consistency-aware reasoning to facilitate HIU. Our network consists of three components, a backbone CNN to extract image features, a factor graph network to learn third-order interactive relations among participants, and a consistency-aware reasoning module to enforce labeling and grouping consistencies. Our key observation is that the consistency-aware-reasoning bias for HIU can be embedded into an energy, minimizing which delivers consistent predictions. An efficient mean-field inference algorithm is proposed, such that all modules of our network could be trained jointly in an end-to-end manner. Experimental results show that our approach achieves leading performance on three benchmarks. Code is available at https://git.io/CAGNet.
Zhenhua Wang 0003, Jiajun Meng, Dongyan Guo, Jianhua Zhang 0002, Qinfeng Shi, Shengyong Chen
ICCV4
2021 Collaborative Visual Inertial SLAM for Multiple Smart Phones
abstract
The efficiency and accuracy of mapping are crucial in a large scene and long-term AR applications. Multi-agent cooperative SLAM is the precondition of multi-user AR interaction. The cooperation of multiple smart phones has the potential to improve efficiency and robustness of task completion and can complete tasks that a single agent cannot do. However, it depends on robust communication, efficient location detection, robust mapping, and efficient information sharing among agents. We propose a multi-intelligence collaborative monocular visual-inertial SLAM deployed on multiple ios mobile devices with a centralized architecture. Each agent can independently explore the environment, run a visual-inertial odometry module online, and then send all the measurement information to a central server with higher computing resources. The server manages all the information received, detects overlapping areas, merges and optimizes the map, and shares information with the agents when needed. We have verified the performance of the system in public datasets and real environments. The accuracy of mapping and fusion of the proposed system is comparable to VINS-Mono which requires higher computing resources.
Ruyu Liu, Kaiqi Chen 0001, Jianhua Zhang 0002, Dongyan Guo
ICRA4
2021 Detection and Segmentation of Unlearned Objects in Unknown Environment
abstract
Detecting and segmenting unlearned objects in unknown environment is a very important visual perception ability to enhance industrial intelligence. In this article, we present a novel conditional random field model integrating unimodal and cross-modal terms for detecting and segmenting object instances without knowing their categories and without sampling extra proposals. This model takes a paired image and point cloud as input, from which we first develop a set of novel category-independent features to distinguish objects. Then, a set of unary, pairwise, and higher order potentials are designed according to these category-independent features, and the cross-modal potential is introduced as a novel global constraints to keep the spatial consistency in both 2-D and 3-D modalities. In this novel model, the unlearned object detection and segmentation is treated as the process of pixel labeling. Thus, adjacent or occlusion object instances can also be separated efficiently from a labeled map. By comparison with the baseline methods, experimental results on a public RGB+D dataset show that the proposed model can obtain better performance with improved precision and recall rate. Moreover, we use the proposed method in a real industrial scene and achieve satisfactory performance.
Jianhua Zhang 0002, Jingbo Chen, Shengyong Chen, Zhenhua Wang 0003, Jianwei Zhang 0001
IEEE Trans. Ind. Informatics1
2021 Map Recovery and Fusion for Collaborative Augment Reality of Multiple Mobile Devices
abstract
The map recovery and fusion is a key issue in the application of large scale and long-term augmented reality (AR) scenarios. However, they are still not addressed well in an efficient and precise way, especially for complex industrial environments. In this article, we propose a map recovery and fusion strategy based on vision-inertial simultaneous localization and mapping. We first develop a heuristic strategy that can fast search and match map points among multiple maps, and can be used for efficient map fusion. For map recovery, we leverage the inertial sensors for short time motion estimation, and transform the previous lost map to the current map. Based on this strategy, a novel framework for collaborative AR is implemented and can parallelly run in multiple mobile devices in real time. Extensive experiments have been carried out on a public data set, and the results show that the proposed method can recovery and fuse multiple maps with high completeness and precision.
Jianhua Zhang 0002, Kaiqi Chen 0001, Zhiying Pan, Ruyu Liu, Thomas Yang 0001, Shengyong Chen
IEEE Trans. Ind. Informatics1
2021 Human Interaction Understanding With Joint Graph Decomposition and Node Labeling
abstract
The task of human interaction understanding involves both recognizing the action of each individual in the scene and decoding the interaction relationship among people, which is useful to a series of vision applications such as camera surveillance, video-based sports analysis and event retrieval. This paper divides the task into two problems including grouping people into clusters and assigning labels to each of them, and presents an approach to solving these problems in a joint manner. Our method does not assume the number of groups is known beforehand as this will substantially restrict its application. With the observation that the two challenges are highly correlated, the key idea is to model the pairwise interacting relations among people via a complete graph and its associated energy function such that the labeling and grouping problems are translated into the minimization of the energy function. We implement this joint framework by fusing both deep features and rich contextual cues, and learn the fusion parameters from data. An alternating search algorithm is developed in order to efficiently solve the associated inference problem. By combining the grouping and labeling results obtained with our method, we are able to achieve the semantic-level understanding of human interactions. Extensive experiments are performed to qualitatively and quantitatively evaluate the effectiveness of our approach, which outperforms state-of-the-art methods on several important benchmarks. An ablation study is also performed to verify the effectiveness of different modules within our approach.
Zhenhua Wang 0003, Jinchao Ge, Dongyan Guo, Jianhua Zhang 0002, Yanjing Lei, Shengyong Chen
IEEE Trans. Image Process.4
2020 Object-oriented Map Exploration and Construction Based on Auxiliary Task Aided DRL
abstract
Environment exploration by autonomous robots through deep reinforcement learning (DRL) based methods has attracted more and more attention. However, existing methods usually focus on robot navigation to single or multiple fixed goals, while ignoring the perception and construction of external environments. In this paper, we propose a novel environment exploration task based on DRL, which requires a robot fast and completely perceives all objects of interest, and reconstructs their poses in a global environment map, as much as the robot can do. To this end, we design an auxiliary task aided DRL model, which is integrated with the auxiliary object detection and 6-DoF pose estimation components. The outcome of auxiliary tasks can improve the learning speed and robustness of DRL, as well as the accuracy of object pose estimation. Comprehensive experimental results on the indoor simulation platform AI2-THOR have shown the effectiveness and robustness of our method.
Junzhe Xu 0001, Jianhua Zhang 0002, Shengyong Chen, Honghai Liu 0001
ICPR2
2020 CalibRCNN: Calibrating Camera and LiDAR by Recurrent Convolutional Neural Network and Geometric Constraints
abstract
In this paper, we present Calibration Recurrent Convolutional Neural Network (CalibRCNN) to infer a 6 degrees of freedom (DOF) rigid body transformation between 3D LiDAR and 2D camera. Different from the existing methods, our 3D-2D CalibRCNN not only uses the LSTM network to extract the temporal features between 3D point clouds and RGB images of consecutive frames, but also uses the geometric loss and photometric loss obtained by the interframe constraint to refine the calibration accuracy of the predicted transformation parameters. The CalibRCNN aims at inferring the correspondence between projected depth image and RGB image to learn the underlying geometry of 2D-3D calibration. Thus, the proposed calibration model achieves a good generalization ability to adapt to unknown initial calibration error ranges, and other 3D LiDAR and 2D camera pairs with different intrinsic parameters from the training dataset. Extensive experiments have demonstrated that our CalibRCNN can achieve state-of-the-art accuracy by comparison with other CNN based methods.
Jieying Shi, Ziheng Zhu, Jianhua Zhang 0002, Ruyu Liu, Zhenhua Wang 0003, Shengyong Chen, Honghai Liu 0001
IROS3
2020 Spatiotemporal Saliency Detection Based on Maximum Consistency Superpixels Merging for Video Analysis
abstract
Motion objects detection becomes more and more important in the applications of video surveillance, e.g., intrusion detection. The spatiotemporal saliency is an effective feature to describe object motion. However, there is a lot of redundancy in the spatial information preventing to obtain accurate saliency in an effective way. At the same time, temporal information cannot be accurately described because it is affected by uneven brightness, complex background, and fast-moving objects, especially at the edge of moving objects. In this article, we develop a novel method to tackle these problems and obtain more accurate spatiotemporal saliency. The key idea is the superpixel merging based on our maximum consistency model in feature space, through which the redundant spatial information is decreased and inhibit some temporal information errors. Experimental evaluations on the NNT dataset and surveillance videos show that the proposed method achieves better performance by comparing with some state-of-the-art methods, and can effectively detect intrusion entities.
Jianhua Zhang 0002, Jingbo Chen, Shengyong Chen
IEEE Trans. Ind. Informatics1
2019 New Convex Relaxations for MRF Inference With Unknown Graphs
abstract
Treating graph structures of Markov random fields as unknown and estimating them jointly with labels have been shown to be useful for modeling human activity recognition and other related tasks. We propose two novel relaxations for solving this problem. The first is a linear programming (LP) relaxation, which is provably tighter than the existing LP relaxation. The second is a non-convex quadratic programming (QP) relaxation, which admits an efficient concave-convex procedure (CCCP). The CCCP algorithm is initialized by solving a convex QP relaxation of the problem, which is obtained by modifying the diagonal of the matrix that specifies the non-convex QP relaxation. We show that our convex QP relaxation is optimal in the sense that it minimizes the L1 norm of the diagonal modification vector. While the convex QP relaxation is not as tight as the existing and the new LP relaxations, when used in conjunction with the CCCP algorithm for the non-convex QP relaxation, it provides accurate solutions. We demonstrate the efficacy of our new relaxations for both synthetic data and human activity recognition.
Zhenhua Wang 0003, Qinfeng Shi, M. Pawan Kumar, Jianhua Zhang 0002
ICCV5
2019 Joint Grouping and Labeling via Complete Graph Decomposition
Jinchao Ge, Zhenhua Wang 0003, Jiajun Meng, Jianhua Zhang 0002, Shengyong Chen
ICONIP (5)4
2019 Robust High Accuracy Visual-Inertial-Laser SLAM System
abstract
In recent years, many excellent works on visual-inertial SLAM and laser-based SLAM have been proposed. Although inertial measurement unit (IMU) significantly improve the motion estimate performance by reducing the impact of illumination variation or texture-less region on visual tracking, tracking failures occur when in such an environment for a long time. Similarly, when in structure-less environments, laser module will fail since lack of sufficient geometric features. Besides, motion estimation by moving lidar has the problem of distortion since range measurements are received continuously. To solve these problems, we propose a robust and high-accuracy visual-inertial-laser SLAM system. The system starts with a visual-inertial tightly-coupled method for motion estimation, followed by scan matching to further optimize the estimation and register point cloud on the map. Furthermore, we enable modules to be adjusted automatically and flexibly. That is, when one of these modules fails, the remaining modules will undertake the motion-tracking task. For further improving the accuracy, loop closure and proximity detection are implemented to eliminate drift accumulation. When loop or proximity is detected, we perform six degree-of-freedom (6-DOF) pose graph optimization to achieve the global consistency. The performance of our system is verified on public dataset, and the experimental results show that the proposed method achieves superior accuracy against other state-of-the-art algorithms.
Zengyuan Wang, Jianhua Zhang 0002, Shengyong Chen, Conger Yuan, Jingqian Zhang, Jianwei Zhang 0001
IROS2
2019 Towards SLAM-Based Outdoor Localization using Poor GPS and 2.5D Building Models
abstract
In this paper, we address the topic of outdoor localization and tracking using monocular camera setups with poor GPS priors. We leverage 2.5D building maps, which are freely available from open-source databases such as OpenStreetMap. The main contributions of our work are a fast initialization method and a non-linear optimization scheme. The initialization upgrades a visual SLAM reconstruction with an absolute scale. The non-linear optimization uses the 2.5D building model footprint, which further improves the tracking accuracy and the scale estimation. A pose optimization step relates the vision-based camera pose estimation from SLAM to the position information received through GPS, in order to fix the common problem of drift. We evaluate our approach on a set of challenging scenarios. The experimental results show that our approach achieves improved accuracy and robustness with an advantage in run-time over previous setups.
Ruyu Liu, Jianhua Zhang 0002, Shengyong Chen, Clemens Arth
ISMAR2
2019 Automatic segmentation of MR depicted carotid arterial boundary based on local priors and constrained global optimisation
abstract
Segmentation of lumen (LB) and outer wall boundaries (OB) of carotid artery in magnetic resonance (MR) images is essential for carotid atherosclerotic disease diagnosis. However, the limited image signal‐to‐noise ratio, flow artefact, and varied lumen and outer wall become significant obstacles for automatic segmentation. A fully automatic framework is proposed for LB and OB segmentation in MR images. First, the lumen is identified by the support vector machine using a special strategy and LB is segmented by the geodesic star‐shape‐constrained graph cut. Then a novel global optimisation is developed to segment OB based on the graph cut, which consists of shape priors and appearance priors. The shape priors are learned from labelled shapes on LB and OB, while the appearance priors are modelled by Gaussian mixture models. A novel shape constraint is also designed as the constraint term. To evaluate author's method, extensive experiments are carried out from 160 MR images belonging to 16 patients. Experimental results demonstrate that the proposed method can yield high accuracy with fully automatic segmentation. Moreover, the advantages of the proposed method have been shown in terms of high flexibility and accuracy without user interactions in comparison with other methods.
Jianhua Zhang 0002, Zhongzhao Teng, Qiu Guan, Junli He, Wafa Abutaleb, Andrew J. Patterson, Martin J. Graves, Jonathan Gillard 0001, Shengyong Chen
IET Image Process.1
2019 Three dimensional object segmentation based on spatial adaptive projection for solid waste
Jianhua Zhang 0002, Yeqiang Qiu, Jianshuang Guo, Jingbo Chen, Shengyong Chen
Neurocomputing1
2019 Hierarchical Topic Model Based Object Association for Semantic SLAM
abstract
Object-based simultaneous localization and mapping (SLAM) is a more natural and robust way for agents to interact with their surrounding environment. However, it introduces a problem of semantic objects association. Correct object association is the key factor to achieve a successful object SLAM system because object association and SLAM are inherently coupled and have not been well tackled yet. A novel formulation of the object association problem based on a hierarchical Dirichlet process (HDP) is proposed. Through the HDP, we can hierarchically associate the grouped object measurements. This can improve the object association accuracy and computation efficiency. Thanks to the novel formulation, the proposed method is also able to correct failure object associations according to its sampling inference algorithm. Furthermore, we introduce object poses to the processing of pose optimization. The object association and pose optimization are then solved in a tightly coupled way, by which both aspects can promote each other. The proposed method is evaluated on indoor and outdoor datasets and the experimental results show a very impressive improvement with respect to the traditional SLAM.
Jianhua Zhang 0002, Mengping Gui, Ruyu Liu, Junzhe Xu 0001, Shengyong Chen
IEEE Trans. Vis. Comput. Graph.1
2018 Feature Selection Mechanism in CNNs for Facial Expression Recognition
Shuwen Zhao, Haibin Cai, Honghai Liu 0001, Jianhua Zhang 0002, Shengyong Chen
BMVC4
2018 Instant SLAM Initialization for Outdoor Omnidirectional Augmented Reality
abstract
The initialization and absolute scale are two critical issues for an Augmented Reality (AR) system. Most existing methods have to resort to some external sensors or some special steps, to initialize an AR system and to obtain a correct scale. In this paper, we introduce an omnidirectional AR system, which can be instantly initialized, recover the absolute scale without other sensor data, and provide a full 360-degrees field-of-view (Fov) to give users the best immersive feelings. Our system shows how to firstly estimate the absolute orientation using the meaningful line and point cues from a single panorama, and how to estimate the camera global position by aligning a panorama after semantic segmentation with a widely available 2.5D map. Based on resulting absolute pose from a single frame, we subsequently render a depth map to initialize a SLAM system. We can then fuse the virtual elements with a real scene according to the continuous camera motion from the SLAM system. We evaluate the SLAM initialization approach on a challenging dataset. The experiments indicate that localization precision from our method is obviously superior to that from consumer GPS devices and we remain unbeatable in time performance compared to previous methods.
Ruyu Liu, Jianhua Zhang 0002, Kejie Yin, Jia-xin Wu, Ruihao Lin, Shengyong Chen
CASA2
2018 Automatic Re-topology and UV Remapping for 3D Scanned Objects based on Neural Network
abstract
Producing an editable model texture could be a challenging problem if the model is scanned from real world or generated by multi-view reconstruction algorithm. To solve this problem, we present a novel re-topology and UV remapping method based on neural network, which transforms arbitrary models with textured coordinates to a semi-regular meshes, and keeps models texture and removes the influence of lighting information. The main innovation of this paper is to use a neural network to find the appropriate location of the starting and ending points for models in the UV maps. Then each fragmented mesh is projected to the 2D planar domain. After calculating and optimizing the orientation field, a semi-regular mesh for each patch is then generated. Those patches can be projected back to three-dimension space and be spliced to a complete mesh. Experiments show that our method can achieve satisfactory performance.
Zhiying Pan, Make Di, Jianhua Zhang 0002, Suraj Ravi
CASA3
2018 Deep CRF-Graph Learning for Semantic Image Segmentation
Fuguang Ding, Zhenhua Wang 0003, Dongyan Guo, Shengyong Chen, Jianhua Zhang 0002, Zhanpeng Shao
PRICAI5
2018 Absolute Orientation and Localization Estimation from an Omnidirectional Image
Ruyu Liu, Jianhua Zhang 0002, Kejie Yin, Zhiying Pan, Ruihao Lin, Shengyong Chen
PRICAI2
2018 Understanding human activities in videos: A joint action and interaction learning approach
Zhenhua Wang 0003, Jiali Jin, Sheng Liu 0002, Jianhua Zhang 0002, Shengyong Chen, Zhen Zhang 0008, Dongyan Guo, Zhanpeng Shao
Neurocomputing5
2018 Object-level saliency: Fusing objectness estimation and saliency detection into a uniform framework
Jianhua Zhang 0002, Yanzhu Zhao, Shengyong Chen
J. Vis. Commun. Image Represent.1
2017 Joint label-interaction learning for human action recognition
abstract
Human interactions and their action categories preserve strong correlations, and the identification of the interaction configuration is of significant importance to improve the action recognition result. However, interactions are typically estimated using heuristics or treated as latent variables. The former usually produces incorrect interaction configuration while the latter introduces challenging training problem. Hence we propose a framework to jointly learn interactions and actions by designing a potential function using both features learned via deep neural networks and human interaction context. We propose an iterative approach to solve the associated inference problem efficiently and approximately. Experimental results on real datasets demonstrate that the proposed approach outperforms baselines by a large margin, and is competitive compared with the state-of-the-arts.
Jiali Jin, Zhenhua Wang 0003, Sheng Liu 0002, Jianhua Zhang 0002, Shengyong Chen, Qiu Guan
ICIP4
2017 Fast initialization for feature-based monocular slam
abstract
Initial map determines the effect of followed slam tracking. Most feature-based monocular slam initialize their map according to key points matching in close frames. Nevertheless, it will consume lots of computational resources and time. And it is easy to fail in some far scene or close scene. In this paper, we present a fast initialization method to reduce runtime and improve success rate of initialization for feature-based monocular slam. First, vanishing points detection based on line segment detector [1] is adopted. Second, we extract orb key points. And the coordinates of every key points are undistorted and normalized. Third, we generate the corresponding depth for each key point by normalizing its distance to the existing vanishing points or gaussian random number. We compare our method with state-of-the-art on public data sets and ours. The experiments show that our method outperforms on runtime and accuracy.
Shaobo Zhang 0005, Sheng Liu 0002, Jianhua Zhang 0002, Zhenhua Wang 0003, Xiaoyan Wang 0007
ICIP3
2017 A Spatio-Temporal CRF for Human Interaction Understanding
abstract
A better understanding of human interactions in videos can be achieved by simultaneously considering the coarse interactions between people, the action of each individual, and the activity of all people as a whole. We divide the recognition task into two stages. The first stage discriminates interactions and noninteractions, actions and activities based on local image information, while during the second stage, actions and activities are recognized in a global manner based on the local recognition results. A conditional random field (CRF) is designed to model human interactions in the spatio-temporal space. Different from most existing global models which cover either action or activity variables only, our model covers them both by considering the interactions between different types of variables. The graph structure of the CRF is predicted by a model learned from training data, which is different from traditional graph construction methods that typically rely on human heuristics. We learn the parameters of the CRF via structured support vector machine. We propose an efficient inference algorithm to tackle the estimation of labels in long videos containing many people. Our model admits both semantic-level understanding of human interactions in videos and competitive action and activity recognition performance.
Zhenhua Wang 0003, Sheng Liu 0002, Jianhua Zhang 0002, Shengyong Chen, Qiu Guan
IEEE Trans. Circuits Syst. Video Technol.3
2016 Objectness ranking by uniform Bayesian model with multimodal and global cues
abstract
Category‐independent object detection and localisation plays an important role in many computer vision tasks. In this study, an efficient method is proposed for generic objectness ranking by fusing two dimension (2D) or 3D information. A novel Bayesian model is designed to integrate multimodal cues and global cues to estimate object location, scale and number. In the pure trichannel colour space, the authors employ global spatial information as new global cues. From the colour+depth (red, green and blue+D) aspect, the authors compute multimodal saliency and oversegments to find two new multimodal cues. Local and regional depth cues are also explored and combined with them together so that a reliable objectness ranking scheme can be implemented. The proposed method is evaluated on web‐public common 2D and RGB+D datasets. In RGB+D cases, the experimental results show that the proposed method achieves an average 5% improvement over state‐of‐the‐art methods. Furthermore, for achieving the similar recall rates, the authors’ method only needs 30% amounts of sampled windows with respect of other available methods.
Jianhua Zhang 0002, Junhao Xiao 0001, Shengyong Chen, Jianwei Zhang 0001
IET Comput. Vis.1
2013 Discover Novel Visual Categories From Dynamic Hierarchies Using Multimodal Attributes
abstract
Learning novel visual categories from observations and experiences in unexplored environment is a vitally important cognitive ability for human beings. A dynamic category hierarchy that is an inherent structure in a human mind is a key component for this ability. This paper develops a framework to build dynamic category hierarchy based on object attributes and a topic model. Since humans trend to utilize multimodal information to learn novel categories, we also develop an algorithm to learn multimodal object attributes from multimodal data. The new multimodal attributes can describe objects efficiently and can generalize from learned categories to novel ones. By comparison with a state-of-the-art unimodal attribute, the multimodal attributes can achieve 4%-19% improvements on average. We also develop a constrained topic model, which can accurately construct category hierarchies for large-scale categories. Based on them, the novel framework can effectively detect novel categories and relate them with known categories for further category learning. Extensive experiments are conducted using a public multimodal dataset, i.e., color and point cloud data, to evaluate the multimodal attributes and the dynamic category hierarchy. The experimental results show the effectiveness of multimodal attributes to describe objects and the satisfactory performance of the dynamic category hierarchy to discover novel categories. By comparison with state-of-the-art methods, the dynamic category hierarchy achieves 7% improvements.
Jianhua Zhang 0002, Jianwei Zhang 0001, Shengyong Chen
IEEE Trans. Ind. Informatics1
2012 Constructing dynamic category hierarchies for novel visual category discovery
abstract
Category hierarchies are commonly used to compactly represent large numbers of categories and reduce the complexity of the classification problem. In this paper we introduce a novel and extended application of category hierarchies which is a powerful novel framework developed to construct dynamic category hierarchies and automatically discover novel visual categories. The dynamic is a characteristic of category hierarchies which can facilitate an important cognitive ability, the discovering of novel categories. We develop a constrained hierarchical latent Dirichlet allocation to build accurate category hierarchies. We employ object attributes as features to describe objects, which can transfer knowledge across categories and can efficiently describe novel categories. By combining them in the novel framework, novel visual object categories can be efficiently discovered and described. Extensive experiments based on PASCAL VOC 2008 and the LabelMe image database show the satisfactory performance of the proposed framework.
Jianhua Zhang 0002, Jianwei Zhang 0001, Shengyong Chen, Ying Hu 0001, Haojun Guan
IROS1
2012 A Hierarchical Model Incorporating Segmented Regions and Pixel Descriptors for Video Background Subtraction
abstract
Background subtraction is important for detecting moving objects in videos. Currently, there are many approaches to performing background subtraction. However, they usually neglect the fact that the background images consist of different objects whose conditions may change frequently. In this paper, a novel hierarchical background model is proposed based on segmented background images. It first segments the background images into several regions by the mean-shift algorithm. Then, a hierarchical model, which consists of the region models and pixel models, is created. The region model is a kind of approximate Gaussian mixture model extracted from the histogram of a specific region. The pixel model is based on the cooccurrence of image variations described by histograms of oriented gradients of pixels in each region. Benefiting from the background segmentation, the region models and pixel models corresponding to different regions can be set to different parameters. The pixel descriptors are calculated only from neighboring pixels belonging to the same object. The experimental results are carried out with a video database to demonstrate the effectiveness, which is applied to both static and dynamic scenes by comparing it with some well-known background subtraction methods.
Shengyong Chen, Jianhua Zhang 0002, Youfu Li 0001, Jianwei Zhang 0001
IEEE Trans. Ind. Informatics2
2011 Integrate multi-modal cues for category-independent object detection and localization
abstract
To detect and localize objects is an indispensable step for many computer vision tasks. Most of the state-of-the-art methods of object detection and localization are category-dependent. These methods can achieve a significant performance. However, they are useless for detecting and localizing objects belonging to an unknown category when applying them to an unknown environment. In this paper, a method is proposed for detecting and localizing generic objects without specifying their categories. The proposed method combines diverse cues, including multi-scale saliency, superpixels straddling, intensity, depth and global information, into a uniform Bayesian framework to obtain accurate detection and localization. By comparison to state-of-the-art methods, our experiments show the promising performance of the proposed method based on the PASCAL VOC 08 dataset and our indoor scene dataset.
Jianhua Zhang 0002, Junhao Xiao 0001, Jianwei Zhang 0001, Houxiang Zhang, Shengyong Chen
IROS1