EDBT 2026 Demo / reviewers in the wild / expert
Shengyong Chen
dblp:93/2479 · also Sheng-Yong Chen
· DBLP profile ↗
272ranked-venue papers
15as first author
154since 2021 · last 2026
0000-0002-6705-3831ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 127 · 8 first-author · 61 since 2021Graphics, computer vision, multimedia, augmented reality and games · 80 · 1 first-author · 49 since 2021Applied, interdisciplinary, general and emerging computing · 52 · 4 first-author · 36 since 2021Systems, architecture and hardware · 22 · 7 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 11 · 2 first-author · 5 since 2021Computer networks · 8 · 7 since 2021Databases, data management, data science and information retrieval · 6 · 4 since 2021Security and privacy · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A2P-Net: Asymmetric Domain-Adaptive Prototype Network for Cross-Domain Multimodal Sensor RetrievalabstractMultimodal sensor data from inertial measurement units (IMUs), including accelerometers, gyroscopes, and inclinometers, encode environmental conditions that are difficult to capture through images or text. In maritime settings, classifying sea states from such sensor streams is important for autonomous navigation but faces two interacting difficulties. First, training relies heavily on synthetic simulation data whose idealized physics diverge from real ocean measurements, creating a domain gap. Second, extreme sea states are rare in both domains, and the resulting class imbalance is compounded by the much larger volume of synthetic samples, which together skew gradient updates away from the scarce but operationally important real-world tail classes. We propose A2P-Net, an end-to-end framework that tackles both problems jointly. An adaptive heterogeneous encoder with decoupled channel–temporal attention maps variable-dimension sensor inputs into a shared latent space. Domain-adversarial training aligns synthetic and real feature distributions in that space, and prototype-based metric learning builds per-class retrieval anchors while an asymmetric weighting scheme up-weights real-domain samples to correct the optimization bias. On two custom sea state datasets that mix real and simulated ship motion recordings, A2P-Net reaches 98.7% and 98.9% F1 on real-only evaluation, outperforming the strongest baseline by 1.1–1.3 percentage points. It also ranks first on 15 of 30 UEA multivariate time-series benchmarks with an average accuracy of 74.2%. Mengna Liu, Xu Cheng 0003, Fan Shi 0001, Shengyong Chen |
ICMR | 6 |
| 2026 | Temporal-channel decoupled learning for sea state estimation from ship motion data under class imbalance
Mengna Liu, Xu Cheng 0003, Fan Shi 0001, Shengyong Chen |
Eng. Appl. Artif. Intell. | 6 |
| 2026 | Learning invariant representation for light field adversarial salient object detection
Mianzhao Wang, Fan Shi 0001, Xu Cheng 0003, Shengyong Chen |
Eng. Appl. Artif. Intell. | 4 |
| 2026 | Boundary-aware difficulty loss for long-tailed recognition
Yao Zhang 0021, Meng Zhao 0001, Shengyong Chen |
Expert Syst. Appl. | 4 |
| 2026 | BTKD++: Beyond Teachers by Critically Distilling Knowledge from Teacher's BiasabstractAbstract Existing knowledge distillation methods indiscriminately transfer knowledge from teacher networks, including output-level decisional biases, i.e., incorrect final predictions that can mislead student learning and limit student performance. We challenge this paradigm by proposing BTKD++, a framework that systematically filters and rectifies teacher’s output-level biased knowledge into corrective signals. Our approach partitions training data into Easy Tasks (correct teacher predictions) and Hard Tasks (incorrect predictions), then applies bias elimination and rectification modules orchestrated by dynamic learning curriculum. We provide an interpretive information-theoretic abstraction to explain the observed competence-threshold phenomenon, under which bias rectification becomes more effective when teacher errors contain sufficiently structured corrective information. BTKD++ demonstrates broad applicability across classification, detection, and segmentation tasks when task outputs are equipped with suitable probabilistic interfaces, and shows consistent effectiveness across CNNs, Transformers, and State-Space Models. Extensive experiments show consistent student-teacher transcendence, establishing new state-of-the-art results. This work redefines knowledge distillation from blind mimicry to critical learning, proving that students can surpass teachers through principled bias correction. The source code is available at https://github.com/smartyige/BTKD . Jianhua Zhang 0002, Yu He 0001, Xu Cheng 0003, Xiufeng Liu 0001, Shengyong Chen, Houxiang Zhang, Ruyu Liu |
Int. J. Comput. Vis. | 7 |
| 2026 | A Novel Dataset and Lightweight Distillation Baseline for Highlight Transparent Object Detection
Gang Li 0005, Qinghui Chen, Qunshu Zhang, Jin Wan, Maomao Xiong, Cong Bai, Dagang Li 0001, Wenyin Zhang, Jinglin Zhang 0004, Shengyong Chen |
Int. J. Comput. Vis. | 13 |
| 2026 | Few-shot video summarization via cross-video temporal invariance
Tinglong Tang, Fanyuan Wu, Shengyong Chen, Xu Cheng 0003 |
Neurocomputing | 3 |
| 2026 | WAMNet: Wavelet-enhanced asymmetric mamba network for semantic segmentation of multimodal remote sensing images
Fei Wang 0032, Yanhong Yang, Haozheng Zhang, Chengkun Li, Yushan Xue, Shengyong Chen |
Neurocomputing | 7 |
| 2026 | Segmentation guided edge enhanced teacher-student for industrial anomaly detection
Yanhong Yang, Haozheng Zhang, Fei Wang 0032, Shengyong Chen |
Neurocomputing | 5 |
| 2026 | ADASign: Adaptive deformable visual attention for continuous sign language recognition
Xuyan Zhang, Huayu Ma, Wanli Xue, Leming Guo, Tiantian Yuan, Shengyong Chen |
Neurocomputing | 8 |
| 2026 | Unification of Closed-Open Industrial Detection Scenarios: New Large-Scale Benchmarks, Challenges and BaselinesabstractLarge-scale Visual-Language Models (LVLMs) have achieved remarkable success in natural visual tasks, yet their application to industrial defect detection remains challenging due to two fundamental limitations: (i) the scarcity of large-scale industrial datasets that cover diverse defect categories across multiple domains, and (ii) the reliance on manual prompts (points, boxes, masks) that introduce subjective noise and lack text-visual interaction for fine-grained understanding. To address these challenges, we introduce a Large-Scale Multi-Modal Industrial Open-Closed benchmark (MMIOC-1 M) containing over one million samples across 14 super-categories, 29 industrial scenes, and 351 defect subcategories. To our knowledge, MMIOC-1 M is the first unified largest benchmark supporting both open-vocabulary and closed-set industrial detection, providing valuable pre-training data for LVLMs in industrial scenarios. Furthermore, we propose a Refined Text-Visual Prompt Network (RTVPNet) that incorporates three key innovations: (1) an expert-assisted domain projection mechanism that enables rapid adaptation of general vision models to industrial domains, (2) an energy-based sparse sampling strategy that automatically generates refined visual prompts without manual intervention, and (3) a bidirectional text-visual interaction module that enhances cross-modal semantic alignment and understanding. Extensive experiments demonstrate that RTVPNet achieves state-of-the-art performance on MMIOC-1 M, LVIS, and COCO benchmarks while maintaining computational efficiency. Jinglin Zhang 0001, Qinghui Chen, Gang Li 0005, Da Chen 0002, Shuainan Jing, Dagang Li 0001, Cong Liu 0012, Cong Bai, Shengyong Chen |
IEEE Trans. Pattern Anal. Mach. Intell. | 11 |
| 2026 | Instance-level geometric prompting for visual object tracking
Wanli Xue, Huayu Ma, Yangcan Wu, Shengyong Chen |
Pattern Recognit. | 6 |
| 2026 | Multi-branch perturbation learning with constraint simulation for semi-supervised semantic segmentationabstractCurrent semi-supervised semantic segmentation (SSS) methods improve generalization via weak-to-strong pseudo-supervision with image perturbations. However, many methods are limited by employing a single perturbation mode and a specific weak-to-strong learning strategy, restricting exploration of the perturbation space and hindering performance in fine-grained segmentation. While diverse perturbations are intuitively beneficial, simply combining them can lead to inefficient optimization and instability. In this paper, we propose a multi-branch strong perturbation constraint learning framework for SSS. Our framework introduces a novel multi-branch perturbation learning (MSPL) strategy, employing multiple parallel branches with diverse strong augmentations to expand the perturbation space and capture complex semantic variations. We further design a novel constraint simulation loss (CSSL), based on a hierarchical consistency learning structure (weak-to-strong and strong-to-strong), which enforces strong-to-strong consistency between different perturbation branches. CSSL mitigates instability and enhances robustness to perturbation-induced noise, enabling the network to better generalize and achieve more accurate segmentation, especially for fine object boundaries. Extensive evaluations on benchmark datasets (PASCAL VOC 2012, Cityscapes, COCO) demonstrate that our method achieves state-of-the-art performance. Ablation studies further validate the effectiveness of our proposed MSPL and CSSL components. Ruyu Liu, Feng Xiao 0005, Jianhua Zhang 0002, Xiufeng Liu 0001, Xu Cheng 0003, Shengyong Chen, Houxiang Zhang |
Pattern Recognit. | 6 |
| 2026 | Pioneering Video Semantic Segmentation With Light Field Imaging and Spatial-Angular-Temporal Fusion
Fan Shi 0001, Xu Cheng 0003, Shengyong Chen |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2026 | Time-Frequency Collaborative Learning for Imbalanced Ship Motion Data With Missing Labels in Sea State EstimationabstractSemi-supervised learning (SSL) has gained significant attention in the domain of sea state estimation (SSE) due to its capacity to alleviate the reliance of deep learning models on extensive labeled datasets. While existing semi-supervised SSE methodologies leveraging pseudo-labeling have achieved promising results, they often overlook the challenges posed by high class imbalance and the prevalence of missing data in ship motion datasets, which restricts their broader applicability. In this article, we propose a novel SSL approach BalanceSSE based on the class-imbalanced ship motion data for SSE. This approach consists of three main modules: 1) the dynamic imputation (DIT); 2) the imbalance temporal-frequency learning (ITFL); and 3) the ClusterProx classifier (CL). The DIT module dynamically imputes incomplete ship motion data by assigning different weights to various dimensions data. The ITFL module employs time-frequency collaborative learning to generate pseudo-labels and integrate an adaptive confidence strategy to select high confidence pseudo-labels. This process is further enhanced by the CL module to produce better estimates. Experimental tests on UCR datasets and ship motion datasets demonstrate that BalanceSSE outperforms state-of-the-art methods. Ablation studies highlight the critical role of each module in BalanceSSE. Mengna Liu, Xu Cheng 0003, Junhao Xiao 0001, Shengyong Chen |
IEEE Trans. Cybern. | 5 |
| 2026 | Adaptive Kernel Selection Module Combined With Feature Enhanced Perception Network for Camouflaged Object DetectionabstractCamouflaged object detection plays a crucial role in applications such as automatic sorting and defect inspection in industrial production, yet existing methods often struggle to flexibly capture features of diverse shapes, orientations, and scales due to their reliance on fixed receptive fields and rigid windowing schemes. To address these limitations, we propose a dual-branch joint network comprising a reference branch and a segmentation branch. The reference branch learns supplementary cues from salient objects that co-occur with camouflaged targets, guiding the segmentation branch toward more accurate delineation. Within the segmentation branch, we introduce three novel modules: 1) a deformable window interaction mechanism that replaces fixed-size transformer windows with learnable quadrilateral windows to adaptively extract features of arbitrary shape and orientation; 2) a feature enhancement perception module that fuses rich multiscale representations through parallel dilated convolutions at varying rates and channel-/spatial-attention mechanisms; and 3) a receptive field adjustment adaptive module that dynamically adjusts its receptive field size to balance sensitivity to fine details and global context. Comprehensive experiments on COD10 K, NC4K, CAMO, and R2C7K benchmarks demonstrate that our model outperforms the majority of current state-of-the-art approaches, while ablation studies and sensitivity analyses confirm the individual and combined effectiveness of our proposed components. Ruyu Liu, Feng Xiao 0005, Jianhua Zhang 0002, Shengyong Chen |
IEEE Trans. Ind. Informatics | 5 |
| 2026 | Learning Domain-Generalizable Discriminative Representations by Mixing Euclidean Dynamics and Hilbert Statistics for Wind Turbine Blade Icing DetectionabstractAs wind energy grows in importance, blade icing threatens turbine efficiency and safety. Existing approaches based solely on convolutional neural networks (CNNs) for Euclidean-space feature extraction struggle with complex dynamics and domain shifts. To address this, we propose the domain-generalizable network for icing turbines via mixed Euclidean and Hilbert representations (DGMEHIT), which runs a CNN deep module and a Hilbert statistical module (HSM) in parallel to extract complementary local sequential and global statistical features from Euclidean-space and Hilbert space. Raw features are mapped into a high-dimensional Hilbert space to enhance global statistical separability and discriminate superficially similar signals. A channel–temporal mixer module further models dynamic dependencies by fusing multichannel and time-domain information. In addition, a maximum mean discrepancy loss is incorporated into the HSM to improve feature consistency across domains, enabling the learning of domain-generalizable representations and enhancing the model’s adaptability to varying geographical and climatic conditions. Experiments on ten public time-series datasets and one real-world icing dataset show DGMEHIT surpasses state-of-the-art methods, improving F1 by 6.4% and MCC by 12.7%, with generalization tests and online evaluations confirming practical robustness. Mengna Liu, Yunke Li, Xu Cheng 0003, Xiufeng Liu 0001, Shengyong Chen |
IEEE Trans. Ind. Informatics | 6 |
| 2026 | WTCLIP: A Wavelet-Aware CLIP Framework for Boundary-Refined Weakly Supervised Semantic SegmentationabstractSome advanced methods have leveraged the zero-shot recognition capability of the contrastive language–image pretraining (CLIP) model and adapted it to weakly supervised semantic segmentation (WSSS), achieving promising performance. However, they primarily use CLIP as an auxiliary feature extractor, leaving the fundamental limitations of class activation mapping unresolved, particularly in preserving fine-grained object boundaries and achieving precise pixelwise localization under sparse supervision. To address these challenges, this article proposes a novel end-to-end WSSS framework WTCLIP, which aims to fully exploit the potential of CLIP for weakly supervised segmentation tasks. Different from traditional methods that use CLIP only as a static feature extractor, we innovatively introduce a learnable wavelet transform decoder to enhance the information extraction capability and significantly improve the model's perception of object boundaries. We dynamically adjust the weight distribution ratio of the CLIP feature layer, capture multiscale edge information, and make full use of the time–frequency localization characteristics of the wavelet transform to significantly improve the quality of pseudolabels and achieve more accurate semantic segmentation. Experimental results show that our method significantly improves the performance of the WSSS task on two public benchmark datasets, notably by4.0%over the state-of-the-art methods, especially in capturing weakly annotated object boundary details. Feng Xiao 0005, Jianhua Zhang 0002, Peihua Han, Shengyong Chen, Houxiang Zhang |
IEEE Trans. Ind. Informatics | 4 |
| 2026 | A Temporal-Spectral Mixer to Class Imbalanced Ship Motion Data-Based Sea State Estimation for Maritime Intelligent Transportation SystemsabstractAccurate sea state estimation (SSE) is critically important for enabling safe and efficient autonomous maritime transportation systems, yet traditional methods are costly, often exhibit latency, and are less suitable for integration within modern Intelligent Transportation Systems (ITS). While data-driven deep learning offers a promising alternative for real-time SSE within maritime ITS, the inherent class imbalance in naturally occurring sea states poses a significant challenge, hindering robust system development. Deep learning-based SSE methods often underutilize frequency-domain information, struggle with multi-scale wave characteristics, and are biased by class imbalance, limiting ITS effectiveness. To address these limitations and advance deep learning in maritime ITS, this paper proposes a novel class-imbalanced ship motion data-based Temporal-Spectral Mixer model for SSE. This model integrates a Spectral Frequency Adaptive (SFA) module to capture global spectral frequency information, a Temporal Multi-Scale Parallel Convolution (TMSPC) module for extracting local temporal multi-scale features, and an Imbalanced Contrastive Clustering Loss (ICC-Loss) function to mitigate class imbalance and enhance ITS applicability. The TMSPC module captures crucial temporal wave characteristics while the SFA module extracts global spectral context for system-aware SSE. Extensive evaluations demonstrate state-of-the-art performance on diverse datasets, including superior results on both public benchmarks and ship motion data compared to existing methods, including class-imbalance techniques. Ablation and sensitivity studies confirm the effectiveness of each module. The proposed Temporal-Spectral Mixer model offers a robust and promising SSE solution for class-imbalanced scenarios, advancing reliable maritime ITS and holding broader potential for time series classification within transportation systems. Xu Cheng 0003, Shengyong Chen |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2026 | Trustworthy Continuous Sign Language RecognitionabstractContinuous sign language recognition (CSLR) uses visual cues (e.g., hands, face, mouth, and body) to automatically recognize the sign language of deaf people, helping them to actively communicate with hearing people. The effects of these visual cues change dynamically with the demonstration of sign language. However, previous CSLR methods usually model visual information from the entire frame or simple fused visual cues, and thus do not well describe such dynamic change among visual cues. Therefore, we propose the Trustworthy Fusion Network ( TFN) of visual cues for CSLR, which comprises two fundamental modules: Intra-cue Cross-modality Feature Fusion module ( IntraCFF) and Inter-cue Trustworthy Fusion module (InterTF). IntraCFF uses the calibrated joint-belief method to dynamically fuse cross-modality features of RGB and keypoint information, to obtain a robust visual cue feature. InterTF innovatively employs the Dempster-Shafer Theory (DST) to evaluate the uncertainty of different cues in expressing sign movements. Then, the trustworthy fusion via DST is used to adaptively weigh and credibly fuse the visual cues based on uncertainty. In addition, to address the semantic gap when fusing different cues, we design consistency fusion constraints during the training stage. These constraints enhance the semantic consistency of different cues with global sign movements. Experiments on publicly CSLR datasets validate the effectiveness of our TFN. Yan Zhang 0154, Wanli Xue, Leming Guo, Yangcan Wu, Tiantian Yuan, Shengyong Chen |
IEEE Trans. Multim. | 8 |
| 2026 | Hierarchical Spatial-Angular Representation Learning for Point-Supervised Salient Object Detection in Light FieldsabstractLight Field Salient Object Detection (LFSOD) aims to identify visually distinctive regions by leveraging the complementary spatial–angular information inherent in 4D light field imagery. A major challenge lies in modeling angular dependencies and maintaining spatial coherence under sparse supervision. In this article, we propose a weakly supervised network that consists of three interdependent modules. First, the Light Field Division (LFD) module utilizes epipolar geometry to extract direction-aware boundary features, enhancing the encoding of angular disparities. Second, the Light Field Spatial Association (LFSA) module anchors cross-view feature alignment using central-viewpoint annotations, thereby enforcing spatial consistency and mitigating redundant representations. Third, the Light Field Saliency Local Clustering (LFLC) module introduces a joint boundary-appearance modeling strategy that integrates adaptive clustering with error-aware regularization to refine structural predictions. Experiments on three benchmark datasets show that our method consistently outperforms mainstream weakly supervised approaches. It also achieves superior performance compared to several fully supervised methods. Xinbo Geng, Fan Shi 0001, Xu Cheng 0003, Shengyong Chen |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2025 | Dust-Mamba: An Efficient Dust Storm Detection Network with Multiple Data SourcesabstractAccurate detection of dust storms is challenging due to complex meteorological interactions. With the development of deep learning, deep neural networks have been increasingly applied to dust storm detection, offering better learning and generalization capabilities compared to traditional physical modeling. However, existing methods face some limitations, leading to performance bottlenecks in dust storm detection. From the task perspective, existing research focuses on occurrence detection while neglecting intensity detection. From the data perspective, existing research fails to explore the utilization of multi-source data. From the model perspective, most models are built on convolutional neural networks, which have an inherent limitation in capturing long-range dependencies. To address the challenges mentioned, this study proposes Dust-Mamba. To the best of our knowledge, this study is the first attempt to accomplish both the occurrence and intensity detection of dust storms with advanced deep learning technology. In Dust-Mamba, multi-source data is introduced to provide a comprehensive perspective, Mamba and attention are applied to boost feature selection while maintaining long-range modeling capability. Additionally, this study proposes Structure Sharing Transfer Learning Strategies for intensity detection, which further enhances the performance of Dust-Mamba with minimal time cost. As shown by experiments, Dust-Mamba achieves Dice scores of 0.963 for occurrence detection and 0.560 for intensity detection, surpassing several baseline models. In conclusion, this study offers valuable baselines for dust storm detection, with significant reference value and promising application potential. Cong Bai, Zhonghao Lin, Jinglin Zhang 0001, Shengyong Chen |
AAAI | 4 |
| 2025 | Can Students Beyond the Teacher? Distilling Knowledge from Teacher's BiasabstractKnowledge distillation (KD) is a model compression technique that transfers knowledge from a large teacher model to a smaller student model to enhance its performance. Existing methods often assume that the student model is inherently inferior to the teacher model. However, we identify that the fundamental issue affecting student performance is the bias transferred by the teacher. Current KD frameworks transmit both right and wrong knowledge, introducing bias that misleads the student model. To address this issue, we propose a novel strategy to rectify bias and greatly improve the student model's performance. Our strategy involves three steps: First, we differentiate knowledge and design a bias elimination method to filter out biases, retaining only the right knowledge for the student model to learn. Next, we propose a bias rectification method to rectify the teacher model's wrong predictions, fundamentally addressing bias interference. The student model learns from both the right knowledge and the rectified biases, greatly improving its prediction accuracy. Additionally, we introduce a dynamic learning approach with a loss function that updates weights dynamically, allowing the student model to quickly learn right knowledge-based easy tasks initially and tackle hard tasks corresponding to biases later, greatly enhancing the student model's learning efficiency. To the best of our knowledge, this is the first strategy enabling the student model to surpass the teacher model. Experiments demonstrate that our strategy, as a plug-and-play module, is versatile across various mainstream KD frameworks. Jianhua Zhang 0002, Ruyu Liu, Xu Cheng 0003, Houxiang Zhang, Shengyong Chen |
AAAI | 6 |
| 2025 | QCTKD-PU: Quantum Convolutional Transformer with Knowledge Distillation for Efficient and Robust Point Cloud UpsamplingabstractPoint cloud upsampling is crucial for high-fidelity 3D reconstruction in real-time applications such as autonomous systems. Existing methods based on CNNs or Transformers face three limitations: (1) prohibitive computational complexity hindering real-time deployment, (2) insufficient modeling of multi-scale geometric dependencies in sparse data, (3) sensitivity to noise and outliers. To address these challenges, we propose QCTKD-PU, a framework integrating Quantum Convolutional Transformers (QCT) and Knowledge Distillation (KD) for Point cloud Upsampling. The QCT leverages quantum superposition and self-attention to encode high-dimensional features, enabling efficient multi-scale point interaction learning. Simultaneously, KD transfers knowledge from a teacher model to a lightweight student network, reducing computational costs while maintaining accuracy. Experiments on benchmark datasets demonstrate superior performance in geometric accuracy and noise robustness compared to state-of-the-art methods. This work pioneers the synergy of quantum computing and lightweight learning for resource-constrained 3D vision tasks, while the student model achieves real-time and compact deployment, offering a practical solution for collaborative edge systems. Yunrui Zhu, Ruyu Liu, Xu Cheng 0003, Jianhua Zhang 0002, Shengyong Chen |
CSCWD | 6 |
| 2025 | SCSegamba: Lightweight Structure-Aware Vision Mamba for Crack Segmentation in StructuresabstractPixel-level segmentation of structural cracks across various scenarios remains a considerable challenge. Current methods encounter challenges in effectively modeling crack morphology and texture, facing challenges in balancing segmentation quality with low computational resource usage. To overcome these limitations, we propose a lightweight Structure-Aware Vision Mamba Network (SCSegamba), capable of generating high-quality pixel-level segmentation maps by leveraging both the morphological information and texture cues of crack pixels with minimal computational cost. Specifically, we developed a StructureAware Visual State Space module (SAVSS), which incorporates a lightweight Gated Bottleneck Convolution (GBC) and a Structure-Aware Scanning Strategy (SASS). The key insight of GBC lies in its effectiveness in modeling the morphological information of cracks, while the SASS enhances the perception of crack topology and texture by strengthening the continuity of semantic information between crack pixels. Experiments on crack benchmark datasets demonstrate that our method outperforms other state-of-the-art (SOTA) methods, achieving the highest performance with only 2.8M parameters. On the multi-scenario dataset, our method reached 0.8390 in F1 score and 0.8479 in mIoU. The code is available at https://github.com/Karl1109/SCSegamba. Fan Shi 0001, Xu Cheng 0003, Shengyong Chen |
CVPR | 5 |
| 2025 | TGSR: Template-Guided Semantic Resampling against Adversarial Tracking AttacksabstractDeep object tracking has made significant strides, demonstrating impressive accuracy across diverse visual scenarios. However, recent studies revealed that visual object trackers remain vulnerable to adversarial attacks specifically designed to disrupt tracking tasks. While image resampling—reconstructing images using resampled coordinates and bilinear interpolation—has shown promise in enhancing tracker robustness by disrupting adversarial patterns, this naive approach overlooks crucial semantic information and scale variations between frames relative to the initial object template. These variations typically arise from changes in the target object or background. To address this limitation, we propose a template-guided semantic resampling (TGSR) method to counter adversarial tracking attacks. Our approach comprises two key components: template-aware joint semantic and appearance implicit representation (T-SAIR) and template-aware predictive resampling (T-PRES). T-SAIR estimates semantic embeddings and corresponding pixel colors at arbitrary coordinates based on incoming frames and the object template, while T-PRES predicts pixel-wise coordinate shifts in response to scale changes relative to the template. The integration of these modules enables our method to effectively reconstruct frames, neutralizing adversarial perturbations while preserving semantic information relative to the object template. Extensive experimental evaluation against three attacks across typical tracking methods demonstrates the effectiveness of our approach. Xuhong Ren, Jianlang Chen, Wanli Xue, Lei Ma 0003, Qing Guo 0005, Jianjun Zhao 0001, Shengyong Chen |
ICME | 7 |
| 2025 | Multi-Scale Convolutional Networks with Class-Normalized Logit Clipping for Robust Sea State Estimation from Noisy Ship Motion DataabstractAutonomous ships utilize automation systems to achieve unmanned navigation, driving innovation in maritime transportation. However, sea conditions, influenced by dynamic factors such as wave height, wind speed, and ocean currents, present a challenge in accurately assessing these conditions. Traditional classification models often assume accurate labels, but noisy labels are prevalent in real-world applications. Existing methods, such as noise sample filtering or loss function adjustment, have limited applicability and poor generalization when dealing with complex sea condition data. To address this issue, this study proposes an end-to-end neural network model. The model's feature extraction module uses deep representation learning to capture latent patterns in the data, and a loss function is designed to mitigate the impact of outliers. The integration of these components allows the model to perform accurate classification even in the presence of noisy labels. Extensive experiments on public and sea condition datasets validate the effectiveness of this approach, demonstrating that the model exhibits strong generalization capabilities and holds great promise for practical applications. Mengna Liu, Xu Cheng 0003, Xiufeng Liu 0001, Fan Shi 0001, Jianhua Zhang 0002, Shengyong Chen |
ICRA | 7 |
| 2025 | LFMamba: Focal Stack-aware State Space Modeling for Light Field Salient Object DetectionabstractSalient object detection (SOD) in light field data presents unique challenges due to dynamic semantic inconsistencies across focal slices and representation heterogeneity between focal slices and the all-focus image. Existing methods often treat focal slices uniformly or rely on simple fusion strategies, which fail to address focus-induced semantic drift and cross-modal feature misalignment. To tackle these issues, we propose LFMamba, a unified network that jointly models dynamic semantic consistency and adaptive cross-modal fusion. We design the Focal-aware State Space Module (FSSM), which generates focal-aware semantic prompts through low-rank decomposition and adaptively routes them according to focal plane indices, thereby enabling bidirectional semantic propagation across slices through non-causal state transitions. Furthermore, we introduce the Focal-guided Cross-modal Fusion Module (FCFM), which mitigates cross-modal heterogeneity by a two-stage hierarchical strategy, combining structure-aware low-level alignment and gated high-level semantic fusion. Extensive experiments on four public light field SOD benchmarks demonstrate that LFMamba achieves superior performance compared to state-of-the-art methods, with improved robustness and consistency under complex focal variation scenarios. Xinbo Geng, Fan Shi 0001, Xu Cheng 0003, Meng Zhao 0001, Shengyong Chen |
ACM Multimedia | 6 |
| 2025 | LIDAR: Lightweight Adaptive Cue-Aware Fusion Vision Mamba for Multimodal Segmentation of Structural CracksabstractAchieving pixel-level segmentation with low computational cost using multimodal data remains a key challenge in crack segmentation tasks. Existing methods lack the capability for adaptive perception and efficient interactive fusion of cross-modal features. To address these challenges, we propose a Lightweight Adaptive Cue-Aware Vision Mamba network (LIDAR), which efficiently perceives and integrates morphological and textural cues from different modalities under multimodal crack scenarios, generating clear pixel-level crack segmentation maps. Specifically, LIDAR is composed of a Lightweight Adaptive Cue-Aware Visual State Space module (LacaVSS) and a Lightweight Dual Domain Dynamic Collaborative Fusion module (LD3CF). LacaVSS adaptively models crack cues through the proposed mask-guided Efficient Dynamic Guided Scanning Strategy (EDG-SS), while LD3CF leverages an Adaptive Frequency Domain Perceptron (AFDP) and a dual-pooling fusion strategy to effectively capture spatial and frequency-domain cues across modalities. Moreover, we design a Lightweight Dynamically Modulated Multi-Kernel convolution (LDMK) to perceive complex morphological structures with minimal computational overhead, replacing most convolutional operations in LIDAR. Experiments on three datasets demonstrate that our method outperforms other state-of-the-art (SOTA) methods. On the light-field depth dataset, our method achieves 0.8204 in F1 and 0.8465 in mIoU with only 5.35M parameters. Code and datasets are available at https://github.com/Karl1109/LIDAR-Mamba. Fan Shi 0001, Xu Cheng 0003, Mengfei Shi, Xia Xie 0003, Shengyong Chen |
ACM Multimedia | 7 |
| 2025 | Watermark Removal via Boundary-Aware Segmentation and Semantic-Guided Diffusion
Zhenjie Jiang, Ruyu Liu, Jianhua Zhang 0002, Mohammed M. Elmogy, Shengyong Chen |
PRCV (2) | 8 |
| 2025 | DMDMN: distraction mining and dual mutual network for curvilinear structure segmentation
Qingbo Wu 0002, Shuofei Meng, Shengyong Chen |
Appl. Intell. | 4 |
| 2025 | An end-to-end model for time series classification in the presence of missing values
Mengna Liu, Pengshuai Yao, Xu Cheng 0003, Shengyong Chen |
Expert Syst. Appl. | 4 |
| 2025 | EHAN: An explicitly high-order attention network for accurate camouflaged object detection
Qingbo Wu 0002, Guanxing Wu, Shengyong Chen |
Neurocomputing | 3 |
| 2025 | One Stone, Three Birds: Prototype-Enhanced Federated Learning for Mitigating Data Scarcity, Imbalance, and Heterogeneity in Blade Icing Detection Across Distributed Wind FarmsabstractWind turbine blade icing poses a critical challenge to wind power generation in high-latitude regions, necessitating innovative solutions for reliable icing detection. To address this challenge while leveraging the abundance of unlabeled data and preserving data privacy, this study proposes a novel federated semi-supervised prototype learning framework, FedIce. By integrating prototype learning and federated learning, FedIce extracts representative class prototypes at the client level and performs global model updates through federated averaging, significantly enhancing robustness against data heterogeneity. Additionally, it incorporates an advanced separation margin strategy to effectively alleviate the adverse effects of class imbalance. Comprehensive experiments using real-world datasets from 20 wind turbines across two wind farms demonstrate that FedIce outperforms existing methods, achieving a remarkable 95.58% improvement in the$mF_{\beta }$metric and a 33.25% enhancement in the mBA metric compared to FedMatch. Lele Qi, Mengna Liu, Xu Cheng 0003, Xiufeng Liu 0001, Shengyong Chen |
IEEE Internet Things J. | 6 |
| 2025 | Alice-SLAM: Accurate and Lite-Communication Collaborative SLAM for Resource-Constrained Multi-AgentabstractMulti-agent collaborative simultaneous localization and mapping (Mac-SLAM) facilitates mutual localization among multi-agent and mapping in unknown environments. However, Mac-SLAM faces two main practical challenges in resource-constrained situations: heavy communication load and conflicts among multi-source maps. To address these issues, we propose Alice-SLAM: an accurate and lite-communication client-server collaborative SLAM system, reducing communication load while accuracy-guaranteed. Specifically, regarding high communication demand, we optimize communication load by compressing keyframe data and sharing only key map information instead of full map information. For inconsistency among multi-maps, we combine specific bundle adjustments (BA) and an adaptive strategy for active map optimization to enhance the consistency of the global map. A set of experiments demonstrates the superior accuracy and reduced communication load of the proposed Alice-SLAM on the EuRoC dataset and in multi-user augmented reality (AR) experiments conducted in our lab, highlighting its effectiveness in resource-constrained cases. We plan to open-source our code1to encourage further research and collaboration in this area. Kaiqi Chen 0001, Ruyu Liu, Xu Cheng 0003, Jianhua Zhang 0002, Shengyong Chen, Houxiang Zhang, Arash Ajoudani |
IEEE J. Sel. Areas Commun. | 6 |
| 2025 | Lightweight Multi-Stage Aggregation Transformer for robust medical image segmentation
Xiaoyan Wang 0007, Yating Zhu, Dongyan Guo, Pan Mu, Ming Xia 0005, Cong Bai, Zhongzhao Teng, Shengyong Chen |
Medical Image Anal. | 10 |
| 2025 | FocTrack: Focus attention for visual tracking
Sixian Chan 0001, Zhenchao Shi, Cong Bai, Shengyong Chen |
Pattern Recognit. | 5 |
| 2025 | SVD-KD: SVD-based hidden layer feature extraction for Knowledge distillation
Jianhua Zhang 0002, Mian Zhou, Ruyu Liu, Xu Cheng 0003, Sasa Nikolic 0002, Shengyong Chen |
Pattern Recognit. | 7 |
| 2025 | CORE: Multi-link graph attention network with inter-regional collaboration for continuous sign language recognition
Yan Zhang 0154, Wanli Xue, Tiantian Yuan, Shengyong Chen |
Pattern Recognit. | 5 |
| 2025 | Selective directed graph convolutional network for skeleton-based action recognition
Chengyuan Ke, Sheng Liu 0002, Yuan Feng 0002, Shengyong Chen |
Pattern Recognit. Lett. | 4 |
| 2025 | Wavelet-Discrete Cosine Transform Synergy for Ship Motion-Based Sea State Estimation in Autonomous ShipsabstractDeveloping a robust autonomous sea state estimation (SSE) model stands as a pivotal challenge in advancing autonomous ships. Presently, deep learning (DL) methodologies have showcased remarkable efficacy in SSE tasks. Nonetheless, the dynamic nature of ship motion introduces temporal variations alongside frequency domain characteristics like periodic swinging, posing challenges for existing DL approaches. Most prevailing DL techniques, predominantly leveraging Convolutional Neural Networks or Long Short-Term Memory Networks, often fail to effectively harness frequency domain information post feature extraction. To tackle these limitations head-on, this paper introduces a pioneering SSE model. Specifically, in order to solve the frequency-domain feature extraction problem, we design a wavelet transform-based frequency domain encoder to extract relevant frequency-domain features from ship motion data by discriminating the contribution of different frequencies in the signal. Subsequently, in order to better integrate the extracted ship motion features, we designed a Feature Perception module based on discrete cosine transform. This module adeptly merges the extracted feature insights while prioritizing crucial frequency domain features. Following rigorous experimentation, our methodology exhibits superior performance compared to existing baseline techniques in SSE, a capability of profound significance for autonomous ships. Moreover, across diverse public multivariate time series classification datasets, our model outperforms current state-of-the-art approaches, underscoring its scalability across distinct domains. Feng Xiao 0005, Xu Cheng 0003, Sasa Nikolic 0002, Jianhua Zhang 0002, Shengyong Chen |
IEEE Trans Autom. Sci. Eng. | 6 |
| 2025 | Prior Knowledge-Driven Hybrid Prompter Learning for RGB-Event TrackingabstractEvent data can asynchronously capture variations in light intensity, thereby implicitly providing valuable complementary cues for RGB-Event tracking. Existing methods typically employ a direct interaction mechanism to fuse RGB and event data. However, due to differences in imaging mechanisms, the representational disparity between these two data types is not fixed, which can lead to tracking failures in certain challenging scenarios. To address this issue, we propose a novel prior knowledge-driven hybrid prompter learning framework for RGB-Event tracking. Specifically, we develop a frame-event hybrid prompter that leverages prior tracking knowledge from the foundation model as intermediate modal support to mitigate the heterogeneity between RGB and event data. By leveraging its rich prior tracking knowledge, the intermediate modal reduces the gap between the dense RGB and sparse event data interactions, effectively guiding complementary learning between modalities. Meanwhile, to mitigate the internal learning disparities between the lightweight hybrid prompter and the deep transformer model, we introduce a pseudo-prompt learning strategy that lies between full fine-tuning and partial fine-tuning. This strategy adopts a divide-and-conquer approach to assign different learning rates to modules with distinct functions, effectively reducing the dominant influence of RGB information in complex scenarios. Extensive experiments conducted on two public RGB-Event tracking datasets show that the proposed HPL outperforms state-of-the-art tracking methods, achieving exceptional performance. Mianzhao Wang, Fan Shi 0001, Xu Cheng 0003, Shengyong Chen |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | Fine-Grained Modality Relation-Aware Network for Video Moment RetrievalabstractVideo moment retrieval (VMR) involves localizing video segments semantically aligned with given queries within videos. Despite the development of numerous methods for VMR in recent years, there remains a need to better incorporate fine-grained modality relation-aware information both in intra-modality and cross-modality. To address these challenges, we propose a Fine-grained Modality Relation-Aware Network (FMRN) tailored for the video moment retrieval task. FMRN effectively explores fine-grained modality relation-aware information within text queries, videos, and proposals. Our approach begins with a semantic graph encoder to capture deep semantic relations in intra-modality. Besides, we introduce a novel fine-grained cross-modality interaction module comprising a cross-similarity weighting module, an intra-modality weighting module, and an adaptive fusion module. These components comprehensively exploit fine-grained relation information within intra-modality and cross-modality contexts. Specifically, the cross-similarity weighting module leverages similarities between text queries and video snippets, as well as between videos and query words. The intra-modality weighting module determines the importance of words and snippets, while the adaptive fusion module combines cross-similarity weighting and intra-modality weighting. Additionally, we design a proposal relation module to enhance retrieval by capturing fine-grained proposals-relation information in videos. Extensive experiments demonstrate that the proposed method can outperform all state-of-the-art methods on the TACoS dataset and obtain comparable results on the Charades-STA and ActivityNet-Captions datasets. Compared with MCMN (TCSVT2024) and DPHANet (TMM2024), FMRN can achieve average improvements of 3.61 % and 5.44 % on the TACoS dataset, respectively. Yibo Zhao 0001, Zan Gao 0001, Chunjie Ma, Weili Guan, Riwei Wang, Shengyong Chen |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2025 | CESFusion: Cross-Frequency Enhanced Spatial - Spectral Fusion Network for Hyperspectral and Multispectral Image FusionabstractThe fusion of hyperspectral and multispectral images involves integrating high spectral resolution hyperspectral image (HSI) and high spatial resolution multispectral image (MSI) to generate a HSI with high spatial and spectral resolution (HR-HSI). Existing HSI-MSI fusion methods primarily focus on information fusion within the spatial domain; however, few solutions have explored the employment of frequency analysis to enhance spatial resolution, limiting their capability for global perception. In this paper, we propose an efficient and novel paradigm for HSI-MSI fusion through the cross-frequency enhanced spatial-spectral fusion network, named CESFusion, exploring the complementary fusion of information between the spatial and frequency domains. Specifically, we first present the cross-frequency domain fusion module (CFFM) to perform global analysis through the Fourier transform and effectively integrate and enhance the frequency domain information from both HSI and MSI. Subsequently, we propose the spectral modeling module (SpeMM) based on state space model (SMM) to capture long-range spectral dependencies with linear complexity, and integrate it with the spatial residual block-based module (SRM) for joint spatial-spectral feature extraction. Finally, to enable sufficient interaction between the spatial and frequency domains, we adopt the cross-domain interaction module (CDIM), capturing and integrating complementary information from both domains. Moreover, a frequency-based loss function is purposely designed to further improve the restoration of global information. Extensive experiments conducted on both synthetic and real datasets demonstrate the superiority of our CESFusion, as evidenced by both quantitative and qualitative evaluation results. Haozheng Zhang, Yanhong Yang, Yanjie Lu, Guodao Zhang, Shengyong Chen |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2025 | MFCLIP: Multi-Modal Fine-Grained CLIP for Generalizable Diffusion Face Forgery DetectionabstractThe rapid development of photo-realistic face generation methods has raised significant concerns in society and academia, highlighting the urgent need for robust and generalizable face forgery detection (FFD) techniques. Although existing approaches mainly capture face forgery patterns using image modality, other modalities like fine-grained noises and texts are not fully explored, which limits the generalization capability of the model. In addition, most FFD methods tend to identify facial images generated by GAN, but struggle to detect unseen diffusion-synthesized ones. To address the limitations, we aim to leverage the cutting-edge foundation model, contrastive language-image pre-training (CLIP), to achieve generalizable diffusion face forgery detection (DFFD). In this paper, we propose a novel multi-modal fine-grained CLIP (MFCLIP) model, which mines comprehensive and fine-grained forgery traces across image-noise modalities via language-guided face forgery representation learning, to facilitate the advancement of DFFD. Specifically, we devise a fine-grained language encoder (FLE) that extracts fine global language features from hierarchical text prompts. We design a multi-modal vision encoder (MVE) to capture global image forgery embeddings as well as fine-grained noise forgery patterns extracted from the richest patch, and integrate them to mine general visual forgery traces. Moreover, we build an innovative plug-and-play sample pair attention (SPA) method to emphasize relevant negative pairs and suppress irrelevant ones, allowing cross-modality sample pairs to conduct more flexible alignment. Extensive experiments and visualizations show that our model outperforms the state of the arts on different settings like cross-generator, cross-forgery, and cross-dataset evaluations. Our code will be available at https://github.com/Jenine-321/MFCLIP. Tianyi Wang 0006, Zitong Yu, Zan Gao 0001, LinLin Shen, Shengyong Chen |
IEEE Trans. Inf. Forensics Secur. | 6 |
| 2025 | Distortion-Aware Outdoor Panoramic Depth Estimation via Local-Global FusionabstractOutdoor panoramic depth estimation faces significant challenges due to the wide field of view (FoV), complex scene structures, and severe distortion encountered in such environments. Traditional methods, which often use distortion convolution, fall short in capturing global distortion information and extracting rich contextual details from panoramic images. To overcome these limitations, this article introduces a novel dual-branch framework that synergistically merges the advantages of equirectangular projection (ERP) and tangent projection (TP). First, we design a unique dual-branch framework specifically tailored for panoramic depth estimation. In this framework, the convolutional neural networks branch processes ERP images to extract rich local information, enhancing the detail accuracy of depth estimation, while the vision transformers branch processes TPs to capture comprehensive global information, improving the smoothness of depth estimation. Then, we further enhance our method with a distortion-aware weight map module that adapts the influence of different image regions according to their distortion level, thus prioritizing features from areas with less distortion. In addition, we implement a dual attention fusion module to seamlessly integrate features from both branches at corresponding layers. Comprehensive experiments across various outdoor datasets reveal that our method significantly outperforms state-of-the-art techniques in terms of depth estimation accuracy, adeptly balancing the capture of both overarching scene depth and intricate details, potentially revolutionizing applications in industrial informatics, such as autonomous navigation and environmental mapping. Ruyu Liu, Yihao Ying, Xiufeng Liu 0001, Weiguo Sheng 0001, Jianhua Zhang 0002, Shengyong Chen |
IEEE Trans. Ind. Informatics | 8 |
| 2025 | Cascaded State Space and Contrastive Learning for Cross-Domain Few-Shot SegmentationabstractCurrent cross-domain few-shot semantic segmentation (CD-FSS) faces multiple challenges, including inconsistent feature mapping among domains and insufficient utilization of low-level information and background information from the source domain. To address these issues, this article proposes a novel cascade feature enhancement and contrastive learning framework to improve the generalization capability of CD-FSS. Within this framework, we first introduce a cascade feature enhancement module to construct distinctive feature representations, enhancing the model’s transferability across domains. By effectively integrating multilevel feature information from support images, this module strengthens the representation capability of query images. Second, we employ contrastive learning to form positive and negative sample pairs for the foreground and background, capturing rich correlations between them. Finally, the iterative prototype enhancement module we propose gradually refines the correspondence between the support image and the query image through iteration, making full use of the embedded supervisory information in the limited support samples. Experimental results demonstrate that the proposed method outperforms existing approaches on multiple benchmark datasets, achieving up to a 9.7% improvement over state-of-the-art methods. Feng Xiao 0005, Jianhua Zhang 0002, Peihua Han, Shengyong Chen, Houxiang Zhang |
IEEE Trans. Ind. Informatics | 4 |
| 2025 | Cross-Scale Denoising Reverse Distillation for Anomaly DetectionabstractEffective discrepancy representation of anomalies plays a crucial role in visual anomaly detection. Recent advances build upon reverse distillation paradigm that boost the teacher–student model’s discrimination capability on anomalies; however, they are still susceptible to the size variation of unpredictable anomalies. To generalize the anomaly size variation, we propose a new algorithm cross-scale denoising reverse distillation (CDRD), which integrates cross-scale denoising with reverse distillation to exchange multiscale perception and enhance the fine-grained representation of features. Specifically, we introduce a cross-scale anomalous signal suppression procedure in the teacher network to facilitate the interaction of information across different scales, thereby enabling the student network to learn more robust normal data representations. In the knowledge transfer process, a fusion compression module acts as an intermediate transmitter of information, aiming to obtain a compact embedding while abandoning anomaly perturbations. Moreover, we construct a detail supplement module in the student network to prevent the loss of key information in the deconvolution process of the decoder. Experiments on well-known datasets demonstrate that our CDRD brings significant improvements over the next best competitor. Yanhong Yang, Feng Xiao 0005, Jianhua Zhang 0002, Guodao Zhang, Shengyong Chen |
IEEE Trans. Ind. Informatics | 6 |
| 2025 | Multiple Information Prompt Learning for Cloth-Changing Person Re-IdentificationabstractCloth-changing person re-identification is a subject closer to the real world, which focuses on solving the problem of person re-identification after pedestrians change clothes. The primary challenge in this field is to overcome the complex interplay between intra-class and inter-class variations and to identify features that remain unaffected by changes in appearance. Sufficient data collection for model training would significantly aid in addressing this problem. However, it is challenging to gather diverse datasets in practice. Current methods focus on implicitly learning identity information from the original image or introducing additional auxiliary models, which are largely limited by the quality of the image and the performance of the additional model. To address these issues, inspired by prompt learning, we propose a novel multiple information prompt learning (MIPL) scheme for cloth-changing person ReID, which learns identity robust features through the common prompt guidance of multiple messages. Specifically, the clothing information stripping (CIS) module is designed to decouple the clothing information from the original RGB image features to counteract the influence of clothing appearance. The bio-guided attention (BGA) module is proposed to increase the learning intensity of the model for key information. A dual-length hybrid patch (DHP) module is employed to make the features have diverse coverage to minimize the impact of feature bias. Extensive experiments demonstrate that the proposed method outperforms all state-of-the-art methods on the LTCC, CelebreID, Celeb-reID-light, and CSCC datasets, achieving rank-1 scores of 74.8%, 73.3%, 66.0%, and 88.1%, respectively. When compared to AIM (CVPR23), ACID (TIP23), and SCNet (MM23), MIPL achieves rank-1 improvements of 11.3%, 13.8%, and 7.9%, respectively, on the PRCC dataset. Shengxun Wei, Zan Gao 0001, Chunjie Ma, Yibo Zhao 0001, Weili Guan, Shengyong Chen |
IEEE Trans. Image Process. | 6 |
| 2025 | Semantic Visual Simultaneous Localization and Mapping: A SurveyabstractVisual Simultaneous Localization and Mapping (vSLAM) is a cornerstone technology in computer vision and robotics, underpinning applications such as autonomous vehicles and robot navigation. While traditional vSLAM systems have shown significant progress in indoor or outdoor environments, their performance often degrades in complex scenes, limiting their adaptability and robustness. Semantic vSLAM, which integrates high-level semantic information into vSLAM systems, has emerged as a promising solution to address these limitations by enabling a richer understanding of the environment. In this paper, we provide a comprehensive review of semantic vSLAM, offering a critical analysis of its evolution, methods, and challenges. We begin by revisiting the development of traditional vSLAM, emphasizing its limitations and the motivation for incorporating semantic information. Subsequently, we delve into the core modules of semantic vSLAM, including semantic extraction, object association, semantic loop closing, back-end optimization, and semantic mapping. Then, we present a performance comparison of semantic vSLAM systems under two different datasets, indoor and outdoor, respectively. Furthermore, we also provide a comparative analysis of widely used SLAM datasets to provide guidance for performance testing and validation. To further enrich the discussion, we identify unresolved challenges in semantic vSLAM, such as long-term semantic perception and association, open and unstructured environments. We propose future research directions, including balancing computational resources and quantifying system risk, large model-based navigation and mapping, and embodied AI SLAM. By providing key insights and forward-looking perspectives, this work aims to stimulate future research and improve the capabilities of semantic vSLAM in real-world applications. Kaiqi Chen 0001, Junhao Xiao 0001, Qiyi Tong, Heng Zhang 0023, Ruyu Liu, Jianhua Zhang 0002, Arash Ajoudani, Shengyong Chen |
IEEE Trans. Intell. Transp. Syst. | 9 |
| 2025 | Covariance Propagation-Based Accurate Loop Detection for High Confusion Environment
Kaiqi Chen 0001, Ruyu Liu, Shengyong Chen, Arash Ajoudani, Jianhua Zhang 0002 |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2025 | Leveragable Adaptive Multi-Scale Features and Learnable Prototypes for Imbalanced Sea State Estimation Based on Ship Motion DataabstractThe adoption of autonomous vessels has been accelerated by the flourishing maritime economy and more stringent shipping regulations. Accurate sea state estimation (SSE) is of paramount importance for their safe operation. Traditional SSE methods face several limitations, including subjectivity in manual observation, high radar costs, and the insufficient timeliness of satellite data. In contrast, SSE methods based on ship motion data, particularly deep learning models, offer advantages in capturing complex nonlinear relationships. However, challenges persist, such as imbalanced sea state data, suboptimal performance under extreme conditions, and the complexity of ship motion data, which hampers effective feature extraction. To address these challenges, this paper proposes a novel deep learning model that integrates a dynamic attention mechanism with adaptive selective kernel modules and multi-scale feature fusion to capture spatiotemporal dependencies. Additionally, an enhanced prototype classifier, utilizing cosine similarity and dynamic prototype updating, is introduced to mitigate data imbalance and improve model robustness in extreme conditions. To evaluate the performance of the proposed model, we conducted experiments on 30 benchmark datasets for multivariate time series classification, as well as two sea state datasets. The comparison results demonstrate that our model outperforms most baseline methods in general multivariate time series classification tasks and also surpasses state-of-the-art methods for SSE and class imbalance learning. Furthermore, the experimental results show that the proposed model is robust to data noise and missing values. Mengna Liu, Xiufeng Liu 0001, Xu Cheng 0003, Shengyong Chen |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2025 | Zero-shot object visual navigation using relation of historical objects with target transfer
Jiangpeng Zheng, Fan Shi 0001, Meng Zhao 0001, Shengyong Chen |
J. Supercomput. | 5 |
| 2025 | FRAME: Feature Rectification for Class Imbalance LearningabstractClass imbalance learning is a challenging task in machine learning applications. To balance training data, traditional class imbalance learning approaches, such as class resampling or reweighting, are commonly applied in the literature. However, these methods can have significant limitations, particularly in the presence of noisy data, missing values, or when applied to advanced learning paradigms like semi-supervised or federated learning. To address these limitations, this paper proposes a novel and theoretically-ensured latentFeatureRectification method for clAss iMbalance lEarning (FRAME). The proposed FRAME can automatically learn multiple centroids for each class in the latent space and then perform class balancing. Unlike data-level methods, FRAME balances feature in the latent space rather than the original space. Compared to algorithm-level methods, FRAME can distinguish different classes based on distance without the need to adjust the learning algorithms. Through latent feature rectification, FRAME can effectively mitigate contaminated noises/missing values without worrying about structural variations in the data. In order to accommodate a wider range of applications, this paper extends FRAME to the following three main learning paradigms: fully-supervised learning, semi-supervised learning, and federated learning. Extensive experiments on 10 binary-class datasets demonstrate that our FRAME can achieve competitive performance than the state-of-the-art methods and its robustness to noises/missing values. Xu Cheng 0003, Fan Shi 0001, Yao Zhang 0021, Huan Li 0003, Xiufeng Liu 0001, Shengyong Chen |
IEEE Trans. Knowl. Data Eng. | 6 |
| 2025 | Domain-Division Based Progressive Learning for Source-Free Domain AdaptationabstractWith growing privacy and portability concerns, source-free domain adaptation requires only a source pre-trained model and an unlabeled target domain, allowing for effective adaptation to the target data. Most existing self-training methods focus on selecting and exploiting samples with reliable predictions, often neglecting others. Inspired by the finding that deep models learn clean samples faster than noisy ones, we propose a domain-division based progressive learning method named DPL. Specifically, our approach consists of two alternating stages, each beginning with the division of the target domain into easy-to-adapt and hard-to-adapt subdomains based on adaptation difficulty, followed by neighborhood-based pseudo label assignment. In stage one, we enhance classification accuracy through uncertainty-aware self-training and alignment of corresponding classes between subdomains. Stage two then applies tailored learning strategies to each subdomain, starting with consistency learning on the easy-to-adapt samples and progressing to utilizing local structural information for the more challenging ones, thereby mining the intrinsic properties of the target data. Extensive experiments on several widely used benchmarks validate the effectiveness of our approach, demonstrating superior performance compared to state-of-the-art methods. Jing Li 0132, Meng Zhao 0001, Wanli Xue, Qinghua Hu, Shengyong Chen |
IEEE Trans. Multim. | 6 |
| 2025 | CACP: Covariance-Aware Cross-Domain Prototypes for Domain Adaptive Semantic SegmentationabstractDomain adaptive semantic segmentation aims to reduce domain shifts / discrepancies between source and target domains, improving the source domain model's generalization ability to the target domain. Recently, prototypical methods, which primarily use single-source or single-target domain prototypes as category centers to aggregate features from both domains, have achieved competitive performance in this task. However, due to large domain shifts, single-source domain prototypes have finite generalization ability and not all source domain knowledge is conducive to model generalization. Single-target domain prototypes are noisy because they are prematurely initialized with all features filtered by pseudo labels, which causes error accumulation in the prototypes. To address these issues, we propose a covariance-aware cross-domain prototypes method (CACP) to achieve robust domain adaptation. We propose to use both domain prototypes to dynamically rectify pseudo labels in the target domain, effectively reducing the recognition difficulty of hard target domain samples and narrowing the gap between features of the same category in both domains. In addition, to further generalize the model to the target domain, we propose two modules based on covariance correlation, FSPC (Features Selection by Prototypes Covariances) and WSPC (Weighting Source by Prototypes Coefficients), to learn discriminative characteristics. FSPC selects highly correlated features to update target domain prototypes online, denoising and enhancing discriminativeness between categories. WSPC utilizes the correlation coefficients between target domain prototypes and source domain features to weight each point in the source domain, eliminating the information interference from the source domain. In particular, CACP achieves excellent performance on the GTA5$\to$Cityscapes and SYNTHIA$\to$Cityscapes tasks with minimal computational resources and time. Yanbing Xue, Feifei Zhang 0001, Xianbin Wen, Zan Gao 0002, Shengyong Chen |
IEEE Trans. Multim. | 6 |
| 2025 | A Semantic-Aware Attention and Visual Shielding Network for Cloth-Changing Person Re-IdentificationabstractCloth-changing person re-identification (ReID) is a newly emerging research topic that aims to retrieve pedestrians whose clothes are changed. Since the human appearance with different clothes exhibits large variations, it is very difficult for existing approaches to extract discriminative and robust feature representations. Current works mainly focus on body shape or contour sketches, but the human semantic information and the potential consistency of pedestrian features before and after changing clothes are not fully explored or are ignored. To solve these issues, in this work, a novel semantic-aware attention and visual shielding network for cloth-changing person ReID (abbreviated as SAVS) is proposed where the key idea is to shield clues related to the appearance of clothes and only focus on visual semantic information that is not sensitive to view/posture changes. Specifically, a visual semantic encoder is first employed to locate the human body and clothing regions based on human semantic segmentation information. Then, a human semantic attention (HSA) module is proposed to highlight the human semantic information and reweight the visual feature map. In addition, a visual clothes shielding (VCS) module is also designed to extract a more robust feature representation for the cloth-changing task by covering the clothing regions and focusing the model on the visual semantic information unrelated to the clothes. Most importantly, these two modules are jointly explored in an end-to-end unified framework. Extensive experiments demonstrate that the proposed method can significantly outperform state-of-the-art methods, and more robust features can be extracted for cloth-changing persons. Compared with multibiometric unified network (MBUNet) (published in TIP2023), this method can achieve improvements of 17.5% (30.9%) and 8.5% (10.4%) on the LTCC and Celeb-reID datasets in terms of mean average precision (mAP) (rank-1), respectively. When compared with the Swin Transformer (Swin-T), the improvements can reach 28.6% (17.3%), 22.5% (10.0%), 19.5% (10.2%), and 8.6% (10.1%) on the PRCC, LTCC, Celeb, and NKUP datasets in terms of rank-1 (mAP), respectively. Zan Gao 0001, Hongwei Wei, Weili Guan, Jie Nie, Meng Wang 0001, Shengyong Chen |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2025 | Learning Self-Corrective Network via Adaptive Self-Labeling and Dynamic NMS for High-Performance Long-Term TrackingabstractThis article presents a self-corrective network-based long-term tracker (SCLT) including a self-modulated tracking reliability evaluator (STRE) and a self-adjusting proposal postprocessor (SPPP). The targets in the long-term sequences often suffer from severe appearance variations. Existing long-term trackers often online update their models to adapt the variations, but the inaccurate tracking results introduce cumulative error into the updated model that may cause severe drift issue. To this end, a robust long-term tracker should have the self-corrective capability that can judge whether the tracking result is reliable or not, and then it is able to recapture the target when severe drift happens caused by serious challenges (e.g., full occlusion and out-of-view). To address the first issue, the STRE designs an effective tracking reliability classifier that is built on a modulation subnetwork. The classifier is trained using the samples with pseudo labels generated by an adaptive self-labeling strategy. The adaptive self-labeling can automatically label the hard negative samples that are often neglected in existing trackers according to the statistical characteristics of target state, and the network modulation mechanism can guide the backbone network to learn more discriminative features without extra training data. To address the second issue, after the STRE has been triggered, the SPPP follows it with a dynamic NMS to recapture the target in time and accurately. In addition, the STRE and the SPPP demonstrate good transportability ability, and their performance is improved when combined with multiple baselines. Compared to the commonly used greedy NMS, the proposed dynamic NMS leverages an adaptive strategy to effectively handle the different conditions of in view and out of view, thereby being able to select the most probable object box that is essential to accurately online update the basic tracker. Extensive evaluations on four large-scale and challenging benchmark datasets including VOT2021LT, OxUvALT, TLP, and LaSOT demonstrate superiority of the proposed SCLT to a variety of state-of-the-art long-term trackers in terms of all measures. Source codes and demos can be found at https://github.com/TJUT-CV/SCLT. Wanli Xue, Kaihua Zhang 0001, Bo Liu 0005, Chengwei Zhang 0001, Jingen Liu, Shengyong Chen |
IEEE Trans. Neural Networks Learn. Syst. | 7 |
| 2025 | TDSF-Net: Tensor Decomposition-Based Subspace Fusion Network for Multimodal Medical Image ClassificationabstractData from multimodalities bring complementary information for deep learning-based medical image classification models. However, data fusion methods simply concatenating features or images barely consider the correlations or complementarities among different modalities and easily suffer from exponential growth in dimensions and computational complexity when the modality increases. Consequently, this article proposes a subspace fusion network with tensor decomposition (TD) to heighten multimodal medical image classification. We first introduce a Tucker low-rank TD module to map the high-level dimensional tensor to the low-rank subspace, reducing the redundancy caused by multimodal data and high-dimensional features. Then, a cross-tensor attention mechanism is utilized to fuse features from the subspace into a high-dimension tensor, enhancing the representation ability of extracted features and constructing the interaction information among components in the subspace. Extensive comparison experiments with state-of-the-art (SOTA) methods are conducted on one self-established and three public multimodal medical image datasets, verifying the effectiveness and generalization ability of the proposed method. The code is available at https://github.com/1zhang-yi/TDSFNet. Yi Zhang 0111, Guoxia Xu, Meng Zhao 0001, Hao Wang 0003, Fan Shi 0001, Shengyong Chen |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2025 | Representation Learning Based on Co-Evolutionary Combined With Probability Distribution Optimization for Precise Defect LocationabstractVisual defect detection methods based on representation learning play an important role in industrial scenarios. Defect detection technology based on representation learning has made significant progress. However, existing defect detection methods still face three challenges: first, the extreme scarcity of industrial defect samples makes training difficult. Second, due to the characteristics of industrial defects, such as blur and background interference, it is challenging to obtain fuzzy defect separation edges and context information. Third, industrial defects cannot obtain accurate positioning information. This article proposes feature co-evolution interaction architecture (CIA) and glass container defect dataset to address the above challenges. Specifically, the contributions of this article are as follows: first, this article designs a glass container image acquisition system that combines RGB and polarization information to create a glass container defect dataset containing more than 60000 samples to alleviate the sample scarcity problem in industrial scenarios. Subsequently, this article designs the CIA. CIA optimizes the probability distribution of features through the co-evolution of edge and context features, thereby improving detection accuracy in blurred defects and noisy environments. Finally, this article proposes a novel inforced IoU loss (IIoU loss), which can obtain more accurate position information by being aware of the scale changes of the predicted box. Defect detection experiments in three mainstream industrial manufacturing categories (Northeastern University (NEU)-Det, glass containers, wood) show that CIA only uses 22.5 GFLOPs, and mean average precision (mAP) (NEU-Det: 88.74%, glass containers: 95.38%, wood: 68.42%) outperforms state-of-the-art methods. Jinglin Zhang 0001, Qinghui Chen, Gang Li 0005, Shijiao Ding, Maomao Xiong, Shengyong Chen |
IEEE Trans. Neural Networks Learn. Syst. | 9 |
| 2025 | A Novel Robustness-Enhancing Adversarial Defense Approach to AI-Powered Sea State Estimation for Autonomous Marine VesselsabstractSea state information is significant for the guide of maritime activities of autonomous vessels. The sea state estimation (SSE) model, powered by artificial intelligence (AI), has shown great effectiveness but is susceptible to malicious data attacks. These attacks can lead to significant declines in the system’s performance and result in incorrect predictions about the sea state. This study introduces SecureSSE, a strategy for protecting SSE models in autonomous marine vessels from adversarial attacks. This approach incorporates three main components: 1) the multiscale feature extraction learning (MFEL) module; 2) the feature convolution aggregation learning (FCAL) module; and 3) the perturbation examples training (PET) module. The PET module is specifically crafted to create perturbation examples that are in line with unaltered data, leveraging the capabilities of both the MFEL and FCAL modules to efficiently extract and integrate detailed features from ship motion data. Our proposed SecureSSE approach is shown to significantly improve the resilience of deep learning models against potential attacks. Through experimental testing, we have validated the effectiveness of this method in enhancing SSE. Additional ablation studies highlight the critical role of each module within the SecureSSE framework. To our knowledge, this is the first study to address adversarial attacks in this context and to propose a comprehensive defense mechanism for SSE systems in autonomous marine vessels. Xu Cheng 0003, Fan Shi 0001, Hanwei Zhang 0001, Hongning Dai, Houxiang Zhang, Shengyong Chen |
IEEE Trans. Syst. Man Cybern. Syst. | 7 |
| 2025 | Dual-stage temporal perception network for continuous sign language recognition
Wanli Xue, Jinlu Sun, Yazhou Wu, Tiantian Yuan, Shengyong Chen |
Vis. Comput. | 7 |
| 2024 | Intentional Evolutionary Learning for Untrimmed Videos with Long Tail DistributionabstractHuman intention understanding in untrimmed videos aims to watch a natural video and predict what the person’s intention is. Currently, exploration of predicting human intentions in untrimmed videos is far from enough. On the one hand, untrimmed videos with mixed actions and backgrounds have a significant long-tail distribution with concept drift characteristics. On the other hand, most methods can only perceive instantaneous intentions, but cannot determine the evolution of intentions. To solve the above challenges, we propose a loss based on Instance Confidence and Class Accuracy (ICCA), which aims to alleviate the prediction bias caused by the long-tail distribution with concept drift characteristics in video streams. In addition, we propose an intention-oriented evolutionary learning method to determine the intention evolution pattern (from what action to what action) and the time of evolution (when the action evolves). We conducted extensive experiments on two untrimmed video datasets (THUMOS14 and ActivityNET v1.3), and our method has achieved excellent results compared to SOTA methods. The code and supplementary materials are available at https://github.com/Jennifer123www/UntrimmedVideo. Xiujie Wang, Jianhua Zhang 0002, Shengyong Chen |
AAAI | 8 |
| 2024 | Cross-Modality Consistency Mining For Continuous Sign Language Recognition with Text-Domain EquivalentsabstractContinuous Sign Language Recognition (CSLR) approaches share similarities with conventional NLP approaches in which language understanding is involved. However, CSLR approaches face the significant challenge of limited scale and vocabulary in existing datasets. Unlike language models that benefit from extensive training datasets, CSLR models often contend with data constraints, hindering their ability to generalize effectively and consistently capture the rich sign language expressions. To leverage the strong contextual and memorial capabilities of pre-trained language models, in this work, we propose Cross-Modality Consistency (XMC) loss to mine the alignment between the visual model and pre-trained language model, enabling the direct alignment at the gloss level. Towards this, we construct a small-scale gloss description corpus named DCSLG with rich descriptive text. Accompanied by CSL-Daily, text-domain equivalents are made for each video in the dataset, making fine-level alignment possible. The experiment results show that the proposed XMC Loss significantly improves the activations, producing more spatio-temporally accurate and relevant activations. Our approach achieves an average reduction in WER by 3.5%. The code and our gloss description corpus named DCSLG are made publicly available on GitHub1. Zhenghao Ke, Sheng Liu 0002, Chengyuan Ke, Yuan Feng 0002, Shengyong Chen |
ICME | 5 |
| 2024 | DA-LGNet: Enhancing Spatial-Spectral feature representation with Dual-Attention Local-General Network for Hyperspectral images and Multispectral images FusionabstractHyperspectral image and Multispectral image (HSI-MSI) fusion aims to fuse a registered high-resolution multi-spectral image (HR-MSI) with a low-resolution hyperspectral image (LR-HSI) to generate a spatially enhanced HSI with high spectral resolution. Current fusion methods often make insufficient utilization of spatial and spectral prior information, including spatial self-similarity and inter-spectral correlations, resulting in a degradation in image fidelity. Therefore, we propose a novel HSI-MSI fusion network, called DA-LGNet, designed to learn spatial-spectral priors with a dual-attention mechanism for feature enhancement and utilize the advantages of the large kernel attention for global and local feature representation. Specifically, the large kernel attention module (LKAM) can efficiently extract and integrate long-range dependencies and detailed textual information. The dual-attention enhancement module (DAEM) incorporates the position attention mechanism with the channel attention mechanism, flexibly capturing spatial and spectral prior information, which are crucial for restoring high-resolution HSI (HR-HSI). Extensive experiments on two datasets demonstrate that our DA-LGNet importantly outper-forms other state-of-the-art methods. Haozheng Zhang, Yanhong Yang, Zhixuan Jing, Shengyong Chen |
ICME | 4 |
| 2024 | Semi-Supervised Camouflaged Object Detection: Multi Information Fusion Combined with Adaptive Receptive Field Selection Network
Feng Xiao 0005, Ruyu Liu, Jianhua Zhang 0002, Shengyong Chen |
PRCV (12) | 6 |
| 2024 | DSNet: A dynamic squeeze network for real-time weld seam image segmentation
Fan Shi 0001, Mounir Kaaniche, Meng Zhao 0001, Yan Jing, Shengyong Chen |
Eng. Appl. Artif. Intell. | 7 |
| 2024 | DFCNet +: Cross-modal dynamic feature contrast net for continuous sign language recognition
Yuan Feng 0002, Nuoyi Chen, Yumeng Wu, Caoyu Jiang, Sheng Liu 0002, Shengyong Chen |
Image Vis. Comput. | 6 |
| 2024 | MLKAF-Net: Multiscale Large Kernel Attention Network for Hyperspectral and Multispectral Image FusionabstractThe fusion of a low spatial resolution hyperspectral image (LR-HSI) with a high spatial resolution multispectral image (HR-MSI) aims to synthesize a high-resolution hyperspectral image (HR-HSI), enabling a broader range of applications for hyperspectral images (HSIs). However, existing fusion methods struggle to capture both long-range dependencies and fine-grained spatial features, resulting in block artifacts and spatial distortions in the reconstructed HR-HSIs. Therefore, we introduce MLKAF-Net, a multiscale HSI-MSI fusion method, which effectively formulates cross-modality fused features in both spatial and spectral domains. MLKAF-Net mainly consists of three modules: the multiscale large kernel attention module (MLKAM), the spatial information aggregation module (SIAM), and the spectral attention module (SPAM). Specifically, the MLKAM incorporates a multiscale mechanism into the large kernel decomposition, adaptively capturing both long-range dependencies and local granular information. We develop the SIAM to establish the spatial quality of the reconstructed HR-HSIs by aggregating abundant spatial information. The SPAM introduces the channel attention to effectively mitigate spectral distortion through preserving beneficial spectral information. Extensive experiments demonstrate that our MLKAF-Net importantly enhances the fusion performance compared to state-of-the-art methods. Haozheng Zhang, Yanhong Yang, Jianhua Zhang 0002, Shengyong Chen |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2024 | Identity-Guided Collaborative Learning for Cloth-Changing Person ReidentificationabstractCloth-changing person reidentification (ReID) is a newly emerging research topic aimed at addressing the issues of large feature variations due to cloth-changing and pedestrian view/pose changes. Although significant progress has been achieved by introducing extra information (e.g., human contour sketching information, human body keypoints, and 3D human information), cloth-changing person ReID remains challenging because pedestrian appearance representations can change at any time. Moreover, human semantic information and pedestrian identity information are not fully explored. To solve these issues, we propose a novel identity-guided collaborative learning scheme (IGCL) for cloth-changing person ReID, where the human semantic is effectively utilized and the identity is unchangeable to guide collaborative learning. First, we design a novel clothing attention degradation stream to reasonably reduce the interference caused by clothing information where clothing attention and mid-level collaborative learning are employed. Second, we propose a human semantic attention and body jigsaw stream to highlight the human semantic information and simulate different poses of the same identity. In this way, the extraction features not only focus on human semantic information that is unrelated to the background but are also suitable for pedestrian pose variations. Moreover, a pedestrian identity enhancement stream is proposed to enhance the identity importance and extract more favorable identity robust features. Most importantly, all these streams are jointly explored in an end-to-end unified framework, and the identity is utilized to guide the optimization. Extensive experiments on six public clothing person ReID datasets (LaST, LTCC, PRCC, NKUP, Celeb-reID-light, and VC-Clothes) demonstrate the superiority of the IGCL method. It outperforms existing methods on multiple datasets, and the extracted features have stronger representation and discrimination ability and are weakly correlated with clothing. Zan Gao 0001, Shengxun Wei, Weili Guan, Lei Zhu 0002, Meng Wang 0001, Shengyong Chen |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2024 | Hunt-inspired Transformer for visual object tracking
Wanli Xue, Kaihua Zhang 0001, Shengyong Chen |
Pattern Recognit. | 5 |
| 2024 | Rethinking Inconsistent Context and Imbalanced Regression in Depression Severity PredictionabstractAs one of the world's most prevalent mental illnesses, depression is not easy to detect since it affects different people in different ways. Recently, linguistic features extracted from transcribed texts have been widely explored in depression detection because they contain a variety of cues about psychological activities. However, the detection performance is limited due to the following two reasons: 1) the dialogue structure is ignored, which causes the Inconsistent Context problem; and 2) Imbalanced Regression occurs due to the long-tailed distribution of depression datasets. To this end, in this paper we investigate the relationship between the local topic and global context in interview transcripts, and bridge the gap between depression symptoms and depression severity. In particular, we propose a model called Conditional Variational Topic-enriched Auto-Encoder (CVTAE), which can capture the spatial features from local topics via variational inference, and the temporal features from the global context with attention mechanism. Besides, we apply the re-weighting strategies to assigning weights to the depression labels with different values. Extensive experiments on the DAIC-WOZ dataset in English and a self-constructed database NCUDID in Chinese demonstrate the effectiveness and robustness of CVTAE, while the comprehensive ablation study and case study show its interpretability. Guanhe Huang, Jing Li 0027, Heli Lu, Shengyong Chen |
IEEE Trans. Affect. Comput. | 5 |
| 2024 | A Semantic Perception and CNN-Transformer Hybrid Network for Occluded Person Re-IdentificationabstractThe objective of the occluded person re-identification (ReID) task is to capture the same person from different camera angles when the pedestrian’s body is partially occluded. In this task, there are two main challenges: 1) pedestrians are often occluded by other persons or objects, and 2) pedestrians change poses. Moreover, these two issues often simultaneously occur. Although many occluded person ReID algorithms have been proposed, many existing methods can often only solve one of these issues well, and the other issue is often ignored. In this work, a novel semantic perception and CNN-transformer hybrid network (abbreviated as SPH) is proposed for occluded person ReID, which consists of a CNN-based human semantic perception stream and a transformer-based pose perception stream. In the former, a human semantic auxiliary module and a human semantic perception module are designed to obtain human semantic information where multi-granularity region features of the human body are extracted to solve the issues of occlusion. In the latter, we propose a token-based pose integration module to obtain the corresponding patch for each pose key-point and the relative position information to solve the change in pedestrian pose. Moreover, these two streams are jointly optimized in a unified framework. In addition, to further solve the issue of occlusion, the human completion strategy is proposed for the query sample where the gallery samples are used to complete the missing parts of the query. Extensive experimental results on three public occluded person ReID datasets, Occluded-DukeMTMC, P-DukeMTMC-reID, and Occluded-REID, demonstrate that the proposed method can outperform all SOTA occluded person ReID methods in terms of the mAP and Rank-1. Compared with PAT (CVPR21) on the Occluded-DukeMTMC and Occluded-REID datasets, the improvements in mAP/Rank-1 reached 10.1%/7.4%, and 10%/1%, respectively. Moreover, when TransReID (ICCV21) was used, SPH achieved improvements of 4.5% (mAP) and 5.5% (Rank-1) on the Occluded-DukeMTMC dataset. Zan Gao 0001, Peng Chen 0047, Tao Zhuo, Meng Liu 0006, Lei Zhu 0002, Meng Wang 0001, Shengyong Chen |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2024 | A Snippets Relation and Hard-Snippets Mask Network for Weakly-Supervised Temporal Action LocalizationabstractWeakly-supervised temporal action localization (WTAL) is a problem learning an action localization model with only video-level labels available. In recent years, many WTAL methods have developed. However, hard-to-predict snippets near action boundaries are often not considered in these existing approaches, causing action incompleteness and action over-complete issues. To solve these issues, in this work, an end-to-end snippets relation and hard-snippets mask network (SRHN) is proposed. Specifically, a hard-snippets mask module is applied to mask the hard-to-predict snippets adaptively, and in this way, the trained model focuses more on those snippets with low uncertainty. Then, a snippets relation module is designed to capture the relationship among snippets and can make hard-to-predict snippets easy to predict by aggregating the information of multiple temporal receptive fields. Finally, a snippet enhancement loss is further developed to reduce the action probabilities that are not present in videos for hard-to-predict snippets and other snippets, enlarging the action probabilities that exist in videos. Extensive experiments on THUMOS14, ActivityNet1.2, and ActivityNet1.3 datasets demonstrate the effectiveness of the SRHN method. Yibo Zhao 0001, Hua Zhang 0003, Zan Gao 0001, Weili Guan, Meng Wang 0001, Shengyong Chen |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2024 | Gaze Estimation by Attention-Induced Hierarchical Variational Auto-EncoderabstractAppearance-based gaze estimation has been widely studied recently with promising performance. The majority of appearance-based gaze estimation methods are developed under the deterministic frameworks. However, the deterministic gaze estimation methods suffer from large performance drop upon challenging eye images in low-resolution, darkness, partial occlusions, etc. To alleviate this problem, in this article, we alternatively reformulate the appearance-based gaze estimation problem under a generative framework. Specifically, we propose a variational inference model, that is, variational gaze estimation network (VGE-Net), to generate multiple gaze maps as complimentary candidates simultaneously supervised by the ground-truth gaze map. To achieve robust estimation, we adaptively fuse the gaze directions predicted on these candidate gaze maps by a regression network through a simple attention mechanism. Experiments on three benchmarks, that is, MPIIGaze, EYEDIAP, and Columbia, demonstrate that our VGE-Net outperforms state-of-the-art gaze estimation methods, especially on challenging cases. Comprehensive ablation studies also validate the effectiveness of our contributions. The code will be publicly released. Guanhe Huang, Jingyue Shi, Jun Xu 0019, Jing Li 0027, Shengyong Chen, Yingjun Du, Xiantong Zhen, Honghai Liu 0001 |
IEEE Trans. Cybern. | 5 |
| 2024 | High-Order Spatial Interactions Enhanced Lightweight Model for Optical Remote Sensing Image-Based Small Ship DetectionabstractAccurate and reliable optical remote sensing image-based small-ship detection is crucial for maritime surveillance systems, but existing methods often struggle with balancing detection performance and computational complexity. In this article, we propose a novel lightweight framework called HSI-ShipDetectionNet that is based on high-order spatial interactions (HSIs) and is suitable for deployment on resource-limited platforms, such as satellites and unmanned aerial vehicles. HSI-ShipDetectionNet includes a prediction branch specifically for tiny ships and a lightweight hybrid attention block (LHAB) for reduced complexity. In addition, the use of an HSI module improves advanced feature understanding and modeling ability. Our model is evaluated using the public Kaggle and FAIR1M marine ship detection datasets and compared with multiple state-of-the-art models including small object detection models, lightweight detection models, and ship detection models. The results show that HSI-ShipDetectionNet outperforms the other models in terms of detection performance while being lightweight and suitable for deployment on resource-limited platforms. Xu Cheng 0003, Fan Shi 0001, Xiufeng Liu 0001, Huan Huo, Shengyong Chen |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2024 | Prediction of LncRNA-Protein Interactions Based on Kernel Combinations and Graph Convolutional NetworksabstractThe complexes of long non-coding RNAs bound to proteins can be involved in regulating life activities at various stages of organisms. However, in the face of the growing number of lncRNAs and proteins, verifying LncRNA-Protein Interactions (LPI) based on traditional biological experiments is time-consuming and laborious. Therefore, with the improvement of computing power, predicting LPI has met new development opportunity. In virtue of the state-of-the-art works, a framework called LncRNA-Protein Interactions based on Kernel Combinations and Graph Convolutional Networks (LPI-KCGCN) has been proposed in this article. We first construct kernel matrices by taking advantage of extracting both the lncRNAs and protein concerning the sequence features, sequence similarity features, expression features, and gene ontology. Then reconstruct the existent kernel matrices as the input of the next step. Combined with known LPI interactions, the reconstructed similarity matrices, which can be used as features of the topology map of the LPI network, are exploited in extracting potential representations in the lncRNA and protein space using a two-layer Graph Convolutional Network. The predicted matrix can be finally obtained by training the network to produce scoring matrices w.r.t. lncRNAs and proteins. Different LPI-KCGCN variants are ensemble to derive the final prediction results and testify on balanced and unbalanced datasets. The 5-fold cross-validation shows that the optimal feature information combination on a dataset with 15.5% positive samples has an AUC value of 0.9714 and an AUPR value of 0.9216. On another highly unbalanced dataset with only 5% positive samples, LPI-KCGCN also has outperformed the state-of-the-art works, which achieved an AUC value of 0.9907 and an AUPR value of 0.9267. Dongdong Mao, Jijun Tang, Zhijun Liao, Shengyong Chen |
IEEE J. Biomed. Health Informatics | 5 |
| 2024 | Selective Feature Fusion and Irregular-Aware Network for Pavement Crack DetectionabstractRoad cracks on highways and main roads are among the most prominent defects. Given the inherent inaccuracy, time-consuming nature, and labor intensiveness of manual road crack detection, there’s a compelling need for automated solutions. The irregular shape of cracks, along with complex background conditions encompassing varying lighting, tree shadows, and dark stains, poses a significant challenge for computer vision-based approaches. Most cracks exhibit irregular edge patterns, which are pivotal features for accurate detection. In response to recent advancements in deep learning within the realm of computer vision, this paper introduces an innovative neural network architecture termed the ‘Selective Feature Fusion and Irregular-Aware Network (SFIAN)’ designed specifically for crack detection on pavements. The proposed network selectively integrates features from multiple levels, enhancing and controlling the flow of valuable information at each stage while effectively modeling irregular crack objects. In an extensive evaluation, this paper conducts experiments on five distinct crack datasets and compares the results with twelve state-of-the-art crack detection methods, including the latest edge detection and semantic segmentation techniques. The experimental findings demonstrate the superior performance of the proposed method, surpassing baseline methods by a notable margin, with an increase of approximately 13.3% in the F1-score, all without introducing additional time complexity. Furthermore, the model achieves real-time processing, achieving a remarkable speed of 35 frames per second (FPS) on images at 320$\times$480 pixels, facilitated by NVIDIA 3090 hardware. Xu Cheng 0003, Fan Shi 0001, Meng Zhao 0001, Xiufeng Liu 0001, Shengyong Chen |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2024 | Asymmetric Dual-Decoder U-Net for Joint Rain and Haze RemovalabstractThis work studies the multi-weather restoration problem. In real-life scenarios, rain and haze, two often co-occurring common weather phenomena, can greatly degrade the clarity and quality of the scene images, leading to a performance drop in the visual applications, such as autonomous driving. However, jointly removing the rain and haze in scene images is ill-posed and challenging, where the existence of haze and rain and the change of atmosphere light, can both degrade the scene information. Current methods focus on the contamination removal part, thus ignoring the restoration of the scene information affected by the change of atmospheric light. We propose a novel deep neural network, named Asymmetric Dual-decoder U-Net (ADU-Net), to address the aforementioned challenge. The ADU-Net produces both the contamination residual and the scene residual to efficiently remove the contamination while preserving the fidelity of the scene information. Extensive experiments show our work outperforms the existing state-of-the-art methods by a considerable margin in both synthetic data and real-world data benchmarks, including RainCityscapes, BID Rain, and SPA-Data. For instance, we improve the state-of-the-art PSNR value by 2.26/4.57 on the RainCityscapes/SPA-Data, respectively. Codes will be made available freely to the research community. Yuan Feng 0002, Yaojun Hu, Pengfei Fang, Sheng Liu 0002, Yanhong Yang, Shengyong Chen |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2024 | A Prototype-Empowered Kernel-Varying Convolutional Model for Imbalanced Sea State Estimation in IoT-Enabled Autonomous ShipabstractSea State Estimation (SSE) is essential for Internet of Things (IoT)-enabled autonomous ships, which rely on favorable sea conditions for safe and efficient navigation. Traditional methods, such as wave buoys and radars, are costly, less accurate, and lack real-time capability. Model-driven methods, based on physical models of ship dynamics, are impractical due to wave randomness. Data-driven methods are limited by the data imbalance problem, as some sea states are more frequent and observable than others. To overcome these challenges, we propose a novel data-driven approach for SSE based on ship motion data. Our approach consists of three main components: a data preprocessing module, a parallel convolution feature extractor, and a theoretical-ensured distance-based classifier. The data preprocessing module aims to enhance the data quality and reduce sensor noise. The parallel convolution feature extractor uses a kernel-varying convolutional structure to capture distinctive features. The distance-based classifier learns representative prototypes for each sea state and assigns a sample to the nearest prototype based on a distance metric. The efficiency of our model is validated through experiments on two SSE datasets and the UEA archive, encompassing thirty multivariate time series classification tasks. The results reveal the generalizability and robustness of our approach. Mengna Liu, Xu Cheng 0003, Fan Shi 0001, Xiufeng Liu 0001, Hongning Dai, Shengyong Chen |
IEEE Trans. Sustain. Comput. | 6 |
| 2024 | KSRB-Net: a continuous sign language recognition deep learning strategy based on motion perception mechanism
Feng Xiao 0005, Yunrui Zhu, Ruyu Liu, Jianhua Zhang 0002, Shengyong Chen |
Vis. Comput. | 5 |
| 2023 | Distilling Cross-Temporal Contexts for Continuous Sign Language RecognitionabstractContinuous sign language recognition (CSLR) aims to recognize glosses in a sign language video. State-of-the-art methods typically have two modules, a spatial perception module and a temporal aggregation module, which are jointly learned end-to-end. Existing results in [9, 20, 25, 36] have indicated that, as the frontal component of the over-all model, the spatial perception module used for spatial feature extraction tends to be insufficiently trained. In this paper, we first conduct empirical studies and show that a shallow temporal aggregation module allows more thor-ough training of the spatial perception module. However, a shallow temporal aggregation module cannot well capture both local and global temporal context information in sign language. To address this dilemma, we propose a cross-temporal context aggregation (CTCA) model. Specifically, we build a dual-path network that contains two branches for perceptions of local temporal context and global temporal context. We further design a cross-context knowledge distil-lation learning objective to aggregate the two types of con-text and the linguistic prior. The knowledge distillation en-ables the resultant one-branch temporal aggregation mod-ule to perceive local-global temporal and semantic context. This shallow temporal perception module structure facili-tates spatial perception module learning. Extensive exper-iments on challenging CSLR benchmarks demonstrate that our method outperforms all state-of-the-art methods. Leming Guo, Wanli Xue, Qing Guo 0005, Bo Liu 0005, Kaihua Zhang 0001, Tiantian Yuan, Shengyong Chen |
CVPR | 7 |
| 2023 | 'Skimming-Perusal' Detection: A Simple Object Detection Baseline in GigaPixel-level ImagesabstractObject detection has achieved amazing performance in regular-sized images, but with the emergence of gigapixel-level images, even the most advanced object detection methods cannot be directly used to process them quickly and efficiently. Therefore, this paper proposes a simple baseline for gigapixel-level images object detection called Skimming-Perusal Detection (SPDet). The SPDet consists mainly of two parts, a skimming model and a perusal model. The skimming model is based on an efficient global-to-local search strategy to detect possible regions containing objects. Non-object regions are merged through a skimming iterative merging strategy to generate skimming patch candidates. The perusal model adaptive scales the skimming patch candidates guided by the coarse detection of the skimming model. Extensive evaluations on the PANDA dataset demonstrate that the SPDet boosts detection speed on gigapixel-level images by 6× while achieving better performance than a variety of state-of-the-art methods. The source code is released at https://github.com/TJUT-CV/SPDet. Wanli Xue, Kaihua Zhang 0001, Shengyong Chen |
ICME | 4 |
| 2023 | Enhancing Ocean Scene Video Captioning with Multimodal Pre-Training and Video-Swin-TransformerabstractWith the success of multimodal pre-training models in the video-language field and various downstream tasks, previous multimodal models used 3DCNN networks as video feature extractors, which have limitations in interacting and fusing with text features. This paper proposes a multimodal pre-training model that utilizes a Video-Swin-Transformer-based network to encode both video and text data, to achieve better performance in video understanding. The model consists of four modules: video encoder, text encoder, interact encoder, and caption decoder to accomplish the task of ocean scene video captioning. A dataset of ocean scene videos, including various content types such as sea surfaces and shores, is also constructed. The training process is divided into two stages: pre-training and fine-tuning. Pre-training is performed on the Howto100m dataset to allow the model to learn video captions in natural scenes and complete video-language matching tasks. The fine-tuning stage is then performed on the ocean1000 dataset to better understand the events and content in ocean scene videos and generate captions that conform to ocean scene video descriptions. The model achieves satisfying results on both the public dataset YouCook2 and the proprietary dataset Ocean1000, demonstrating its ability in video-text information fusion and interaction. Meng Zhao 0001, Fan Shi 0001, Meng'en Zhang, Yu He 0001, Shengyong Chen |
IECON | 6 |
| 2023 | SANet: A novel segmented attention mechanism and multi-level information fusion network for 6D object pose estimation
Xinbo Geng, Fan Shi 0001, Xu Cheng 0003, Mianzhao Wang, Shengyong Chen, Hongning Dai |
Comput. Commun. | 6 |
| 2023 | Human Interaction Understanding With Consistency-Aware LearningabstractCompared with the progress made on human activity classification, much less success has been achieved on human interaction understanding (HIU). Apart from the latter task is much more challenging, the main causation is that recent approaches learn human interactive relations via shallow graphical representations, which are inadequate to model complicated human interactive-relations. This paper proposes a deep consistency-aware framework aiming at tackling the grouping and labelling inconsistencies in HIU. This framework consists of three components, including a backbone CNN to extract image features, a factor graph network to implicitly learn higher-order consistencies among labelling and grouping variables, and a consistency-aware reasoning module to explicitly enforcing consistencies. The last module is inspired by our key observation that the consistency-aware reasoning bias can be embedded into an energy function or a particular loss function, minimizing which delivers consistent predictions. An efficient mean-field inference algorithm is proposed, such that all modules of our network could be trained in an end-to-end fashion. Experimental results demonstrate that the two proposed consistency-learning modules complement each other, and both make considerable contributions in achieving leading performance on three benchmarks of HIU. The effectiveness of the proposed approach is further validated by experiments on detecting human-object interactions. Jiajun Meng, Zhenhua Wang 0003, Kaining Ying, Jianhua Zhang 0002, Dongyan Guo, Zhen Zhang 0008, Qinfeng Shi, Shengyong Chen |
IEEE Trans. Pattern Anal. Mach. Intell. | 8 |
| 2023 | Predicting Demands of COVID-19 Prevention and Control Materials via Co-Evolutionary Transfer LearningabstractThe novel coronavirus pneumonia (COVID-19) has created great demands for medical resources. Determining these demands timely and accurately is critically important for the prevention and control of the pandemic. However, even if the infection rate has been estimated, the demands of many medical materials are still difficult to estimate due to their complex relationships with the infection rate and insufficient historical data. To alleviate the difficulties, we propose a co-evolutionary transfer learning (CETL) method for predicting the demands of a set of medical materials, which is important in COVID-19 prevention and control. CETL reuses material demand knowledge not only from other epidemics, such as severe acute respiratory syndrome (SARS) and bird flu but also from natural and manmade disasters. The knowledge or data of these related tasks can also be relatively few and imbalanced. In CETL, each prediction task is implemented by a fuzzy deep contractive autoencoder (CAE), and all prediction networks are cooperatively evolved, simultaneously using intrapopulation evolution to learn task-specific knowledge in each domain and using interpopulation evolution to learn common knowledge shared across the domains. Experimental results show that CETL achieves high prediction accuracies compared to selected state-of-the-art transfer learning and multitask learning models on datasets during two stages of COVID-19 spreading in China. Qin Song, Yujun Zheng 0001, Weiguo Sheng 0001, Shengyong Chen |
IEEE Trans. Cybern. | 6 |
| 2023 | A Multitemporal Scale and Spatial-Temporal Transformer Network for Temporal Action LocalizationabstractTemporal action localization plays an important role in video analysis, which aims to localize and classify actions in untrimmed videos. Previous methods often predict actions on a feature space of a single temporal scale. However, the temporal features of a low-level scale lack sufficient semantics for action classification, while a high-level scale cannot provide the rich details of the action boundaries. In addition, the long-range dependencies of video frames are often ignored. To address these issues, a novel multitemporal-scale spatial–temporal transformer (MSST) network is proposed for temporal action localization, which predicts actions on a feature space of multiple temporal scales. Specifically, we first use refined feature pyramids of different scales to pass semantics from high-level scales to low-level scales. Second, to establish the long temporal scale of the entire video, we use a spatial–temporal transformer encoder to capture the long-range dependencies of video frames. Then, the refined features with long-range dependencies are fed into a classifier for coarse action prediction. Finally, to further improve the prediction accuracy, we propose a frame-level self-attention module to refine the classification and boundaries of each action instance. Most importantly, these three modules are jointly explored in a unified framework, and MSST has an anchor-free and end-to-end architecture. Extensive experiments show that the proposed method can outperform state-of-the-art approaches on the THUMOS14 dataset and achieve comparable performance on the ActivityNet1.3 dataset. Compared with A2Net (TIP20, Avg{0.3:0.7}), Sub-Action (CSVT2022, Avg{0.1:0.5}), and AFSD (CVPR21, Avg{0.3:0.7}) on the THUMOS14 dataset, the proposed method can achieve improvements of 12.6%, 17.4%, and 2.2%, respectively. Zan Gao 0001, Xinglei Cui, Tao Zhuo, Zhiyong Cheng 0001, Anan Liu, Meng Wang 0001, Shengyong Chen |
IEEE Trans. Hum. Mach. Syst. | 7 |
| 2023 | Visual Object Tracking Based on Light-Field Imaging in the Presence of Similar DistractorsabstractVisual object tracking is of great importance in the field of computer vision. One of the main challenges is the difficulty of identifying moving targets from nearby similar distractors with a single-view image of the scene. To overcome this challenge, in this article, we acquire multiview images of the scenes by using a light-field camera. The multiview images are able to capture the 4-D structure instead of the 2-D plane of the objects but are more difficult to process. Therefore, we propose a novel representation for multiview images, i.e., the macro-epipolar plane image (macro-EPI), which highlights both spatial topological and angular information of the target and distractors. It is obtained by slicing the original multiview images into pieces and properly restacking these pieces in an ordinal manner. The resulting macro-EPI is mapped into the 2-D space; therefore, we adapt a modified autoencoder network to train a macro-EPI feature extractor. Thereafter, we design a composite framework of two-pattern convolution filters based on a discriminative correlation filter for object tracking, which successfully discriminates the target from the distractors by merging the macro-EPI features and the single-view image features. The experiments also show that our method outperforms the state-of-the-art methods in the presence of similar distractors. Mianzhao Wang, Fan Shi 0001, Xu Cheng 0003, Meng Zhao 0001, Yao Zhang 0021, Shengyong Chen |
IEEE Trans. Ind. Informatics | 8 |
| 2023 | A Novel Class-Imbalanced Ship Motion Data-Based Cross-Scale Model for Sea State EstimationabstractSea state estimation (SSE) is significant to the development of autonomous ships, which can enhance the sustainable development of maritime transportation. Traditional model-based methods are limited by their drawbacks, such as high costs and inaccurate estimations. The deep learning model shows superior performance, but it requires that the sample quantity for each sea state should be almost the same. Since the occurrence probability of each state is different, the ships mainly work in low sea states, and the collected ship motion data for different sea states are highly imbalanced. This work proposes a novel class-imbalanced ship motion data-based cross-scale model for SSE. The model consists of three major components: a multi-scale feature learning module, a cross-scale feature learning module, and a prototype classifier module. The multi-scale and cross-scale feature learning modules are designed to learn abundant coarse and fine-level features from the ship motion data. The prototype classifier is utilized to overcome the limitation of the conventional softmax classifier to produce better estimates. Our research highlights our model’s remarkable scalability and versatility with 30 publicly available datasets in time series classification, demonstrating superior performance over baseline methods in 21 cases. Notably, it outperformed ShapeNet by 5.72% and EDI by 26.3%. We further validated our model’s proficiency using ship motion datasets, consistently surpassing eight state-of-the-art baselines and five class-imbalanced learning methods. Ablation and sensitivity studies, emphasize the critical role of each model component. Our findings underscore the model’s robustness and its potential to advance time series classification in diverse domains. Xu Cheng 0003, Xiufeng Liu 0001, Fan Shi 0001, Zhengru Ren, Shengyong Chen |
IEEE Trans. Intell. Transp. Syst. | 7 |
| 2023 | Coordination and Optimization Control Framework for Vessels Platooning in Inland Waterborne Transportation SystemabstractVessels sailing in a single platoon could reduce resistance from the perspective of the whole platoon and the individual vessel, and contribute to improving energy benefits. Moreover, transportation energy costs and traffic efficiency are essential indicators for measuring waterborne transportation systems. We attempt to minimize transportation energy costs by coordinating platoon formation using a distributed framework of controllers. A large-scale coordinated vessel platooning program is proposed to minimize transportation energy costs and optimize traffic efficiency while guaranteeing safety. The control framework covers routing, energy consumption-dependent cooperative platooning decision and speed optimization based on graph search algorithm, cluster analysis, optimal control approach and model predictive control. Firstly, a local scheduling strategy combined with the leader vessel selection algorithm is adopted. Furthermore, we used cluster analysis to create a series of mergeable vessel platooning sets. Then, we used the mathematical planning method and a two-step hybrid optimal control approach to calculate the improvement and optimization of each vessel platoon’s path and speed. Finally, the scalability of the scheduling strategy is elucidated. In a simulation of large scale inland waterborne network, savings surpassed 3.5% when six hundreds vessels participated in the system. These simulation results reveal that the scheduling strategy coordinating vessels into vessel platooning, which improves transportation efficiency as well as descends cost, comparing to a fixed origin route in the waterway network. Man Zhu, Shengyong Chen, Xu Cheng 0003, Yuanqiao Wen, Weidong Zhang 0004, Rudy R. Negenborn, Yusong Pang |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2023 | DSP-Based Traffic Target Detection for Intelligent TransportationabstractInternet of Things (IoT)-based intelligent transportation is attracting more and more attention. As a key component of intelligent transportation, traffic video monitoring is very important, in which vehicle and pedestrian detection on the road is a crucial task. Although vehicle and pedestrian detection through deep learning (DL) may achieve high accuracy, it tends to require high computing resources, which hinders its use on IoT devices. As an important class of IoT devices, digital signal processor (DSP) has the characteristics of low energy consumption, small size, and strong performance, which has been widely used in intelligent transportation. In order to use DL on DSP for accurate vehicle and pedestrian detection, we first propose a series of general tactics to optimize the object detection convolutional neural network (CNN) model, including convolution layer optimization, cache optimization, compiler optimization, intrinsics optimization and direct memory access (DMA) acceleration, and then a parallel scheme to extend the model to run on multicore, and further quantize the implementation of the model. We evaluate it on UA-DETRAC and KITTI datasets. Experimental results show that our method achieves a faster speed than running the same CNN model on a mainstream desktop CPU, with only 0.06% accuracy loss. Jianhua Zhang 0002, Rucen Wang, Ruyu Liu, Dongyan Guo, Bo Li 0090, Shengyong Chen |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2023 | A Novel Action Saliency and Context-Aware Network for Weakly-Supervised Temporal Action LocalizationabstractTemporal action localization is a challenging task in computer vision, and it tries to find the start time and the end time of the actions and predict their categories. However, compared to temporal action localization, weakly supervised temporal action localization (WTAL) is a more challenging task due to its poor annotations. With only video-level annotation, some background frames, similar to actions, would be classified as actions and produce inaccurate results. In addition, the two-stream fusion problem, ignored previously, also needs to be further considered. To resolve these issues, we propose a novel action saliency and context-aware network (ASCN) for weakly supervised temporal action localization tasks. Specifically, the temporal saliency and context module is designed to enhance the global saliency and context information of the RGB and the flow features to suppress the backgrounds and enhance the actions. In addition, a hybrid attention mechanism using frame differences and two-stream attention is designed to model the local action context information and further enlarge the scores of the potential action regions and suppress the background regions. Finally, to obtain two-stream consistency and solve the fusion problem, we use the similarity loss and a channel self-attention module to adaptively fuse the enhanced RGB and flow features. Extensive experiments demonstrate that ASCN can outperform all of the SOTA WTAL methods on the THUMOS14 dataset and the ActivityNet1.3 dataset with an average mAP that can reach 37.2% on the THUMOS14 dataset and attains an average mAP of 26.3% on the ActivityNet1.3 dataset. On the ActivityNet1.2 dataset, ASCN can also obtain comparable results. Compared with AdapNet (TNNLS20), MMSD (TIP22), and FTCL (CVPR22) on the THUMOS14 dataset, ASCN can outperform them by 13.5%, 2.9%, and 2.8%, respectively. Yibo Zhao 0001, Hua Zhang 0003, Zan Gao 0001, Wen Gao 0001, Meng Wang 0001, Shengyong Chen |
IEEE Trans. Multim. | 6 |
| 2023 | Adaptively Customizing Activation Functions for Various LayersabstractTo enhance the nonlinearity of neural networks and increase their mapping abilities between the inputs and response variables, activation functions play a crucial role to model more complex relationships and patterns in the data. In this work, a novel methodology is proposed to adaptively customize activation functions only by adding very few parameters to the traditional activation functions such as Sigmoid, Tanh, and rectified linear unit (ReLU). To verify the effectiveness of the proposed methodology, some theoretical and experimental analysis on accelerating the convergence and improving the performance is presented, and a series of experiments are conducted based on various network models (such as AlexNet, VggNet, GoogLeNet, ResNet and DenseNet), and various datasets (such as CIFAR10, CIFAR100, miniImageNet, PASCAL VOC, and COCO). To further verify the validity and suitability in various optimization strategies and usage scenarios, some comparison experiments are also implemented among different optimization strategies (such as SGD, Momentum, AdaGrad, AdaDelta, and ADAM) and different recognition tasks such as classification and detection. The results show that the proposed methodology is very simple but with significant performance in convergence speed, precision, and generalization, and it can surpass other popular methods such as ReLU and adaptive functions such as Swish in almost all experiments in terms of overall performance. Haigen Hu, Aizhu Liu, Qiu Guan, Hanwang Qian, Xiaoxin Li 0001, Shengyong Chen, Qianwei Zhou |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2022 | PA-AWCNN: Two-stream Parallel Attention Adaptive Weight Network for RGB-D Action RecognitionabstractDue to overly relying on appearance information or adopting direct static feature fusion, most of the existing action recognition methods based on multi-modality have poor robustness and insufficient consideration of modality differences. To address these problems, we propose a two-stream adaptive weight integration network with a three-dimensional parallel attention module, PA-AWCNN. Firstly, a three-dimensional Parallel Attention (PA) module is proposed to effectively extract features of spatial, temporal and channel dimensions and reduce the cross-dimensional interference, to achieve better robustness. Secondly, a Common Feature-driven (CFD) feature integration module is proposed to dynamically integrate appearance and depth features with adaptive weights, utilizing modality differences to redeem the lack of each feature, thereby balance the influence of both. The proposed PA-AW CNN uses the representative integrated feature generated by attention enhancement and feature integration for action recognition; it can not only get higher recognition accuracy but also improve the performance of distinguishing similar actions. Experiments illustrate that the proposed method achieves com-parable performances to state-of-the-art methods and obtains the accuracy of 92.76% and 95.65% on NTU RGB+D Dataset and SBU Kinect Interaction Dataset, respectively. The code is publicly available at: https://github.com/Luu-Yao/PA-AWCNN. Sheng Liu 0002, Chaonan Li, Siyu Zou, Shengyong Chen, Diyi Guan |
ICRA | 5 |
| 2022 | HMD-former: a Transformer-based Human Mesh Deformer with Inter-layer Semantic ConsistencyabstractWe present a transformer-based network, Human Mesh Deformer (HMD-former), to tackle the problem of 3D human mesh reconstruction from a single RGB image. HMD-former applies a pre-trained CNN to extract image grid features and a transformer decoder to gradually warp the template 3D mesh to the deformed mesh. On each decoder layer, the fine-grained local information of grid features is well utilized using cross-attention by softly and content-dependently transforming the grid features to vertex embeddings. Auxiliary losses and proposed bi-directional mapping layers inherently ensure semantic consistency throughout the whole decoder, which free the network from learning unnecessary embedding transformation between layers. This further induces each layer of the decoder to focus on refining vertex embeddings and makes the whole network work in a progressively refining manner. Experiments on different public datasets Human3.6M and 3DPW show better reconstruction accuracy and faster inference speed than previous state-of-the-art methods, demonstrating the effectiveness and generalizability of HMD-former. Code is publicly available at https://github.com/siyuzou/HMD-former. Siyu Zou, Sheng Liu 0002, Chaonan Li, Shengyong Chen |
ICRA | 5 |
| 2022 | Accelerating Motion Perception Model Mimics the Visual Neuronal Ensemble of CrababstractIn nature, crabs have a panoramic vision for the localization and perception of accelerating motion from local segments to global view in order to guide reactive behaviours including escape. The visual neuronal ensemble in crab plays crucial roles in such capability, however, has never been investigated and modelled as an artificial vision system. To bridge this gap, we propose an accelerating motion perception model (AMPM) mimicking the visual neuronal ensemble in crab. The AMPM includes two main parts, wherein the pre-synaptic network from the previous modelling work simulates 16 MLGI neurons covering the entire view to localize moving objects. The emphasis herein is laid on the original modelling of MLGIs' post-synaptic network to perceive accelerating motions from a global view, which employs a novel spatial-temporal difference encoder (STDE), and an adaptive spiking threshold temporal difference encoder (AT-TDE). Specifically, the STDE transforms “time-to-travel” between activations of two successive segments of MLG1 into excitatory post-synaptic current (EPSC), which decays with the elapse of time. The AT-TDE in two directional, i.e., counter-clockwise and clockwise accelerating detectors guarantees “non-firing” to constant movements. Accordingly, the accelerating motion can be effectively localized and perceived by the whole network. The systematic experiments verified the feasibility and robustness of the proposed method. The model responses to translational accelerating motion also fit many of the explored physiological features of direction selective neurons in the lobula complex of crab (i.e. lobula complex direction cells, LCDCs). This modelling study not only provides a reasonable hypothesis for such biological neural pathways, but is also critical for developing a new neuromorphic sensor strategy. Mu Hua, Shigang Yue, Shengyong Chen, Qinbing Fu |
IJCNN | 5 |
| 2022 | Human Interaction Recognition with Skeletal Attention and Shift Graph ConvolutionabstractHuman interaction recognition has wide applications including intelligent surveillance, intelligent transportation and the analysis of sports videos. In recent years, benefiting from the development of action recognition based on deep learning, the performance of human interaction recognition has been boosted. This paper tackles two vital issues in recognizing human interactions, namely target missing and inadequate feature expression. To this end, we first design a data preprocessing method using skeleton estimation and multi-object tracking, which effectively reduces the chance of missing detection. Second, we propose a two-stream network composing of an appearance branch and a pose branch. The appearance branch extracts features enhanced via part affinity maps and part confidences maps, while the pose branch trains a customized Shift-GCN to extract skeletal features from people-pairs. Appearance and pose features are then fused to generate a more powerful representation of human interactions. Extensive experiments on two existing benchmarks, UT and BIT-Interaction, as well as a new dataset crafted by us, namely Campus-Interaction (CI), demonstrate the superior performance of the proposed approach over the state-of-the-arts. Zhenhua Wang 0003, Jiajun Meng, Sheng Liu 0002, Jianhua Zhang 0002, Shengyong Chen |
IJCNN | 6 |
| 2022 | LFBCNet: Light Field Boundary-aware and Cascaded Interaction Network for Salient Object DetectionabstractIn light field imaging techniques, the abundance of stereo spatial information aids in improving the performance of salient object detection. In some complex scenes, however, applying the 4D light field boundary structure to discriminate salient objects from background regions is still under-explored. In this paper, we propose a light field boundary-aware and cascaded interaction network based on light field macro-EPI, named LFBCNet. Firstly, we propose a well-designed light field multi-epipolar-aware learning (LFML) module to learn rich salient boundary cues by perceiving the continuous angle changes from light field macro-EPI. Secondly, to fully excavate the correlation between salient objects and boundaries at different scales, we design multiple light field boundary interactive (LFBI) modules and cascade them to form a light field multi-scale cascade interaction decoder network. Each LFBI is assigned to predict exquisite salient objects and boundaries by interactively transmitting the salient object and boundary features. Meanwhile, the salient boundary features are forced to gradually refine the salient object features during the multi-scale cascade encoding. Furthermore, a light field multi-scale-fusion prediction (LFMP) module is developed to automatically select and integrate multi-scale salient object features for final saliency prediction. The proposed LFBCNet can accurately distinguish tiny differences between salient objects and background regions. Comprehensive experiments on large benchmark datasets prove that the proposed method achieves competitive performance over 2-D, 3-D, and 4-D salient object detection methods. Mianzhao Wang, Fan Shi 0001, Xu Cheng 0003, Meng Zhao 0001, Yao Zhang 0021, Shengyong Chen |
ACM Multimedia | 8 |
| 2022 | Multi-level Temporal Relation Graph for Continuous Sign Language Recognition
Wanli Xue, Leming Guo, Tiantian Yuan, Shengyong Chen |
PRCV (3) | 5 |
| 2022 | Synthetic-to-real: instance segmentation of clinical cluster cells with unlabeled synthetic trainingabstractMOTIVATION: The presence of tumor cell clusters in pleural effusion may be a signal of cancer metastasis. The instance segmentation of single cell from cell clusters plays a pivotal role in cluster cell analysis. However, current cell segmentation methods perform poorly for cluster cells due to the overlapping/touching characters of clusters, multiple instance properties of cells, and the poor generalization ability of the models. RESULTS: In this article, we propose a contour constraint instance segmentation framework (CC framework) for cluster cells based on a cluster cell combination enhancement module. The framework can accurately locate each instance from cluster cells and realize high-precision contour segmentation under a few samples. Specifically, we propose the contour attention constraint module to alleviate over- and under-segmentation among individual cell-instance boundaries. In addition, to evaluate the framework, we construct a pleural effusion cluster cell dataset including 197 high-quality samples. The quantitative results show that the numeric result of APmask is > 90%, a more than 10% increase compared with state-of-the-art semantic segmentation algorithms. From the qualitative results, we can observe that our method rarely has segmentation errors. Meng Zhao 0001, Fan Shi 0001, Xuguo Sun, Shengyong Chen |
Bioinform. | 6 |
| 2022 | Joint Classification and Regression for Visual Tracking with Fully Convolutional Siamese NetworksabstractAbstract Visual tracking of generic objects is one of the fundamental but challenging problems in computer vision. Here, we propose a novel fully convolutional Siamese network to solve visual tracking by directly predicting the target bounding box in an end-to-end manner. We first reformulate the visual tracking task as two subproblems: a classification problem for pixel category prediction and a regression task for object status estimation at this pixel. With this decomposition, we design a simple yet effective Siamese architecture based classification and regression framework, termed SiamCAR, which consists of two subnetworks: a Siamese subnetwork for feature extraction and a classification-regression subnetwork for direct bounding box prediction. Since the proposed framework is both proposal- and anchor-free, SiamCAR can avoid the tedious hyper-parameter tuning of anchors, considerably simplifying the training. To demonstrate that a much simpler tracking framework can achieve superior tracking results, we conduct extensive experiments and comparisons with state-of-the-art trackers on a few challenging benchmarks. Without bells and whistles, SiamCAR achieves leading performance with a real-time speed. Furthermore, the ablation study validates that the proposed framework is effective with various backbone networks, and can benefit from deeper networks. Code is available at https://github.com/ohhhyeahhh/SiamCAR . Dongyan Guo, Yanyan Shao, Zhenhua Wang 0003, Chunhua Shen, Liyan Zhang 0001, Shengyong Chen |
Int. J. Comput. Vis. | 7 |
| 2022 | SMS-Net: Sparse multi-scale voxel feature aggregation network for LiDAR-based 3D object detection
Sheng Liu 0002, Yifeng Cao, Dingda Li, Shengyong Chen |
Neurocomputing | 5 |
| 2022 | A differential evolution with adaptive neighborhood mutation and local search for multi-modal optimization
Mengmeng Sheng, Shengyong Chen, Weibo Liu 0001, Jiafa Mao, Xiaohui Liu 0001 |
Neurocomputing | 2 |
| 2022 | A Driving Assistance System for Solving A-Pillar Blind Spots Through Gaze Detection and Field-of-View EstimationabstractThe blind spots brought by a car’s A-pillar are main reasons of most of accidents. In this work, a vision system that focuses on eliminating A-pillar blind spot of a car without any affection of driver’s operation is investigated. The driver’s facial features are captured by a binocular vision system that is mounted on A-pillar, head poses and gaze line directions are reconstructed. The generated blind spot by A-pillar is then simultaneously calculated according to the position of driver’s gaze. A field-of-view of the blind spots is displayed in a screen system mounted on the A-pillar. The screen conjointly with front and side windows thus provides a full view field for the driver, which can effectively reduce the occurrence of accidents. Minling Yang, Tianjun Li, Shengyong Chen |
Int. J. Pattern Recognit. Artif. Intell. | 5 |
| 2022 | CMRDF: A Real-Time Food Alerting System Based on Multimodal DataabstractA healthy diet is a major concern for everyone, especially for those with specific diseases, such as diabetes. Meanwhile, with the rapid development of new technologies, it is feasible for us to detect the deep latent relationship between daily meals and wellbeing. Advanced Internet of Things devices, such as smart bracelets and wearable cameras, make it possible for people to know how food is related to health at any time. However, it is still arduous for individuals to memorize all the health information and utilize them to regulate their diet. To deal with such problems, we propose a novel system called cross-modal retrieval on diabetogenic food (CMRDF) which realizes a real-time dietary notice based on multimodal data captured from wearable devices. In this system, we propose a new graph-based cross-modal retrieval method named graph correlation analysis with ranking loss that finds the latent information in multimodal data. We use graph convolutional networks to dig the deep latent information in modalities and represent the data in finer granularity. It uses visual and physiological information to estimate whether the food that a user tries to obtain is diabetogenic or not, and feeds back the reasons in detail. Extensive experiments on the MSCOCO data set and the new proposed multimodal diabetogenic food database real-life diabetogenic show that the proposed cross-modal retrieval method outperforms state-of-the-art methods and CMRDF can achieve reliable results on preventing diabetic patients from inappropriate food. Cong Bai, Shengyong Chen |
IEEE Internet Things J. | 4 |
| 2022 | Light field imaging for computer vision: a surveyabstractLight field (LF) imaging has attracted attention because of its ability to solve computer vision problems. In this paper we briefly review the research progress in computer vision in recent years. For most factors that affect computer vision development, the richness and accuracy of visual information acquisition are decisive. LF imaging technology has made great contributions to computer vision because it uses cameras or microlens arrays to record the position and direction information of light rays, acquiring complete three-dimensional (3D) scene information. LF imaging technology improves the accuracy of depth estimation, image segmentation, blending, fusion, and 3D reconstruction. LF has also been innovatively applied to iris and face recognition, identification of materials and fake pedestrians, acquisition of epipolar plane images, shape recovery, and LF microscopy. Here, we further summarize the existing problems and the development trends of LF imaging in computer vision, including the establishment and evaluation of the LF dataset, applications under high dynamic range (HDR) conditions, LF image enhancement, virtual reality, 3D display, and 3D movies, military optical camouflage technology, image recognition at micro-scale, image processing method based on HDR, and the optimal relationship between spatial resolution and four-dimensional (4D) LF information acquisition. LF imaging has achieved great success in various studies. Over the past 25 years, more than 180 publications have reported the capability of LF imaging in solving computer vision problems. We summarize these reports to make it easier for researchers to search the detailed methods for specific solutions. Fan Shi 0001, Meng Zhao 0001, Shengyong Chen |
Frontiers Inf. Technol. Electron. Eng. | 4 |
| 2022 | A particle swarm optimizer with multi-level population sampling and dynamic p-learning mechanisms for large-scale optimization
Mengmeng Sheng, Zidong Wang 0001, Weibo Liu 0001, Shengyong Chen, Xiaohui Liu 0001 |
Knowl. Based Syst. | 5 |
| 2022 | Rainformer: Features Extraction Balanced Network for Radar-Based Precipitation NowcastingabstractPrecipitation nowcasting is one of the fundamental challenges in natural hazard research. High-intensity rainfall, especially the rainstorm, will lead to the enormous loss of people’s property. Existing methods usually utilize convolution operation to extract rainfall features and increase the network depth to expand the receptive field to obtain fake global features. Although this scheme is simple, only local rainfall features can be extracted leading to insensitivity to high-intensity rainfall. This letter proposes a novel precipitation nowcasting framework named Rainformer, in which, two practical components are proposed: the global features extraction unit and the gate fusion unit (GFU). The former provides robust global features learning ability depending on the window-based multi-head self-attention (W-MSA) mechanism, while the latter provides a balanced fusion of local and global features. Rainformer has a simple yet efficient architecture and significantly improves the accuracy of rainfall prediction, especially on high-intensity rainfall. It offers a potential solution for real-world applications. The experimental results show that Rainformer outperforms seven state of the arts methods on the benchmark database and provides more insights into the high-intensity rainfall prediction task. Cong Bai, Jinglin Zhang 0001, Shengyong Chen |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2022 | Robust Visual-Lidar Simultaneous Localization and Mapping System for UAVabstractObtaining 3-D data by LIDAR from unmanned aerial vehicles (UAVs) is vital for the field of remote sensing; however, the highly dynamic movement of UAVs and narrow viewpoint of LIDAR pose a great challenge to the self-localization for UAVs based on solely LIDAR sensor. To this end, we propose a robust simultaneous localization and mapping (SLAM) system, which combines the image data obtained by vision sensor and point clouds obtained by LIDAR. In the front-end of the proposed system, the more stable line and plane features are extracted from point clouds through clustering. Then the relative pose between two consecutive frames is computed by the least squares iterative closest point algorithm. Afterward, a novel direct odometry algorithm is developed by combining the image frames and sparse point clouds, where the relative pose is used as a prior. In the back-end, the pose estimation is refined and the 3-D map with texture information is built at a lower frequency. Extensive experiments show that our method can achieve robust and highly precise localization and mapping for UAVs. Kaiqi Chen 0001, Qinying Chen, Yanhong Yang, Jianhua Zhang 0002, Shengyong Chen |
IEEE Geosci. Remote. Sens. Lett. | 6 |
| 2022 | Hyperspectral Image Restoration via Subspace-Based Nonlocal Low-Rank Tensor ApproximationabstractIn this letter, we present a subspace-based nonlocal low-rank tensor approximation framework (SNLRTA) for hyperspectral image (HSI) restoration. The proposed method consists of a subspace learning method to achieve an accurate subspace characterization of HSI and a nonlocal low-rank tensor approximation to take spatial nonlocal self-similarity into consideration. Specifically, the HSI first exploits residual statistics on median filtered image to estimate a robust subspace. Laplacian scale mixture (LSM) modeling is then investigated to model tensor coefficients from overlapping cubes in low-rank subspace. Both the hidden scale parameters and the sparse coefficients therein are adaptively shrink, characterizing the sparsity of similar patches. Meanwhile, the$\ell _{1}$data fidelity facilitates the implicit detection of outliers after median filtering. Substantiated by extensive experimental results, the proposed method outperforms several state-of-the-art approaches on mixed noise removal, qualitatively and quantitatively. Yanhong Yang, Yuan Feng 0002, Jianhua Zhang 0002, Shengyong Chen |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2022 | SSA-Net: Spatial self-attention network for COVID-19 pneumonia infection segmentation with semi-supervised few-shot learning
Xiaoyan Wang 0007, Yiwen Yuan, Dongyan Guo, Ming Xia 0005, Zhenhua Wang 0003, Cong Bai, Shengyong Chen |
Medical Image Anal. | 9 |
| 2022 | Unsupervised adversarial image retrieval
Ling Huang 0003, Cong Bai, Yijuan Lu, Shaobo Zhang 0005, Shengyong Chen |
Multim. Syst. | 5 |
| 2022 | Res2-UNeXt: a novel deep learning framework for few-shot cell image segmentation
Sixian Chan 0001, Cong Bai, Shengyong Chen |
Multim. Tools Appl. | 5 |
| 2022 | Semantic association enhancement transformer with relative position for image captioning
Yunbo Wang, Yuxin Peng 0001, Shengyong Chen |
Multim. Tools Appl. | 4 |
| 2022 | Dual attention granularity network for vehicle re-identification
Jianhua Zhang 0002, Jingbo Chen, Jiewei Cao, Ruyu Liu, Linjie Bian, Shengyong Chen |
Neural Comput. Appl. | 6 |
| 2022 | Online multiple object tracking using joint detection and embedding network
Sixian Chan 0001, Yangwei Jia, Xiaolong Zhou 0001, Cong Bai, Shengyong Chen, Xiaoqin Zhang 0002 |
Pattern Recognit. | 5 |
| 2022 | DARTSRepair: Core-failure-set guided DARTS for network robustness to common corruptions
Xuhong Ren, Jianlang Chen, Felix Juefei-Xu, Wanli Xue, Qing Guo 0005, Lei Ma 0003, Jianjun Zhao 0001, Shengyong Chen |
Pattern Recognit. | 8 |
| 2022 | Multi-Level View Associative Convolution Network for View-Based 3D Model RetrievalabstractWith the continuous improvement of image processing capabilities, a three-dimensional (3D) model that can contain rich information is becoming the fourth type of multimedia data (in addition to sound, image, and video). Moreover, since there is a wide range of applications of 3D models, how to quickly and effectively obtain the correct target model from the massive data has become a key issue. To date, 3D model retrieval approaches have been proposed, and in these approaches, view-based 3D model retrieval methods can achieve satisfactory performance. In the 3D model retrieval task, the latent relationship mining of all images in a 3D model, the adaptive fusion of different images, and the discriminative feature extraction are the main challenges, but in most existing solutions, these issues are separately performed and they are not explored in an end-to-end network architecture. To solve these issues, in this work, we propose a novel and effective multi-level view associative convolution network (MLVACN) to realize view-based 3D model retrieval, where the relationship exploration of multiple-view images, the fusion of different images, and the feature discrimination learning are realized in a unified end-to-end framework. Specifically, we design the group association layer and the block association layer to study the latent relationships among different views from the view-level and the block-level, respectively. Moreover, the weight fusion layer is further designed to adaptively fuse different views in a 3D model. In addition, these three layers are embedded into theMLVACN. Finally, the pairwise discrimination loss function is proposed to learn the discriminative features of the 3D model. Extensive experimental results on three 3D model retrieval datasets including ModelNet40, ModelNet10, and ShapeNetCore55 demonstrate thatMLVACNcan outperform state-of-the-art methods in term of mAP. When the ModelNet40 dataset is used, the mAP ofMLVACNis improved by 13.25%, 7.75%, 3.95%, and 0.61% as compared to those of the MVCNN, GVCNN, PVNet, and MLVCNN methods, respectively. Zan Gao 0001, Yan Zhang 0154, Hua Zhang 0003, Weili Guan, Shengyong Chen |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2022 | Fast Tensor Nuclear Norm for Structured Low-Rank Visual InpaintingabstractLow-rank modeling has achieved great success in visual data completion. However, the low-rank assumption of original visual data may be in approximate mode, which leads to suboptimality for the recovery of underlying details, especially when the missing rate is extremely high. In this paper, we go further by providing a detailed analysis about the rank distributions in Hankel structured and clustered cases, and figure out both non-local similarity and patch-based structuralization play a positive role. This motivates us to develop a new Hankel low-rank tensor recovery method that is competent to truthfully capture the underlying details with sacrifice of slightly more computational burden. First, benefiting from the correlation of different spectral bands and the smoothness of local spatial neighborhood, we divide the visual data into overlapping 3D patches and group the similar ones into individual clusters exploring the non-local similarity. Second, the 3D patches are individually mapped to the structured Hankel tensors for better revealing low-rank property of the image. Finally, we solve the tensor completion model via the well-known alternating direction method of multiplier (ADMM) optimization algorithm. Due to the fact that size expansion happens inevitably in Hankelization operation, we further propose a fast randomized skinny tensor singular value decomposition (rst-SVD) to accelerate the per-iteration running efficiency. Extensive experimental results on real world datasets verify the superiority of our method compared to the state-of-the-art visual inpainting approaches. Honghui Xu 0002, Jianwei Zheng 0001, Xiaomin Yao, Yuchao Feng, Shengyong Chen |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2022 | Stable Linear Structures and Seam Measurements for Parallax Image StitchingabstractParallax tolerance is a fundamental problem in image stitching. To solve this problem, increasing effort has been devoted to the spatially varying and seam cutting. However, there still exist some issues that need to be adequately addressed. First, the implementation of the spatially varying warping requires a restricted premise that the overlapping region can be aligned: if this premise is not met, ghosting caused by misalignment will emerge. In addition, the spatially varying warping may cause distortion due to the issue of inconsistent homographies. Second, conventional seam cutting will lead to objects being cropped and duplicated. Therefore, in this paper, we propose a stable framework for stitching images: the framework consists of a uniform linear structure model that is able to mitigate the distortion of projection and perspective, while preserving the structures of objects in the non-overlapping region; and a stable hybrid actor-critic that estimates stable seam measurements in the overlapping region to diminish the parallax. Comparison experiments show that the proposed method is superior to some conventional methods with respect to mitigating ghosting and preserving structure. Wanli Xue, Weilun Xie, Yao Zhang 0021, Shengyong Chen |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2022 | Object-Aware Ghost Identification and Elimination for Dynamic Scene MosaicabstractComposite ghost is a common phenomenon that widely exists in dynamic scene image mosaic and significantly affects the naturalness of mosaic. To remove the ghost effectively and produce visually natural mosaic, we propose a novel image mosaic method by jointly identifying composite ghost and eliminating ghost regions without distorting, splitting, and duplicating objects. Specifically, our main contributions are three-fold:First, we propose themotion-awarecomposite ghost identification to localize the potential composite ghosts in the mosaic region (i.e., overlapping area between two images to be stitched) by detecting the salient-moving objects in two stitched images.Second, we design theobject-awarealternative region selection strategy to produce ghostless regions that can replace the localized composite ghosts while avoiding object distortion, object separation, and object repetition.Third, we realize theimage interpolation-basedcomposite ghost elimination that can generate natural stitched image by eliminating the composite ghost of the initial blending result with the selected image source. We validate the proposed method on challenging datasets and show that our method outperform the state-of-the-art methods. Zhe Zhang 0039, Xuhong Ren, Wanli Xue, Chengwei Zhang 0001, Qing Guo 0005, Shengyong Chen |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2022 | Multi-Task Convolution Operators With Object Detection for Visual TrackingabstractRecently, multi-task correlation filters has drawn much attention in the object tracking field, which utilizes the multi-task learning (MTL) approach to explore the interdependencies among deep features for object tracking. However, the existing multi-task correlation filters based method fails to consider the relations between the correlation filters. To address this problem, a novel correlation filters based visual tracking method is proposed in this paper, with the integration of multi-task convolution operators and object detection. In our method, convolution and correction filters are jointly learnt through using the MTL technique, with the purpose of exploring not only the interdependencies of deep features but also the internal relevance of the convolution filters. In addition, object detection is introduced into our algorithm to handle the problem of object missing to ensure a better performance of our tracking method. Experiments on five benchmark datasets demonstrate that the proposed visual tracking method outperforms existing state-of-the-art approaches. Yuhui Zheng, Xinyan Liu 0002, Bin Xiao 0002, Xu Cheng 0003, Yi Wu 0001, Shengyong Chen |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2022 | LSCIDMR: Large-Scale Satellite Cloud Image Database for Meteorological ResearchabstractPeople can infer the weather from clouds. Various weather phenomena are linked inextricably to clouds, which can be observed by meteorological satellites. Thus, cloud images obtained by meteorological satellites can be used to identify different weather phenomena to provide meteorological status and future projections. How to classify and recognize cloud images automatically, especially with deep learning, is an interesting topic. Generally speaking, large-scale training data are essential for deep learning. However, there is no such cloud images database to date. Thus, we propose a large-scale cloud image database for meteorological research (LSCIDMR). To the best of our knowledge, it is the first publicly available satellite cloud image benchmark database for meteorological research, in which weather systems are linked directly with the cloud images. LSCIDMR contains 104 390 high-resolution images, covering 11 classes with two different annotation methods: 1) single-label annotation and 2) multiple-label annotation, called LSCIDMR-S and LSCIDMR-M, respectively. The labels are annotated manually, and we obtain a total of 414 221 multiple labels and 40 625 single labels. Several representative deep learning methods are evaluated on the proposed LSCIDMR, and the results can serve as useful baselines for future research. Furthermore, experimental results demonstrate that it is possible to learn effective deep learning models from a sufficiently large image database for the cloud image classification. Cong Bai, Minjing Zhang, Jinglin Zhang 0001, Jianwei Zheng 0001, Shengyong Chen |
IEEE Trans. Cybern. | 5 |
| 2022 | A Novel Multiple-View Adversarial Learning Network for Unsupervised Domain Adaptation Action RecognitionabstractAbstract-domain adaptation action recognition is a hot research topic in machine learning and some effective approaches have been proposed. However, samples in the target domain with label information are often required by these approaches. Moreover, domain-invariant discriminative feature learning, feature fusion, and classifier module learning have not been explored in an end-to-end framework. Thus, in this study, we propose a novel end-to-end multiple-view adversarial learning network (MAN) for unsupervised domain adaptation action recognition in which the fusion of RGB and optical-flow features, domain-invariant discrimination feature learning, and action recognition is conducted in a unified framework. Specifically, a robust spatiotemporal feature extraction network, including a spatial transform network and an adaptive intrachannel weight network, is proposed to improve the scale invariance and robustness of the method. Then, a self-attention mechanism fusion module is designed to adaptively fuse the RGB and optical-flow features. Moreover, a multiview adversarial learning loss is developed to obtain domain-invariant discriminative features. In addition, three benchmark datasets are constructed for unsupervised domain adaptation action recognition, for which all actions and samples are carefully collected from public action datasets, and their action categories are hierarchically augmented, which can guide how to extend existing action datasets. We conduct extensive experiments on four benchmark datasets, and the experimental results demonstrate that our proposed MAN can outperform several state-of-the-art unsupervised domain adaptation action recognition approaches. When the SDAI Action II-6 and SDAI Action II-11 datasets are used, MAN can achieve 3.7% ( H → U ) and 6.1% ( H → U ) improvements over the temporal attentive adversarial adaptation network (published in ICCV 2019) module, respectively. As an added contribution, the SDAI Action II-6, SDAI Action II-11, and SDAI Action II-16 datasets will be released to facilitate future research on domain adaptation action recognition. Zan Gao 0001, Yibo Zhao 0001, Hua Zhang 0003, Da Chen 0002, Anan Liu, Shengyong Chen |
IEEE Trans. Cybern. | 6 |
| 2022 | A Differential Evolution Algorithm With Adaptive Niching and K-Means Operation for Data ClusteringabstractClustering, as an important part of data mining, is inherently a challenging problem. This article proposes a differential evolution algorithm with adaptive niching and k -means operation (denoted as DE_ANS_AKO) for partitional data clustering. Within the proposed algorithm, an adaptive niching scheme, which can dynamically adjust the size of each niche in the population, is devised and integrated to prevent premature convergence of evolutionary search, thus appropriately searching the space to identify the optimal or near-optimal solution. Furthermore, to improve the search efficiency, an adaptive k -means operation has been designed and employed at the niche level of population. The performance of the proposed algorithm has been evaluated on synthetic as well as real datasets and compared with related methods. The experimental results reveal that the proposed algorithm is able to reliably and efficiently deliver high quality clustering solutions and generally outperforms related methods implemented for comparisons. Weiguo Sheng 0001, Zidong Wang 0001, Qi Li 0021, Yujun Zheng 0001, Shengyong Chen |
IEEE Trans. Cybern. | 6 |
| 2022 | A Blockchain-Empowered Cluster-Based Federated Learning Model for Blade Icing Estimation on IoT-Enabled Wind TurbineabstractWind energy is a fast-growing renewable energy but faces blade icing. Data-driven methods provide talented solutions for blade icing detection, but a considerable amount of Internet of Things data needs to be collected to a central server, which may lead to the leakage of sensitive business data. To address this limitation, this article proposesBLADE, a Blockchain-empowered imbalanced federated learning (FL) model for blade icing detection. With the help of the Blockchain, the conventional FL is improved without worrying about the failure of the single centralized server and boosts the privacy preserving. A validation mechanism is introduced into the Blockchain to enhance the defense against poisoning attacks. In addition, a novel imbalanced learning algorithm is integrated into BLADE to solve the class imbalance problem in the sensor data. BLADE is evaluated on ten wind turbines from two wind farms. The experimental results verify the effectiveness, superiority, and feasibility of the proposed BLADE. Xu Cheng 0003, Fan Shi 0001, Meng Zhao 0001, Shengyong Chen, Hao Wang 0003 |
IEEE Trans. Ind. Informatics | 5 |
| 2022 | Variational Abnormal Behavior Detection With Motion ConsistencyabstractAbnormal crowd behavior detection has recently attracted increasing attention due to its wide applications in computer vision research areas. However, it is still an extremely challenging task due to the great variability of abnormal behavior coupled with huge ambiguity and uncertainty of video contents. To tackle these challenges, we propose a new probabilistic framework named variational abnormal behavior detection (VABD), which can detect abnormal crowd behavior in video sequences. We make three major contributions: (1) We develop a new probabilistic latent variable model that combines the strengths of the U-Net and conditional variational auto-encoder, which also are the backbone of our model; (2) We propose a motion loss based on an optical flow network to impose the motion consistency of generated video frames and input video frames; (3) We embed a Wasserstein generative adversarial network at the end of the backbone network to enhance the framework performance. VABD can accurately discriminate abnormal video frames from video sequences. Experimental results on UCSD, CUHK Avenue, IITB-Corridor, and ShanghaiTech datasets show that VABD outperforms the state-of-the-art algorithms on abnormal crowd behavior detection. Without data augmentation, our VABD achieves 72.24% in terms of AUC on IITB-Corridor, which surpasses the state-of-the-art methods by nearly 5%. Jing Li 0027, Qingwang Huang, Yingjun Du, Xiantong Zhen, Shengyong Chen, Ling Shao 0001 |
IEEE Trans. Image Process. | 5 |
| 2022 | Dual-View 3D Reconstruction via Learning Correspondence and Dependency of Point Cloud RegionsabstractMulti-view 3D reconstruction generally adopts the feature fusion strategy to guide the generation of 3D shape for objects with different views. Empirically, the correspondence learning of object regions across different views enables better feature fusion. However, such idea has not been fully exploited in existing methods. Furthermore, current methods fail to explore the intrinsic dependency among regions within a 3D shape, leading to a rough reconstruction result. To address the above issues, we propose a Dual-View 3D Point Cloud reconstruction architecture named DVPC, which takes two views images as inputs, and progressively generates a refined 3D point cloud. First, a point cloud generation network is assigned to generate a coarse point cloud for each input view. Second, a dual-view point clouds synthesis network is presented in DVPC. It constructs a regional attention mechanism to learn a high-quality correspondence among regions across two coarse point clouds in different views, so that our DVPC can achieve feature fusion accurately. And then it develops a point cloud deformation module to produce a relatively-precise point cloud via establishing the communication between the coarse point cloud and the fused feature. Lastly, a point-region transformer network is devised to model the dependency among regions within the relatively-precise point cloud. With the dependency, the relatively-precise point cloud is refined into a desirable 3D point cloud with rich details. Qualitative and quantitative experiments on the ShapeNet and Pix3D datasets demonstrate that the proposed DVPC outperforms the state-of-the-art methods in terms of reconstruction quality. Jianxin Wang 0001, Shourui Yang, Yunbo Wang, Jianhua Zhang 0002, Yuxin Peng 0001, Shengyong Chen |
IEEE Trans. Image Process. | 6 |
| 2022 | A Temporal-Aware Relation and Attention Network for Temporal Action LocalizationabstractTemporal action localization is currently an active research topic in computer vision and machine learning due to its usage in smart surveillance. It is a challenging problem since the categories of the actions must be classified in untrimmed videos and the start and end of the actions need to be accurately found. Although many temporal action localization methods have been proposed, they require substantial amounts of computational resources for the training and inference processes. To solve these issues, in this work, a novel temporal-aware relation and attention network (abbreviated as TRA) is proposed for the temporal action localization task. TRA has an anchor-free and end-to-end architecture that fully uses temporal-aware information. Specifically, a temporal self-attention module is first designed to determine the relationship between different temporal positions, and more weight is given to features within the actions. Then, a multiple temporal aggregation module is constructed to aggregate the temporal domain information. Finally, a graph relation module is designed to obtain the aggregated graph features, which are used to refine the boundaries and classification results. Most importantly, these three modules are jointly explored in a unified framework, and temporal awareness is always fully used. Extensive experiments demonstrate that the proposed method can outperform all state-of-the-art methods on the THUMOS14 dataset with an average mAP that reaches 67.6% and obtain a comparable result on the ActivityNet1.3 dataset with an average mAP that reaches 34.4%. Compared with A2Net (TIP20), PCG-TAL (TIP21), and AFSD (CVPR21) TRA can achieve improvements of 11.7%, 4.4%, and 1.8%, respectively on the THUMOS14 dataset. Yibo Zhao 0001, Hua Zhang 0003, Zan Gao 0001, Weili Guan, Jie Nie, Anan Liu, Meng Wang 0001, Shengyong Chen |
IEEE Trans. Image Process. | 8 |
| 2022 | Data-Driven Modeling for Transferable Sea State Estimation Between Marine SystemsabstractSea state estimation is beneficial for marine systems to enhance on-board decision-making and improve work efficiency. In the era of ship intelligence, artificial intelligence has greatly promoted the technology of sensing environment, such as by using the deep learning. However, it is difficult to collect enough motion data from a marine system to train a deep learning model. In addition, the model for sea state estimation is trained using the data from a specific marine system; applying the model directly to another marine system may result in performance degradation. In this paper, a supervised transfer learning based framework for sea state estimation (STLSSE) is proposed. The STLSSE focuses on knowledge transfer when the collected data for the source marine system is sufficient but the collected data of the target marine system is scarce. In STLSSE, a data pairing algorithm is proposed to determine the relationship of the source and the target marine system. Based on these paired data, a Siamese convolutional neural network, including a new proposed residual fully convolutional network and two novel attention modules, is designed for the semantic alignment. Moreover, the conventional contrastive loss is improved to characterize the distributions when there are only few samples in the target marine system. The extensive comparisons between STLSSE and state-of-the-art transfer learning approaches show its superior performance. The comparisons with state-of-the-art attention modules has verified the competitiveness of the proposed attention modules. The key parameters and each component of STLSSE are emphasized in the ablation and sensitivity studies. Xu Cheng 0003, Guoyuan Li, Peihua Han, Robert Skulstad, Shengyong Chen, Houxiang Zhang |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2022 | Guest Editorial Introduction to the Special Issue on Intelligent Transportation Systems in Epidemic AreasabstractThe COVID-19 pandemic has posed significant challenges to transportation systems in various aspects, such as transferring patients and medical resources, enforcing physical distancing in public transportation, and controlling virus transmission through transportation networks. To address these challenges, a variety of artificial intelligence technologies, such as autonomous driving, big data analytics, intelligent vehicle routing and scheduling, and intelligent traffic control, have been employed in the design of intelligent transportation systems. This Special Issue provides a forum for researchers and practitioners to present the most recent advances in presenting and applying intelligent technologies to promote transportation systems in large-scale epidemics. Yujun Zheng 0001, Honghai Liu 0001, Houxiang Zhang, Shengyong Chen |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2022 | A Novel Deep Class-Imbalanced Semisupervised Model for Wind Turbine Blade Icing DetectionabstractWind energy is of great importance for future energy development. In order to fully exploit wind energy, wind farms are often located at high latitudes, a practice that is accompanied by a high risk of icing. Traditional blade icing detection methods are usually based on manual inspection or external sensors/tools, but these techniques are limited by human expertise and additional costs. Model-based methods are highly dependent on prior domain knowledge and prone to misinterpretation. Data-driven approaches can offer promising solutions but require a massive amount of labeled training data, which are not generally available. In addition, the data collected for icing detection tend to be imbalanced because, most of the time, wind turbines operate under normal conditions. To address these challenges, this article presents a novel deep class-imbalanced semisupervised (DCISS) model for estimating blade icing conditions. DCISS integrates class-imbalanced and semisupervised learning (SSL) using a prototypical network that can rebalance features and measure the similarities between labeled and unlabeled samples. In addition, a channel calibration attention module is proposed to improve the ability to extract features from raw data. The proposed model has been evaluated using the blade icing datasets of three wind turbines. Compared to the classical anomaly detection and state-of-the-art SSL algorithms, DCISS shows significant advantages in terms of accuracy. Compared to five different class-imbalanced loss functions, the proposed DCISS is competitive. The generalization and practicability of the proposed model are further verified in the use case of online estimation. Xu Cheng 0003, Fan Shi 0001, Xiufeng Liu 0001, Meng Zhao 0001, Shengyong Chen |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2022 | Pairwise Two-Stream ConvNets for Cross-Domain Action Recognition With Small DataabstractIn this work, we target cross-domain action recognition (CDAR) in the video domain and propose a novel end-to-end pairwise two-stream ConvNets (PTC) algorithm for real-life conditions, in which only a few labeled samples are available. To cope with the limited training sample problem, we employ pairwise network architecture that can leverage training samples from a source domain and, thus, requires only a few labeled samples per category from the target domain. In particular, a frame self-attention mechanism and an adaptive weight scheme are embedded into the PTC network to adaptively combine the RGB and flow features. This design can effectively learn domain-invariant features for both the source and target domains. In addition, we propose a sphere boundary sample-selecting scheme that selects the training samples at the boundary of a class (in the feature space) to train the PTC model. In this way, a well-enhanced generalization capability can be achieved. To validate the effectiveness of our PTC model, we construct two CDAR data sets (SDAI Action I and SDAI Action II) that include indoor and outdoor environments; all actions and samples in these data sets were carefully collected from public action data sets. To the best of our knowledge, these are the first data sets specifically designed for the CDAR task. Extensive experiments were conducted on these two data sets. The results show that PTC outperforms state-of-the-art video action recognition methods in terms of both accuracy and training efficiency. It is noteworthy that when only two labeled training samples per category are used in the SDAI Action I data set, PTC achieves 21.9% and 6.8% improvement in accuracy over two-stream and temporal segment networks models, respectively. As an added contribution, the SDAI Action I and SDAI Action II data sets will be released to facilitate future research on the CDAR task. Zan Gao 0001, Leming Guo, Tongwei Ren, Anan Liu, Zhiyong Cheng 0001, Shengyong Chen |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2022 | Temporal Encoding and Multispike Learning Framework for Efficient Recognition of Visual PatternsabstractBiological systems under a parallel and spike-based computation endow individuals with abilities to have prompt and reliable responses to different stimuli. Spiking neural networks (SNNs) have thus been developed to emulate their efficiency and to explore principles of spike-based processing. However, the design of a biologically plausible and efficient SNN for image classification still remains as a challenging task. Previous efforts can be generally clustered into two major categories in terms of coding schemes being employed: rate and temporal. The rate-based schemes suffer inefficiency, whereas the temporal-based ones typically end with a relatively poor performance in accuracy. It is intriguing and important to develop an SNN with both efficiency and efficacy being considered. In this article, we focus on the temporal-based approaches in a way to advance their accuracy performance by a great margin while keeping the efficiency on the other hand. A new temporal-based framework integrated with the multispike learning is developed for efficient recognition of visual patterns. Different approaches of encoding and learning under our framework are evaluated with the MNIST and Fashion-MNIST data sets. Experimental results demonstrate the efficient and effective performance of our temporal-based approaches across a variety of conditions, improving accuracies to higher levels that are even comparable to rate-based ones but importantly with a lighter network structure and far less number of spikes. This article attempts to extend the advanced multispike learning to the challenging task of image recognition and bring state of the arts in temporal-based approaches to a novel level. The experimental results could be potentially favorable to low-power and high-speed requirements in the field of artificial intelligence and contribute to attract more efforts toward brain-like computing. Qiang Yu 0005, Shiming Song 0001, Chenxiang Ma, Jianguo Wei, Shengyong Chen, Kay Chen Tan |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2021 | RCGA-Net: An Improved Multi-hybrid Attention Mechanism Network in Biomedical Image SegmentationabstractDrawing support from an effective Medical Image Segmentation (MIS) is conducive to a substantial diagnostic basis for the physicians to identify the focus lesion in the patient body and give the subsequent clinical assessment of the patient status. Although various works have tried the challenging quantitative analysis problem, it is still difficult to conduct precise automatic segmentation, especially the soft tissue organs. In this decade, with the increased amount of available datasets, deep learning-based networks have achieved remarkable performance in image processing. Inspired by the state-of-the-art deep learning works, in this paper, we propose an end-to-end multi-layer network named RCGA-Net. It consists of an encoder-decoder backbone that integrates a coordinate attention mechanism based on space and channel and a global context extraction module to highlight more valuable information. To evaluate the performance of RCGA-Net, we apply it to different kinds of clinical and experimental MIS tasks to testify its generalization ability. Extensive experiments represent that our schema has taken the outperform or compatible results among the comparison methods group. Specifically, the numeric result of RCGA-Net on the pulmonary dataset has achieved a 99.12% optimum F1-score. Feng Xiao 0005, Shengyong Chen, Zhijun Liao, Jijun Tang |
BIBM | 5 |
| 2021 | Consistency-Aware Graph Network for Human Interaction UnderstandingabstractCompared with the progress made on human activity classification, much less success has been achieved on human interaction understanding (HIU). Apart from the latter task is much more challenging, the main cause is that recent approaches learn human interactive relations via shallow graphical models, which is inadequate to model complicated human interactions. In this paper, we propose a consistency-aware graph network, which combines the representative ability of graph network and the consistency-aware reasoning to facilitate HIU. Our network consists of three components, a backbone CNN to extract image features, a factor graph network to learn third-order interactive relations among participants, and a consistency-aware reasoning module to enforce labeling and grouping consistencies. Our key observation is that the consistency-aware-reasoning bias for HIU can be embedded into an energy, minimizing which delivers consistent predictions. An efficient mean-field inference algorithm is proposed, such that all modules of our network could be trained jointly in an end-to-end manner. Experimental results show that our approach achieves leading performance on three benchmarks. Code is available at https://git.io/CAGNet. Zhenhua Wang 0003, Jiajun Meng, Dongyan Guo, Jianhua Zhang 0002, Qinfeng Shi, Shengyong Chen |
ICCV | 6 |
| 2021 | A Novel Patch Convolutional Neural Network for View-based 3D Model RetrievalabstractIn industrial enterprises, effective retrieval of three-dimensional (3-D) computer-aided design (CAD) models can greatly save time and cost in new product development and manufacturing, thus, many researchers have focused on it. Recently, many view-based 3D model retrieval methods have been proposed and have achieved state-of-the-art performance. However, most of these methods focus on extracting more discriminative view-level features and effectively aggregating the multi-view images of a 3D model, and the latent relationship among these multi-view images is not fully explored. Thus, we tackle this problem from the perspective of exploiting the relationships between patch features to capture long-range associations among multi-view images. To capture associations among views, in this work, we propose a novel patch convolutional neural network (PCNN ) for view-based 3D model retrieval. Specifically, we first employ a CNN to extract patch features of each view image separately. Second, a novel neural network module named PatchConv is designed to exploit intrinsic relationships between neighboring patches in the feature space to capture long-range associations among multi-view images. Then, an adaptive weighted view layer is further embedded into PCNN to automatically assign a weight to each view according to the similarity between each view feature and the view-pooling feature. Finally, a discrimination loss function is employed to extract the discriminative 3D model feature, which consists of softmax loss values generated by the fusion classifier and the specific classifier. Extensive experimental results on two public 3D model retrieval benchmarks, namely, the ModelNet40, and ModelNet10, demonstrate that our proposed PCNN can outperform state-of-the-art approaches, with mAP values of 93.67%, and 96.23%, respectively. Zan Gao 0001, Yuxiang Shao, Weili Guan, Meng Liu 0006, Zhiyong Cheng 0001, Shengyong Chen |
ACM Multimedia | 6 |
| 2021 | MRAC-Net: Multi-resolution Anisotropic Convolutional Network for 3D Point Cloud Completion
Sheng Liu 0002, Dingda Li, Yifeng Cao, Shengyong Chen |
PRICAI (3) | 5 |
| 2021 | Box Regression-Guided Anchor-free for Robust Visual TrackingabstractThe Siamese tracker-based approach has achieved significant success in recent years. However, these approaches do not consider the different requirements for input feature in classification and regression branches. The regression branch needs feature information slightly larger than the object region, while the classification branch needs to avoid classification failure caused by the introduction of background information. In this paper, we present a novel Box Regression-Guided Anchor-free for Robust Visual Tracking. Firstly, a scale-aware regression module is designed to satisfy the feature requirements of the regression branch, which can capture feature information of various scales. Secondly, regression-guided classification module is applied to aligning the feature between the regression result and correlation feature, thereby avoiding the introduction of background information to classification branch. In addition, the new correlation operation is introduced to gain more superb correlation feature. Comparsion experimental exhibits that the proposed tracker achieves promising results in five challenging benchmark tests, including GOT-10K, OTB-2015, VOT-2018, VOT-2019 and TrackingNet, and run at an average speed of 60 FPS in real-time. Sixian Chan 0001, Xiaolong Zhou 0001, Cong Bai, Hua Gao, Shengyong Chen |
SMC | 6 |
| 2021 | CRB-Net: A Sign Language Recognition Deep Learning Strategy Based on Multi-modal Fusion with Attention MechanismabstractAt present, sign language recognition (SLR) researchers are mainly committed to establishing a sign language recognition model based on single-mode data. Nevertheless, this manipulation often leads to a defective understanding of the sign language semantics and ignoring some visual information. In a nutshell, the challenges locate redundancy removing and the alignment of the sign language data with the given tag. To solve the conundrum, this paper proposes a deep learning strategy called CRB-Net, which has used a kind of multimodal fusion attention mechanism. We first extract the features from RGB video and depth video, respectively, then conduct multi-modal fusion. Finally, the fused feature information is fed into an encoder-decoder network to achieve the goal of end-to-end continuous SLR. We verify the effectiveness of our method on three datasets, including the German dataset RWTH-Phoenix-Weather-2014, the Chinese dataset USTC-CSL and the Chinese dataset TJUT-SLRT. As shown by experimental results, the accuracy of 98.5% of our framework CRB-Net has outperformed the state-of-the-art works in the comparison, both in accuracy and algorithm execution efficiency. Feng Xiao 0005, Tiantian Yuan, Shengyong Chen |
SMC | 4 |
| 2021 | Tensor completion using patch-wise high order Hankelization and randomized tensor ring initialization
Jianwei Zheng 0001, Honghui Xu 0002, Yuchao Feng, Peijun Chen, Shengyong Chen |
Eng. Appl. Artif. Intell. | 6 |
| 2021 | End-to-end feature fusion Siamese network for adaptive visual trackingabstractAbstract According to observations, different visual objects have different salient features in different scenarios. Even for the same object, its salient shape and appearance features may change greatly from time to time in a long‐term tracking task. Motivated by them, an end‐to‐end feature fusion framework was proposed based on the Siamese network, named FF‐Siam, which can effectively fuse different features for adaptive visual tracking. The framework consists of four layers. A feature extraction layer is designed to extract the different features of the target region and search region. The extracted features are then put into a weight generation layer to obtain the channel weights, which indicate the importance of different feature channels. Both features and the channel weights are utilised in a template generation layer to generate a discriminative template. Finally, the corresponding response maps created by the convolution of the search region features and the template are applied with a fusion layer to obtain the final response map for locating the target. Experimental results demonstrate that the proposed framework achieves state‐of‐the‐art performance on the popular Temple‐Colour, OTB50 and UAV123 benchmarks. Dongyan Guo, Weixuan Zhao, Zhenhua Wang 0003, Shengyong Chen |
IET Image Process. | 6 |
| 2021 | Hybrid-attention guided network with multiple resolution features for person re-identification
Guoqing Zhang 0002, Junchuan Yang, Yuhui Zheng, Yi Wu 0001, Shengyong Chen |
Inf. Sci. | 6 |
| 2021 | Enhanced low-rank constraint for temporal subspace clustering and its acceleration scheme
Jianwei Zheng 0001, Guojiang Shen, Shengyong Chen |
Pattern Recognit. | 4 |
| 2021 | Detection and Segmentation of Unlearned Objects in Unknown EnvironmentabstractDetecting and segmenting unlearned objects in unknown environment is a very important visual perception ability to enhance industrial intelligence. In this article, we present a novel conditional random field model integrating unimodal and cross-modal terms for detecting and segmenting object instances without knowing their categories and without sampling extra proposals. This model takes a paired image and point cloud as input, from which we first develop a set of novel category-independent features to distinguish objects. Then, a set of unary, pairwise, and higher order potentials are designed according to these category-independent features, and the cross-modal potential is introduced as a novel global constraints to keep the spatial consistency in both 2-D and 3-D modalities. In this novel model, the unlearned object detection and segmentation is treated as the process of pixel labeling. Thus, adjacent or occlusion object instances can also be separated efficiently from a labeled map. By comparison with the baseline methods, experimental results on a public RGB+D dataset show that the proposed model can obtain better performance with improved precision and recall rate. Moreover, we use the proposed method in a real industrial scene and achieve satisfactory performance. Jianhua Zhang 0002, Jingbo Chen, Shengyong Chen, Zhenhua Wang 0003, Jianwei Zhang 0001 |
IEEE Trans. Ind. Informatics | 3 |
| 2021 | Map Recovery and Fusion for Collaborative Augment Reality of Multiple Mobile DevicesabstractThe map recovery and fusion is a key issue in the application of large scale and long-term augmented reality (AR) scenarios. However, they are still not addressed well in an efficient and precise way, especially for complex industrial environments. In this article, we propose a map recovery and fusion strategy based on vision-inertial simultaneous localization and mapping. We first develop a heuristic strategy that can fast search and match map points among multiple maps, and can be used for efficient map fusion. For map recovery, we leverage the inertial sensors for short time motion estimation, and transform the previous lost map to the current map. Based on this strategy, a novel framework for collaborative AR is implemented and can parallelly run in multiple mobile devices in real time. Extensive experiments have been carried out on a public data set, and the results show that the proposed method can recovery and fuse multiple maps with high completeness and precision. Jianhua Zhang 0002, Kaiqi Chen 0001, Zhiying Pan, Ruyu Liu, Thomas Yang 0001, Shengyong Chen |
IEEE Trans. Ind. Informatics | 8 |
| 2021 | Evolutionary Human-UAV Cooperation for Transmission Network RestorationabstractPower transmission networks are vulnerable to natural or man-made disasters, and it is of critical importance to efficiently restore damaged power supply in disaster-affected areas. A large-scale damaged transmission network can contain many faults that are initially uninspected/unlocated. Using unmanned aerial vehicles (UAVs) to inspect these faults can significantly improve the efficiency of subsequent restoration performed by human operators. Such a cooperative human-UAV scheduling problem is highly complex due to the correlation between UAV schedules and human-team schedules. In this article, we propose a cooperative evolutionary algorithm that simultaneously evolves two populations, one of UAV scheduling solutions (U-solutions) and the other of human-team scheduling solutions (H-solutions), which cooperate by determining a best matching U-solution for each H-solution and evaluating U-solutions based on a surrogate objective function that is iteratively improved by feedback from H-solutions. Our algorithm exhibits significant performance advantages over the state-of-the-arts on various test instances and an application to transmission network restoration in the 2017 Jiuzhaigou earthquake. Yujun Zheng 0001, Yi-Chen Du, Zhenglian Su, Haifeng Ling, Min-Xia Zhang, Shengyong Chen |
IEEE Trans. Ind. Informatics | 6 |
| 2021 | A Pairwise Attentive Adversarial Spatiotemporal Network for Cross-Domain Few-Shot Action Recognition-R2abstractAction recognition is a popular research topic in the computer vision and machine learning domains. Although many action recognition methods have been proposed, only a few researchers have focused on cross-domain few-shot action recognition, which must often be performed in real security surveillance. Since the problems of action recognition, domain adaptation, and few-shot learning need to be simultaneously solved, the cross-domain few-shot action recognition task is a challenging problem. To solve these issues, in this work, we develop a novel end-to-end pairwise attentive adversarial spatiotemporal network (PASTN) to perform the cross-domain few-shot action recognition task, in which spatiotemporal information acquisition, few-shot learning, and video domain adaptation are realised in a unified framework. Specifically, the Resnet-50 network is selected as the backbone of the PASTN, and a 3D convolution block is embedded in the top layer of the 2D CNN (ResNet-50) to capture the spatiotemporal representations. Moreover, a novel attentive adversarial network architecture is designed to align the spatiotemporal dynamics actions with higher domain discrepancies. In addition, the pairwise margin discrimination loss is designed for the pairwise network architecture to improve the discrimination of the learned domain-invariant spatiotemporal feature. The results of extensive experiments performed on three public benchmarks of the cross-domain action recognition datasets, including SDAI Action I, SDAI Action II and UCF50-OlympicSport, demonstrate that the proposed PASTN can significantly outperform the state-of-the-art cross-domain action recognition methods in terms of both the accuracy and computational time. Even when only two labelled training samples per category are considered in the office1 scenario of the SDAI Action I dataset, the accuracy of the PASTN is improved by 6.1%, 10.9%, 16.8%, and 14% compared to that of the $TA^{3}N$ , TemporalPooling, I3D, and P3D methods, respectively. Zan Gao 0001, Leming Guo, Weili Guan, Anan Liu, Tongwei Ren, Shengyong Chen |
IEEE Trans. Image Process. | 6 |
| 2021 | Human Interaction Understanding With Joint Graph Decomposition and Node LabelingabstractThe task of human interaction understanding involves both recognizing the action of each individual in the scene and decoding the interaction relationship among people, which is useful to a series of vision applications such as camera surveillance, video-based sports analysis and event retrieval. This paper divides the task into two problems including grouping people into clusters and assigning labels to each of them, and presents an approach to solving these problems in a joint manner. Our method does not assume the number of groups is known beforehand as this will substantially restrict its application. With the observation that the two challenges are highly correlated, the key idea is to model the pairwise interacting relations among people via a complete graph and its associated energy function such that the labeling and grouping problems are translated into the minimization of the energy function. We implement this joint framework by fusing both deep features and rich contextual cues, and learn the fusion parameters from data. An alternating search algorithm is developed in order to efficiently solve the associated inference problem. By combining the grouping and labeling results obtained with our method, we are able to achieve the semantic-level understanding of human interactions. Extensive experiments are performed to qualitatively and quantitatively evaluate the effectiveness of our approach, which outperforms state-of-the-art methods on several important benchmarks. An ablation study is also performed to verify the effectiveness of different modules within our approach. Zhenhua Wang 0003, Jinchao Ge, Dongyan Guo, Jianhua Zhang 0002, Yanjing Lei, Shengyong Chen |
IEEE Trans. Image Process. | 6 |
| 2021 | Deep High-Resolution Representation Learning for Cross-Resolution Person Re-IdentificationabstractPerson re-identification (re-ID) tackles the problem of matching person images with the same identity from different cameras. In practical applications, due to the differences in camera performance and distance between cameras and persons of interest, captured person images usually have various resolutions. This problem, named Cross-Resolution Person Re-identification, presents a great challenge for the accurate person matching. In this paper, we propose a Deep High-Resolution Pseudo-Siamese Framework (PS-HRNet) to solve the above problem. Specifically, we first improve the VDSR by introducing existing channel attention (CA) mechanism and harvest a new module, i.e., VDSR-CA, to restore the resolution of low-resolution images and make full use of the different channel information of feature maps. Then we reform the HRNet by designing a novel representation head, HRNet-ReID, to extract discriminating features. In addition, a pseudo-siamese framework is developed to reduce the difference of feature distributions between low-resolution images and high-resolution images. The experimental results on five cross-resolution person datasets verify the effectiveness of our proposed approach. Compared with the state-of-the-art methods, the proposed PS-HRNet improves the Rank-1 accuracy by 3.4%, 6.2%, 2.5%,1.1% and 4.2% on MLR-Market-1501, MLR-CUHK03, MLR-VIPeR, MLR-DukeMTMC-reID, and CAVIAR datasets, respectively, which demonstrates the superiority of our method in handling the Cross-Resolution Person Re-ID task. Our code is available at https://github.com/zhguoqing. Guoqing Zhang 0002, Zhicheng Dong 0001, Hao Wang 0101, Yuhui Zheng, Shengyong Chen |
IEEE Trans. Image Process. | 6 |
| 2021 | Parallel Connected LSTM for Matrix Sequence Prediction with Elusive CorrelationsabstractThis article is about a challenging problem called matrix sequence prediction, which is motivated from the application of taxi order prediction. Remarkably, the problem differs greatly from previous sequence prediction tasks in the sense that the time-wise correlations are quite elusive; namely, distant entries could be strongly correlated and nearby entries are unnecessarily related. Such distinct specifics make prevalent convolution-recurrence-based methods inadequate to apply. To remedy this trouble, we propose a novel architecture called Parallel Connected LSTM (PcLSTM), which integrates two new mechanisms, Multi-channel Linearized Connection (McLC) and Adaptive Parallel Unit (APU), into the framework of LSTM. Benefiting from the strengths of McLC and APU, our PcLSTM is able to handle well both the elusive correlations within each timestamp and the temporal dependencies across different timestamps, achieving state-of-the-art performance in a set of experiments demonstrated on synthetic and real-world datasets. Qi Zhao 0013, Chuqiao Chen, Guangcan Liu, Qingshan Liu 0001, Shengyong Chen |
ACM Trans. Intell. Syst. Technol. | 5 |
| 2021 | DCR: A Unified Framework for Holistic/Partial Person ReIDabstractPerson reidentification (ReID) is a very popular research topic in machine learning and computer vision. According to the occlusions, it can be divided into holistic person ReID and partial person ReID tasks. Occlusions commonly exist in the partial person ReID task but pose or observation perspective changes often occur in the holistic person ReID task; thus, many different algorithms or different network architectures have been designed for each task. However, this approach increases the cost in practice and hinders the development of ReID techniques. To solve this problem, in this work, a unified framework is proposed for holistic/partial person ReID, which can effectively and efficiently address changes in pose or observation perspective and the occlusions in both tasks. In detail, we first employ a fully convolutional network (FCN) to generate feature maps for an arbitrarily sized image and then use spatial pyramid pooling (SPP) to obtain its spatial pyramid feature. Thereafter, to efficiently solve the matching problem between the query image and gallery images, we build a deep spatial pyramid feature collaborative reconstruction model (DCR). In DCR, the reconstruction errors mainly come from similar blocks (uncovered parts), and the influence of the reconstruction errors of dissimilar blocks (covered parts or changed parts) is minimized. In addition, we also use the deep mutual learning approach to jointly learn the features in the training process and promote model training. Experimental results on two partial person ReID datasets and three holistic person ReID datasets demonstrate that the DCR outperforms the state-of-the-art approaches on both tasks and all datasets. Specifically, it outperforms all competitors with a large margin and achieves an improvement of 9.07% and 5.95% over the DSR method (published in CVPR18) on the Partial ReID and Partial-iLIDS datasets with Rank-1, respectively. Similarly, it also achieves an improvement of 5.08% over the VPM method (published in CVPR19) on the DukeMTMC-ReID dataset with Rank-1. Additionally, the running time of our method for each query is more than 7 faster than that of the DSR or DuATM methods. Zan Gao 0001, Li-Shuai Gao, Hua Zhang 0003, Zhiyong Cheng 0001, Richang Hong, Shengyong Chen |
IEEE Trans. Multim. | 6 |
| 2021 | 3D Tensor Auto-encoder with Application to Video CompressionabstractAuto-encoder has been widely used to compress high-dimensional data such as the images and videos. However, the traditional auto-encoder network needs to store a large number of parameters. Namely, when the input data is of dimension n , the number of parameters in an auto-encoder is in general O ( n ). In this article, we introduce a network structure called 3D Tensor Auto-Encoder (3DTAE). Unlike the traditional auto-encoder, in which a video is represented as a vector, our 3DTAE considers videos as 3D tensors to directly pass tensor objects through the network. The weights of each layer are represented by three small matrices, and thus the number of parameters in 3DTAE is just O ( n 1/3). The compact nature of 3DTAE fits well the needs of video compression. Given an ensemble of high-dimensional videos, we represent them as 3DTAE networks plus some small core tensors, and we further quantize the network parameters and the core tensors to get the final compressed data. Experimental results verify the efficiency of 3DTAE. Yang Li 0039, Guangcan Liu, Yubao Sun, Qingshan Liu 0001, Shengyong Chen |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2020 | SiamCAR: Siamese Fully Convolutional Classification and Regression for Visual TrackingabstractBy decomposing the visual tracking task into two subproblems as classification for pixel category and regression for object bounding box at this pixel, we propose a novel fully convolutional Siamese network to solve visual tracking end-to-end in a per-pixel manner. The proposed framework SiamCAR consists of two simple subnetworks: one Siamese subnetwork for feature extraction and one classification-regression subnetwork for bounding box prediction. Different from state-of-the-art trackers like Siamese-RPN, SiamRPN++ and SPM, which are based on region proposal, the proposed framework is both proposal and anchor free. Consequently, we are able to avoid the tricky hyper-parameter tuning of anchors and reduce human intervention. The proposed framework is simple, neat and effective. Extensive experiments and comparisons with state-of-the-art trackers are conducted on challenging benchmarks including GOT-10K, LaSOT, UAV123 and OTB-50. Without bells and whistles, our SiamCAR achieves the leading performance with a considerable real-time speed. The code is available at https://github.com/ohhhyeahhh/SiamCAR. Dongyan Guo, Zhenhua Wang 0003, Shengyong Chen |
CVPR | 5 |
| 2020 | Fused 3-Stage Image Segmentation for Pleural Effusion Cell ClustersabstractThe appearance of tumor cell clusters in pleural effusion is usually a vital sign of cancer metastasis. Segmentation, as an indispensable basis, is of crucial importance for diagnosing, chemical treatment, and prognosis in patients. However, accurate segmentation of unstained cell clusters containing more detailed features than the fluorescent staining images remains to be a challenging problem due to the complex background and the unclear boundary. Therefore, in this paper, we propose a fused 3-stage image segmentation algorithm, namely Coarse segmentation-Mapping-Fine segmentation (CMF) to achieve unstained cell clusters from whole slide images. Firstly, we establish a tumor cell cluster dataset consisting of 107 sets of images, with each set containing one unstained image, one stained image, and one ground-truth image. Then, according to the features of the unstained and stained cell clusters, we propose a three-stage segmentation method: 1) Coarse segmentation on stained images to extract suspicious cell regions-Region of Interest (ROI); 2) Mapping this ROI to the corresponding unstained image to get the ROI of the unstained image (UI-ROI); 3) Fine Segmentation using improved automatic fuzzy clustering framework (AFCF) on the UI-ROI to get precise cell cluster boundaries. Experimental results on 107 sets of images demonstrate that the proposed algorithm can achieve better performance on unstained cell clusters with an F1 score of 90.40%. Sike Ma, Meng Zhao 0001, Hao Wang 0003, Fan Shi 0001, Xuguo Sun, Shengyong Chen, Hongning Dai |
ICPR | 6 |
| 2020 | Object-oriented Map Exploration and Construction Based on Auxiliary Task Aided DRLabstractEnvironment exploration by autonomous robots through deep reinforcement learning (DRL) based methods has attracted more and more attention. However, existing methods usually focus on robot navigation to single or multiple fixed goals, while ignoring the perception and construction of external environments. In this paper, we propose a novel environment exploration task based on DRL, which requires a robot fast and completely perceives all objects of interest, and reconstructs their poses in a global environment map, as much as the robot can do. To this end, we design an auxiliary task aided DRL model, which is integrated with the auxiliary object detection and 6-DoF pose estimation components. The outcome of auxiliary tasks can improve the learning speed and robustness of DRL, as well as the accuracy of object pose estimation. Comprehensive experimental results on the indoor simulation platform AI2-THOR have shown the effectiveness and robustness of our method. Junzhe Xu 0001, Jianhua Zhang 0002, Shengyong Chen, Honghai Liu 0001 |
ICPR | 3 |
| 2020 | SpectralSeaNet: Spectrogram and Convolutional Network-based Sea State EstimationabstractSea State is significant to the operations on the sea. The traditional model-based approaches need lots of knowledge of vessels, which limit the real-world use. This paper proposes a spectrogram-based deep learning model for sea state estimation (SpectralNet). In this model, the ship motion data is converted to spectrogram using short time Fourier transform (STFT). Unlike other methods, the spectrogram of each sensor will be combined to a new image. And then, a 2D convolutional neural network (CNN) is built as the classifier and the sea state can be identified. The experimental results show the proposed approach can achieve higher classification accuracy compared these methods applied directly in raw time series data. Through the comparison results of the proposed approach and the combination of spectrogram of different number of sensors, the proposed approach can achieve highest classification accuracy, and the classification accuracy is growing with the number of combined sensors. The sensitivity analysis finds the classification accuracy is easily influenced by the scale factor of images. Xu Cheng 0003, Guoyuan Li, Robert Skulstad, Houxiang Zhang, Shengyong Chen |
IECON | 5 |
| 2020 | CalibRCNN: Calibrating Camera and LiDAR by Recurrent Convolutional Neural Network and Geometric ConstraintsabstractIn this paper, we present Calibration Recurrent Convolutional Neural Network (CalibRCNN) to infer a 6 degrees of freedom (DOF) rigid body transformation between 3D LiDAR and 2D camera. Different from the existing methods, our 3D-2D CalibRCNN not only uses the LSTM network to extract the temporal features between 3D point clouds and RGB images of consecutive frames, but also uses the geometric loss and photometric loss obtained by the interframe constraint to refine the calibration accuracy of the predicted transformation parameters. The CalibRCNN aims at inferring the correspondence between projected depth image and RGB image to learn the underlying geometry of 2D-3D calibration. Thus, the proposed calibration model achieves a good generalization ability to adapt to unknown initial calibration error ranges, and other 3D LiDAR and 2D camera pairs with different intrinsic parameters from the training dataset. Extensive experiments have demonstrated that our CalibRCNN can achieve state-of-the-art accuracy by comparison with other CNN based methods. Jieying Shi, Ziheng Zhu, Jianhua Zhang 0002, Ruyu Liu, Zhenhua Wang 0003, Shengyong Chen, Honghai Liu 0001 |
IROS | 6 |
| 2020 | Deep Adversarial Discrete Hashing for Cross-Modal RetrievalabstractCross-modal hashing has received widespread attentions on cross-modal retrieval task due to its superior retrieval efficiency and low storage cost. However, most existing cross-modal hashing methods learn binary codes directly from multimedia data, which cannot fully utilize the semantic knowledge of the data. Furthermore, they cannot learn the ranking based similarity relevance of data points with multi-label. And they usually use a relax constraint of hash code which causes non-negligible quantization loss in the optimization. In this paper, a hashing method called Deep Adversarial Discrete Hashing (DADH) is proposed to address these issues for cross-modal retrieval. The proposed method uses adversarial training to learn features across modalities and ensure the distribution consistency of feature representations across modalities. We also introduce a weighted cosine triplet constraint which can make full use of semantic knowledge from the multi-label to ensure the precise ranking relevance of item pairs. In addition, we use a discrete hashing strategy to learn the discrete binary codes without relaxation, by which the semantic knowledge from label in the hash codes can be preserved while the quantization loss can be minimized. Ablation experiments and comparison experiments on two cross-modal databases show that the proposed DADH improves the performance and outperforms several state-of-the-art hashing methods for cross-modal retrieval. Cong Bai, Jinglin Zhang 0001, Shengyong Chen |
ICMR | 5 |
| 2020 | Improving auto-encoder novelty detection using channel attention and entropy minimizationabstractNovelty detection is a important research area which mainly solves the classification problem of inliers which usually consists of normal samples and outliers composed of abnormal samples. Auto-encoder is often used for novelty detection. However, the generalization ability of the auto-encoder may cause the undesirable reconstruction of abnormal elements and reduce the identification ability of the model. To solve the problem, we focus on the perspective of better reconstructing the normal samples as well as retaining the unique information of normal samples to improve the performance of auto-encoder for novelty detection. Firstly, we introduce attention mechanism into the task. Under the action of attention mechanism, auto-encoder can pay more attention to the representation of inlier samples through adversarial training. Secondly, we apply the information entropy into the latent layer to make it sparse and constrain the expression of diversity. Experimental results on three public datasets show that the proposed method achieves comparable performance compared with previous popular approaches. Dongyan Guo, Shengyong Chen |
MMAsia | 5 |
| 2020 | Instance Image Retrieval with Generative Adversarial Training
Cong Bai, Ling Huang 0003, Yu-Gang Jiang 0001, Shengyong Chen |
MMM (1) | 5 |
| 2020 | Extreme learning machine with feature mapping of kernel functionabstractKernel‐based extreme learning machine (KELM) solves the problem of random initialisation of extreme learning machine (ELM), and it has a faster learning speed and higher learning accuracy. However, when it comes to a scenario in which the dimensionality of kernel function mapping space is less than the number of samples, the kernel function theoretically cannot be introduced into ELM. To solve this problem, ELM with feature mapping (FM) of kernel function (FM‐KELM) is proposed in this study, in which the random FM between the input layer and hidden layer of ELM is replaced with the FM of the kernel function. Moreover, the authors prove that when the regularised parameter C is close to zero, the solution of introduced kernel function is approximately equal to the correct solution. The proposed algorithm is more robust than KELM for the parameter C. Several experimental results show that the proposed algorithm in this study achieves higher classification accuracy without excessive parameter tuning, and the duration of the training and testing process is significantly reduced. Zhaoxi Wang, Shengyong Chen, Rongwei Guo, Yangbo Feng |
IET Image Process. | 2 |
| 2020 | Reachability Analysis of Networked Finite State Machine With Communication Losses: A Switched PerspectiveabstractNetworked finite state machine takes into account communication losses in industrial communication interfaces due to the limited bandwidth. The reachability analysis of networked finite state machine is a fundamental and important research topic in blocking detection, safety analysis, communication system design and so on. This paper is concerned with the impact of arbitrary communication losses on the reachability of networked finite state machine from a switched perspective. First, to model the dynamics under arbitrary communication losses in communication interfaces (from the controller to the actuator), by resorting to the semi-tensor product (STP) of matrices, a switched algebraic model of networked finite state machine with arbitrary communication losses is proposed, and the reachability analysis can be investigated by using the constructed model under arbitrary switching signal. Subsequently, based on the algebraic expression and its transition matrix, necessary and sufficient conditions for the reachability are derived for networked finite state machine. Finally, some typical numerical examples are exploited to demonstrate the effectiveness of the proposed approach. Note that current results provide valuable clues to design reliable and convergent Internet-of-Things networks. Chengyi Xia, Shengyong Chen, Thomas Yang 0001, Zengqiang Chen 0001 |
IEEE J. Sel. Areas Commun. | 3 |
| 2020 | Cross-domain representation learning by domain-migration generative adversarial network for sketch based image retrieval
Cong Bai, Jian Chen 0009, Pengyi Hao, Shengyong Chen |
J. Vis. Commun. Image Represent. | 5 |
| 2020 | Hyperspectral Image Restoration via Local Low-Rank Matrix Recovery and Moreau-Enhanced Total VariationabstractIn this letter, we present a hyperspectral image (HSI) mixed-noise removal method named Moreau-enhanced total variation (TV) regularized local low-rank matrix recovery (LLRMTV). The rank-fixed matrix recovery is first adopted to separate the low-rank clean HSI patches from the sparse noise. Then, a Moreau-enhanced TV regularized image reconstruction strategy is utilized to ensure the piecewise smoothness of the reconstructed image from the low-rank patches. The proposed Moreau-enhanced TV restoration method involves a nonconvex penalty designed to maintain the convexity of the objective function. Moreover, the proposed model is integrated into an augmented Lagrange multiplier (ALM) algorithm to produce final results, leading to a complete HSI restoration framework. Examples of restoration illustrate the improvement over the typical TV regularization. Yanhong Yang, Jianwei Zheng 0001, Shengyong Chen |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2020 | Factorized weight interaction neural networks for sparse feature prediction
Dafang Zou, Mengmeng Sheng, Hui Yu 0013, Jiafa Mao, Shengyong Chen, Weiguo Sheng 0001 |
Neural Comput. Appl. | 5 |
| 2020 | DV-Net: Dual-view network for 3D reconstruction by fusing multiple sets of gated control point clouds
Shourui Yang, Yuxin Peng 0001, Junchao Zhang 0005, Shengyong Chen |
Pattern Recognit. Lett. | 5 |
| 2020 | Local low-rank matrix recovery for hyperspectral image denoising with ℓ0 gradient constraint
Yanhong Yang, Jianwei Zheng 0001, Shengyong Chen |
Pattern Recognit. Lett. | 3 |
| 2020 | Detection and Recognition for Life State of Cell Cancer Using Two-Stage Cascade CNNsabstractCancer cell detection and its stages recognition of life cycle are an important step to analyze cellular dynamics in the automation of cell based-experiments. In this work, a two-stage hierarchical method is proposed to detect and recognize different life stages of bladder cells by using two cascade Convolutional Neural Networks (CNNs). Initially, a hybrid object proposal algorithm (called EdgeSelective) by combining EdgeBoxes and Selective Search is proposed to generate candidate object proposals instead of a single Selective Search method in Region-CNN (R-CNN), and it can exploit the advantages of different mechanisms for generating proposals so that each cell in the image can be fully contained by at least one proposed region during the detection process. Then, the obtained cells from the previous step are used to train and extract features by employing CNNs for the purpose of cell life stage recognition. Finally, a series of comparison experiments are implemented. The results show that the proposed method can obtain better performance than traditional methods either in the stage of cell detection or cell life stage recognition, and it encourages and suggests the application in the development of new anticancer drug and cytopathology analysis of cancer patients in the near future. Haigen Hu, Qiu Guan, Shengyong Chen, Zhiwei Ji |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2020 | Evolutionary Collaborative Human-UAV Search for Escaped CriminalsabstractThe use of unmanned aerial vehicles (UAVs) for target searching in complex environments has increased considerably in recent years. The numerous studies on UAV search methods have been reported, but few have been conducted on collaborative human-UAV search which is common in many applications. In this paper, we present a problem of collaborative human-UAV search for escaped criminals, the aim of which is to minimize the expected time of capture rather than detection. We show that our problem is much more complex than the problem of pure UAV search. The difficulty of our problem is further increased by the fact that criminals will attempt to avoid detection and capture. To solve the problem, we propose a hybrid evolutionary algorithm (EA) that uses three evolutionary operators, namely, comprehensive learning, variable mutation, and local search, to efficiently explore the solution space. The experimental results demonstrate that the proposed method outperforms some well-known EAs and other popular UAV search methods on test instances. An application of our method to a real-world operation took 311 min to capture a criminal who had escaped for over three days, validating its practicability and performance advantage. This paper provides a good basis for promoting the application of EAs to a wider class of man-machine collaboration scheduling problems. Yujun Zheng 0001, Yi-Chen Du, Haifeng Ling, Weiguo Sheng 0001, Shengyong Chen |
IEEE Trans. Evol. Comput. | 5 |
| 2020 | Guest Editorial: Special Section on Latest Advances on Industrial Intelligent Video Systems and AnalyticsabstractThis Special Section collects the latest developments in video system design, data compression, target detection, object localization, behavior analysis, motion detection, and real-time implementation of industrial video systems and intelligent analytics to bring the latest ideas and solutions of the research community on practical video systems to our audience. The Special Section focuses on several topics that are recently concerned in the community, including, multicamera network, real-time hardware implementation, networked data analytics, bandwidth limited compression, motion pattern analysis, action understanding, three-dimensional (3-D) reconstruction, contextual recognition, object detection and tracking, intelligent robot vision, security surveillance, intelligent transportation, and other industrial applications. The Special Section presents 14 articles on intelligent video systems and analytics. These articles are briefly summarized here. Shengyong Chen, Honghai Liu 0001, Naoyuki Kubota |
IEEE Trans. Ind. Informatics | 1 |
| 2020 | Spatiotemporal Saliency Detection Based on Maximum Consistency Superpixels Merging for Video AnalysisabstractMotion objects detection becomes more and more important in the applications of video surveillance, e.g., intrusion detection. The spatiotemporal saliency is an effective feature to describe object motion. However, there is a lot of redundancy in the spatial information preventing to obtain accurate saliency in an effective way. At the same time, temporal information cannot be accurately described because it is affected by uneven brightness, complex background, and fast-moving objects, especially at the edge of moving objects. In this article, we develop a novel method to tackle these problems and obtain more accurate spatiotemporal saliency. The key idea is the superpixel merging based on our maximum consistency model in feature space, through which the redundant spatial information is decreased and inhibit some temporal information errors. Experimental evaluations on the NNT dataset and surveillance videos show that the proposed method achieves better performance by comparing with some state-of-the-art methods, and can effectively detect intrusion entities. Jianhua Zhang 0002, Jingbo Chen, Shengyong Chen |
IEEE Trans. Ind. Informatics | 4 |
| 2020 | Multi-Task Deep Dual Correlation Filters for Visual TrackingabstractCorrelation filters combined with deep features have delivered impressive results in visual tracking task. However, existing approaches treat deep features produced by different network layers independently, limiting their representation power. To address this issue, this paper proposes a multi-task deep dual correlation filters (MDDCF) based method for robust visual tracking. First, a new multi-task learning scheme is designed to take full advantage of the multi-level features of deep networks, where target representation with individual features is regarded as a single task. As such, the interdependencies between different levels of features can be better explored. Second, we reformulate the objective function of the dual correlation filters and propose a new alternating optimization method, allowing joint training of the correlation filters and network parameters. Third, we design an effective object template update scheme which can well capture the target appearance variations. Extensive experimental evaluations on seven benchmark datasets show that the proposed MDDCF tracker performs favorably against state-ofthe-art methods. Yuhui Zheng, Xinyan Liu 0002, Xu Cheng 0003, Kaihua Zhang 0001, Yi Wu 0001, Shengyong Chen |
IEEE Trans. Image Process. | 6 |
| 2020 | Truncated Low-Rank and Total p Variation Constrained Color Image Completion and its Moreau Approximation AlgorithmabstractRecently, low-rank (LR) and total variation (TV) constrained tensor completion algorithms have been broadly studied for image restoration. These algorithms, however, ignore the difference of the intrinsic properties along spatial structure, spectral correlation, and unfolded mode. In this paper, we go further by providing a detailed comparison of the LR and TV properties in matrix and tensor cases, and figure out the LRTV constraints for pixel matrices are more evident and accordant than for others. This inspires us to develop a simple yet effective multichannel LRTV model that is capable of genuinely discovering the intrinsic properties with reduced computational cost. Moreover, due to the suboptimality of nuclear norm and$l_{1}$norm in approximating the essential low rank and low gradient properties, we employ two enhanced constraints, i.e., truncated nuclear norm (TNN) and total$p$variation${\text{T}}_{p}\text{V}$, for a better performance. This results in a challenging problem since that both TNN and${\text{T}}_{p}\text{V}$are nonsmooth and nonconvex. Observing that the Moreau approximation of${\text{T}}_{p}\text{V}$constraint is a continuous difference-of-convex function, we then develop a first-order method by repeatedly computing two simple proximal operators. Under mild assumption, we further prove that the sequence generated by our method clusters at a stationary point. Extensive experimental results on color image completion show the efficacy and efficiency of our method over state-of-the-art competitors. Jianwei Zheng 0001, Xi Yang 0006, Shengyong Chen |
IEEE Trans. Image Process. | 4 |
| 2020 | Structural Analysis of Attributes for Vehicle Re-Identification and RetrievalabstractVehicle re-identification plays an important role in video surveillance applications. Despite the efforts made on this problem in the past few years, it remains a challenging task due to various factors such as pose variation, illumination changes, and subtle inter-class difference. We believe that the key information for identification has not been well explored in the literature. In this paper, we first collect a vehicle dataset `VAC21' which contains 7129 images of five types of vehicles. Then, we carefully label the 21 classes of structural attributes hierarchically with bounding boxes. To our knowledge, this is the first dataset with several detailed attributes labeled. Based on this dataset, we use the state-of-the-art one-stage detection method, Single-shot Detection, as a baseline model for detecting attributes. Subsequently, we make a few important modifications tailored for this application to improve accuracy: 1) adding more proposals from low-level layers to improve the accuracy of detecting small objects and 2) employing the focal loss to improve the mean average precision. Furthermore, the results of the attribute detection can be applied to a series of vision tasks that focus on analyzing the images of vehicles. Finally, we propose a novel region of interests (ROIs)-based vehicle re-identification and retrieval method in which the ROIs' deep features are used as discriminative identifiers, encoding the structure information of a vehicle. These deep features are input to a boosting model to improve the accuracy. A set of experiments are conducted on the dataset VehicleID and the experimental results show that our method outperforms the state-of-the-art methods. Yanzhu Zhao, Chunhua Shen, Huibing Wang, Shengyong Chen |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2019 | MC-Unet: Multi-scale Convolution Unet for Bladder Cancer Cell Segmentation in Phase-Contrast Microscopy ImagesabstractOwing to the high density, low contrast, deformable cell shapes, low inter-cellular shape and appearance variation, and occlusion of the cells by division or fusion especially in phase-contrast microscopy images, it is still a challenging task to segment cells from the complex background. In this work, we proposed a multi-scale convolution Unet (MC-Unet) for bladder cancer cell segmentation in Phase-Contrast microscopy images. More specifically, the second 3x3 convolution of each layer in the standard Unet is replaced with a multi-scale convolution (MC) block with different kernel sizes, such as 1x1, 3x3, and 5x5. To verify the effectiveness of the proposed method, a series of experiments are conducted on the bladder cancer T24 dataset and the MoNuSeg dataset, and the results shows the proposed MC-Unet can obtain better comprehensive performance than the standard Unet. Haigen Hu, Yixing Zheng, Qianwei Zhou, Jie Xiao 0003, Shengyong Chen, Qiu Guan |
BIBM | 5 |
| 2019 | Learning A 3D Gaze Estimator with Improved Itracker Combined with Bidirectional LSTMabstractFree-head 3D gaze estimation which outputs gaze vector in 3D space has wide application in human-computer interaction. In this paper, we propose a novel 3D gaze estimator by improving the Itracker and employing a many-to-one bidirectional LSTM (bi-LSTM). First, we improve the conventional Itracker by removing the face-grid and reducing one network branch via concatenating the two-eye region images to predict the subject's gaze of a single frame. Then, we employ the bi-LSTM to fit the temporal information between frames to estimate gaze vector for video sequence. Experimental results show that our improved Itracker obtains 11.6% significant improvement over the state-of-the-art methods on MPIIGaze dataset (single image frame) and has robust estimation accuracy for different image resolutions. Moreover, experimental results on EyeDiap dataset (video sequence) further bring 3% accuracy improvement by employing the bi-LSTM. Xiaolong Zhou 0001, Jianing Lin, Shengyong Chen |
ICME | 4 |
| 2019 | Joint Grouping and Labeling via Complete Graph Decomposition
Jinchao Ge, Zhenhua Wang 0003, Jiajun Meng, Jianhua Zhang 0002, Shengyong Chen |
ICONIP (5) | 5 |
| 2019 | Modeling and Analysis of Motion Data from Dynamically Positioned Vessels for Sea State EstimationabstractDeveloping a reliable model to identify the sea state is significant for the autonomous ship. This paper introduces a novel deep neural network model (SeaStateNet) to estimate the sea state based on the ship motion data from dynamically positioned vessels. The SeaStateNet mainly consists of three components: an Long-Short-Term Memory (LSTM) recurrent neural network to capture the long dependency in the ship motion data; a convolutional neural network (CNN) to extract time-invariant features; and a Fast Fourier Transform (FFT) block to extract frequency features. A feature fusion layer is designed to learn the degree affected by each component. The proposed model is applied directly to the raw time series data, without needing of any hand-engineered features. A sensitivity analysis (SA) method is applied to assess the influence of data preprocessing. Through benchmark test and experiment on ship motion dataset, SeaStateNet is verified effective for sea state estimation. The investigation on real-time test further shows the practicality of the proposed model. Xu Cheng 0003, Guoyuan Li, Robert Skulstad, Shengyong Chen, Hans Petter Hildre, Houxiang Zhang |
ICRA | 4 |
| 2019 | Robust High Accuracy Visual-Inertial-Laser SLAM SystemabstractIn recent years, many excellent works on visual-inertial SLAM and laser-based SLAM have been proposed. Although inertial measurement unit (IMU) significantly improve the motion estimate performance by reducing the impact of illumination variation or texture-less region on visual tracking, tracking failures occur when in such an environment for a long time. Similarly, when in structure-less environments, laser module will fail since lack of sufficient geometric features. Besides, motion estimation by moving lidar has the problem of distortion since range measurements are received continuously. To solve these problems, we propose a robust and high-accuracy visual-inertial-laser SLAM system. The system starts with a visual-inertial tightly-coupled method for motion estimation, followed by scan matching to further optimize the estimation and register point cloud on the map. Furthermore, we enable modules to be adjusted automatically and flexibly. That is, when one of these modules fails, the remaining modules will undertake the motion-tracking task. For further improving the accuracy, loop closure and proximity detection are implemented to eliminate drift accumulation. When loop or proximity is detected, we perform six degree-of-freedom (6-DOF) pose graph optimization to achieve the global consistency. The performance of our system is verified on public dataset, and the experimental results show that the proposed method achieves superior accuracy against other state-of-the-art algorithms. Zengyuan Wang, Jianhua Zhang 0002, Shengyong Chen, Conger Yuan, Jingqian Zhang, Jianwei Zhang 0001 |
IROS | 3 |
| 2019 | Towards SLAM-Based Outdoor Localization using Poor GPS and 2.5D Building ModelsabstractIn this paper, we address the topic of outdoor localization and tracking using monocular camera setups with poor GPS priors. We leverage 2.5D building maps, which are freely available from open-source databases such as OpenStreetMap. The main contributions of our work are a fast initialization method and a non-linear optimization scheme. The initialization upgrades a visual SLAM reconstruction with an absolute scale. The non-linear optimization uses the 2.5D building model footprint, which further improves the tracking accuracy and the scale estimation. A pose optimization step relates the vision-based camera pose estimation from SLAM to the position information received through GPS, in order to fix the common problem of drift. We evaluate our approach on a set of challenging scenarios. The experimental results show that our approach achieves improved accuracy and robustness with an advantage in run-time over previous setups. Ruyu Liu, Jianhua Zhang 0002, Shengyong Chen, Clemens Arth |
ISMAR | 3 |
| 2019 | Dictionary Learning and Confidence Map Estimation-Based Tracker for Robot-Assisted Therapy System
Xiaolong Zhou 0001, Sixian Chan 0001, Shengyong Chen, Honghai Liu 0001 |
PRCV (1) | 4 |
| 2019 | Deep learning for multiple object tracking: a surveyabstractDeep learning has been proved effective in multiple object tracking, which confronts the difficulties of frequent occlusions, confusing appearance, in‐and‐out objects, and lack of enough labelled data. Recently, deep learning based multi‐object tracking methods make a rapid progress from representation learning to network modelling due to the development of deep learning theory and benchmark setup. In this study, the authors summarise and analyse deep learning based multi‐object tracking methods which are top‐ranked in the public benchmark test. First, they investigate functionality of deep networks in these methods, and classify the methods into three categories as description enhancement using deep features, deep network embedding, and end‐to‐end deep network construction. Second, they review deep network structures in these methods, and detail the usage and training of these networks for multi‐object tracking problem. Through experimental comparison of tracking results in the benchmarks in total and by group, they finally show the effectiveness of deep networks for tracking employed in different manners, and compare the advantages of these networks and their robustness under different tracking conditions. Moreover, they analyse the limitations of current methods, and draw some useful conclusions to facilitate the exploration of new directions for multi‐object tracking. Yingkun Xu, Xiaolong Zhou 0001, Shengyong Chen, Fenfen Li |
IET Comput. Vis. | 3 |
| 2019 | As-global-as-possible stereo matching with adaptive smoothness priorabstractMore global matching (MGM) overcomes the limitation of one‐dimensional scanline optimisation in semi‐global matching (SGM). Nevertheless, the possible weaknesses of the MGM algorithm are as follows: (i) only two directions are considered for each image traversal direction, which may lead to massive mismatches; (ii) disparity estimation around the object boundaries usually performs terrible since the smoothness term is designed independent of the image prior. In this research, the authors consider all of the four directions for each image traversal direction through a novel model. Besides utilising the prior of neighboured pixels' correlation, adaptive smoothness terms are modelled and augmented into the energy function. These contributions encourage ‘as‐global‐as‐possible (AGAP)’. More importantly, different from the recent works in which the aggregated algorithms have been conducted as the data term of an energy function, conversely, the authors make the energy function as a part of cost aggregation framework. Performance evaluations on Middlebury v.2 and v.3 stereo data sets demonstrate that the proposed AGAP outperforms other four most challenging stereo matching algorithms, and also performs better on Microsoft i2i stereo videos. In addition, under various strategies of parallelisation, the presented AGAP shows a near real‐time execution time. Hua Zhang 0003, Yanbing Xue, Shengyong Chen |
IET Image Process. | 4 |
| 2019 | Automatic segmentation of MR depicted carotid arterial boundary based on local priors and constrained global optimisationabstractSegmentation of lumen (LB) and outer wall boundaries (OB) of carotid artery in magnetic resonance (MR) images is essential for carotid atherosclerotic disease diagnosis. However, the limited image signal‐to‐noise ratio, flow artefact, and varied lumen and outer wall become significant obstacles for automatic segmentation. A fully automatic framework is proposed for LB and OB segmentation in MR images. First, the lumen is identified by the support vector machine using a special strategy and LB is segmented by the geodesic star‐shape‐constrained graph cut. Then a novel global optimisation is developed to segment OB based on the graph cut, which consists of shape priors and appearance priors. The shape priors are learned from labelled shapes on LB and OB, while the appearance priors are modelled by Gaussian mixture models. A novel shape constraint is also designed as the constraint term. To evaluate author's method, extensive experiments are carried out from 160 MR images belonging to 16 patients. Experimental results demonstrate that the proposed method can yield high accuracy with fully automatic segmentation. Moreover, the advantages of the proposed method have been shown in terms of high flexibility and accuracy without user interactions in comparison with other methods. Jianhua Zhang 0002, Zhongzhao Teng, Qiu Guan, Junli He, Wafa Abutaleb, Andrew J. Patterson, Martin J. Graves, Jonathan Gillard 0001, Shengyong Chen |
IET Image Process. | 9 |
| 2019 | Moving object detection via segmentation and saliency constrained RPCA
Yang Li 0039, Guangcan Liu, Qingshan Liu 0001, Yubao Sun, Shengyong Chen |
Neurocomputing | 5 |
| 2019 | A multilevel sampling strategy based memetic differential evolution for multimodal optimization
Mengmeng Sheng, Kangfei Ye, Jiafa Mao, Shengyong Chen, Weiguo Sheng 0001 |
Neurocomputing | 6 |
| 2019 | Three dimensional object segmentation based on spatial adaptive projection for solid waste
Jianhua Zhang 0002, Yeqiang Qiu, Jianshuang Guo, Jingbo Chen, Shengyong Chen |
Neurocomputing | 6 |
| 2019 | Supervised learning based discrete hashing for image retrieval
Cong Bai, Jinglin Zhang 0001, Zhi Liu 0003, Shengyong Chen |
Pattern Recognit. | 5 |
| 2019 | Very large-scale data classification based on K-means clustering and multi-kernel SVM
Tinglong Tang, Shengyong Chen, Meng Zhao 0001, Wei Huang 0015, Jake Luo |
Soft Comput. | 2 |
| 2019 | A hybrid biogeography-based optimization and fuzzy C-means algorithm for image segmentation
Min-Xia Zhang, Weixuan Jiang, Xiao-Han Zhou, Yu Xue 0003, Shengyong Chen |
Soft Comput. | 5 |
| 2019 | A Hierarchical Model for Human Action Recognition From Body-PartsabstractAs increasing attention is paid to human action recognition from skeleton data, this paper focuses on such tasks by proposing a hierarchical model to discover the structure information of body-parts involved in actions for better analysis of human actions in the skeleton data. Considering human actions as simultaneous motions of body-parts of the human skeleton, we propose a hierarchical model to simultaneously apply discriminative body-parts selection at a same scale and group coupling of bundles of body-parts at different scales, while we decompose the human skeleton into a hierarchy of body-parts of varying scales. To represent such hierarchy of body-parts, we accordingly build a hierarchical rotation and relative velocity (HRRV) descriptor. The hierarchical representations encoded by Fisher vectors of the HRRV descriptors are properly formulated into the hierarchical model via the proposed mixed norm, to apply the sparse selection of body-parts and regularize the structure of such hierarchy of body-parts. The extensive evaluations on three challenging datasets demonstrate the effectiveness of our proposed approach, which achieves superior performance compared to the state-of-the-art algorithms on datasets with various sizes, showing it is more widely applicable than existing approaches. Zhanpeng Shao, Youfu Li 0001, Yao Guo 0002, Xiaolong Zhou 0001, Shengyong Chen |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2019 | Hierarchical Topic Model Based Object Association for Semantic SLAMabstractObject-based simultaneous localization and mapping (SLAM) is a more natural and robust way for agents to interact with their surrounding environment. However, it introduces a problem of semantic objects association. Correct object association is the key factor to achieve a successful object SLAM system because object association and SLAM are inherently coupled and have not been well tackled yet. A novel formulation of the object association problem based on a hierarchical Dirichlet process (HDP) is proposed. Through the HDP, we can hierarchically associate the grouped object measurements. This can improve the object association accuracy and computation efficiency. Thanks to the novel formulation, the proposed method is also able to correct failure object associations according to its sampling inference algorithm. Furthermore, we introduce object poses to the processing of pose optimization. The object association and pose optimization are then solved in a tightly coupled way, by which both aspects can promote each other. The proposed method is evaluated on indoor and outdoor datasets and the experimental results show a very impressive improvement with respect to the traditional SLAM. Jianhua Zhang 0002, Mengping Gui, Ruyu Liu, Junzhe Xu 0001, Shengyong Chen |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2018 | Feature Selection Mechanism in CNNs for Facial Expression Recognition
Shuwen Zhao, Haibin Cai, Honghai Liu 0001, Jianhua Zhang 0002, Shengyong Chen |
BMVC | 5 |
| 2018 | Instant SLAM Initialization for Outdoor Omnidirectional Augmented RealityabstractThe initialization and absolute scale are two critical issues for an Augmented Reality (AR) system. Most existing methods have to resort to some external sensors or some special steps, to initialize an AR system and to obtain a correct scale. In this paper, we introduce an omnidirectional AR system, which can be instantly initialized, recover the absolute scale without other sensor data, and provide a full 360-degrees field-of-view (Fov) to give users the best immersive feelings. Our system shows how to firstly estimate the absolute orientation using the meaningful line and point cues from a single panorama, and how to estimate the camera global position by aligning a panorama after semantic segmentation with a widely available 2.5D map. Based on resulting absolute pose from a single frame, we subsequently render a depth map to initialize a SLAM system. We can then fuse the virtual elements with a real scene according to the continuous camera motion from the SLAM system. We evaluate the SLAM initialization approach on a challenging dataset. The experiments indicate that localization precision from our method is obviously superior to that from consumer GPS devices and we remain unbeatable in time performance compared to previous methods. Ruyu Liu, Jianhua Zhang 0002, Kejie Yin, Jia-xin Wu, Ruihao Lin, Shengyong Chen |
CASA | 6 |
| 2018 | A Multi-channel Multi-classifier Method for Classifying Pancreatic Cystic Neoplasms Based on ResNet
Haigen Hu, Kangjie Li, Qiu Guan, Feng Chen 0038, Shengyong Chen, Yicheng Ni |
ICANN (2) | 5 |
| 2018 | AGO: Accelerating Global Optimization for Accurate Stereo Matching
Hua Zhang 0003, Yanbing Xue, Shengyong Chen |
MMM (1) | 4 |
| 2018 | Oscillation Detection and Parameter-Adaptive Hedge Algorithm for Real-Time Visual Tracking
Bolin Lv, Xiaolong Zhou 0001, Shengyong Chen |
PRCV (4) | 3 |
| 2018 | Deep CRF-Graph Learning for Semantic Image Segmentation
Fuguang Ding, Zhenhua Wang 0003, Dongyan Guo, Shengyong Chen, Jianhua Zhang 0002, Zhanpeng Shao |
PRICAI | 4 |
| 2018 | Siamese Network Based Features Fusion for Adaptive Visual Tracking
Dongyan Guo, Weixuan Zhao, Zhenhua Wang 0003, Shengyong Chen, Jian Zhang 0002 |
PRICAI (1) | 5 |
| 2018 | Absolute Orientation and Localization Estimation from an Omnidirectional Image
Ruyu Liu, Jianhua Zhang 0002, Kejie Yin, Zhiying Pan, Ruihao Lin, Shengyong Chen |
PRICAI | 6 |
| 2018 | MSCS: MeshStereo with Cross-Scale Cost Filtering for fast stereo matchingabstractMeshStereo (MS) and cross‐scale cost filtering (CSCF) are two most recently celebrated models for stereo matching. On one hand, MS model enlightens for fast solving the dense stereo correspondence problem according to a region‐based opinion. On the other hand, CSCF model could generate more robust matching cost volumes than single scale. In this study, the authors weave these two models together for attaining greater and faster disparity estimation. With CSCF, more powerful initial volumes of matching cost are computed and they are conducted as the data term of MS energy function model. More importantly, the novel‐fused stereo model also draws a closer connection between multi‐scale aggregated and global algorithms. Integrating the advantages of both stereo models, they name the presented one as MS with cross‐scale (MSCS). Performance evaluations on Middlebury v.2 and v.3 stereo data sets demonstrate that the proposed MSCS outperforms other four most challenging stereo matching algorithms; and also performs better on Microsoft i2i stereo videos. In addition, thanks to this novel‐fused model, MSCS requires fewer iteration times for optimising and makes it surprisingly possesses a much faster execution time. Hua Zhang 0003, Yanbing Xue, Shengyong Chen |
IET Comput. Vis. | 4 |
| 2018 | Optimization of deep convolutional neural network for large scale image retrieval
Cong Bai, Ling Huang 0003, Jianwei Zheng 0001, Shengyong Chen |
Neurocomputing | 5 |
| 2018 | Online classification for object tracking based on superpixel
Sixian Chan 0001, Xiaolong Zhou 0001, Shengyong Chen |
Neurocomputing | 3 |
| 2018 | A fast online multivariable identification method for greenhouse environment control problems
Haigen Hu, Qiu Guan, Xiaoxin Li 0001, Shengyong Chen, Qianwei Zhou |
Neurocomputing | 5 |
| 2018 | Fast subspace segmentation via Random Sample Probing
Yang Li 0039, Yubao Sun, Qingshan Liu 0001, Shengyong Chen |
Neurocomputing | 4 |
| 2018 | Understanding human activities in videos: A joint action and interaction learning approach
Zhenhua Wang 0003, Jiali Jin, Sheng Liu 0002, Jianhua Zhang 0002, Shengyong Chen, Zhen Zhang 0008, Dongyan Guo, Zhanpeng Shao |
Neurocomputing | 6 |
| 2018 | Saliency-based multi-feature modeling for semantic image retrieval
Cong Bai, Jianan Chen 0002, Ling Huang 0003, Kidiyo Kpalma, Shengyong Chen |
J. Vis. Commun. Image Represent. | 5 |
| 2018 | Object-level saliency: Fusing objectness estimation and saliency detection into a uniform framework
Jianhua Zhang 0002, Yanzhu Zhao, Shengyong Chen |
J. Vis. Commun. Image Represent. | 3 |
| 2018 | Generative adversarial network based telecom fraud detection at the receiving bank
Yujun Zheng 0001, Xiao-Han Zhou, Weiguo Sheng 0001, Yu Xue 0003, Shengyong Chen |
Neural Networks | 5 |
| 2017 | Segment-tree based cost aggregation for stereo matching with enhanced segmentation advantageabstractSegment-tree (ST) based cost aggregation algorithm for stereo matching successfully integrates the information of segmentation with non-local cost aggregation framework. The tree structure which is generated by the segmentation strategy directly determines the final results for this kind of algorithms. However, the original strategy performs unreasonable due to its coarse performance and ignores to meet the disparity consistency assumption. To improve these weaknesses we propose a novel segmentation algorithm for constructing a more faithful ST with enhanced segmentation advantage according to a robust initial over-segmentation. Then we implement non-local cost aggregation framework on this new ST structure and obtain improved disparity maps. Performance evaluations on all 31 Middlebury stereo pairs show that the proposed algorithm outperforms than other five state-of-the-art aggregated based algorithms and also keeps time efficiency. Hua Zhang 0003, Yanbing Xue, Mian Zhou, Guangping Xu, Zan Gao 0002, Shengyong Chen |
ICASSP | 7 |
| 2017 | Joint label-interaction learning for human action recognitionabstractHuman interactions and their action categories preserve strong correlations, and the identification of the interaction configuration is of significant importance to improve the action recognition result. However, interactions are typically estimated using heuristics or treated as latent variables. The former usually produces incorrect interaction configuration while the latter introduces challenging training problem. Hence we propose a framework to jointly learn interactions and actions by designing a potential function using both features learned via deep neural networks and human interaction context. We propose an iterative approach to solve the associated inference problem efficiently and approximately. Experimental results on real datasets demonstrate that the proposed approach outperforms baselines by a large margin, and is competitive compared with the state-of-the-arts. Jiali Jin, Zhenhua Wang 0003, Sheng Liu 0002, Jianhua Zhang 0002, Shengyong Chen, Qiu Guan |
ICIP | 5 |
| 2017 | SPMVP: Spatial PatchMatch Stereo with Virtual Pixel Aggregation
Hua Zhang 0003, Yanbing Xue, Shengyong Chen |
ICONIP (3) | 4 |
| 2017 | Compressive tracking with locality sensitive histograms featuresabstractCurrently, Compressive Tracking (CT) method has drawn great attention because of its high efficiency. However, it cannot well deal with some appearance variations due to its limitations of feature expression and it only uses a fixed parameter to update the appearance model. In order to handle such matters, we propose an adaptive CT method that combines the predicted target position with CT based on Locality Sensitive Histograms (LSH) features. Our method significantly improves CT in four aspects. First, the efficient illumination invariant features extracted based on LSH are used to represent an effective appearance model that is robust to illumination changes. Second, the color attributes tracker is adopted to predict the target position for re-building the new weighted discriminant function which brings in the color information to make up for the inadequacy of Haar-like characteristics. Third, a new model update mechanism is proposed to preserve the stable features while avoid the noisy appearance variations during tracking. Fourth, a trajectory rectification method is employed to refine the tracking location when possible inaccurate tracking occurs. Finally, we show that our tracker achieves state-of-the-art performance in a comprehensive evaluation over 47 challenging color sequences. Sixian Chan 0001, Xiaolong Zhou 0001, Zhuo Zhang 0012, Shengyong Chen |
ICRA | 4 |
| 2017 | Object tracking using a convolutional network and a structured output SVMabstractObject tracking has been a challenge in computer vision. In this paper, we present a novel method to model target appearance and combine it with structured output learning for robust online tracking within a tracking-by-detection framework. We take both convolutional features and handcrafted features into account to robustly encode the target appearance. First, we extract convolutional features of the target by kernels generated from the initial annotated frame. To capture appearance variation during tracking, we propose a new strategy to update the target and background kernel pool. Secondly, we employ a structured output SVM for refining the target’s location to mitigate uncertainty in labeling samples as positive or negative. Compared with existing state-of-the-art trackers, our tracking method not only enhances the robustness of the feature representation, but also uses structured output prediction to avoid relying on heuristic intermediate steps to produce labelled binary samples. Extensive experimental evaluation on the challenging OTB-50 video sequences shows competitive results in terms of both success and precision rate, demonstrating the merits of the proposed tracking method. Xiaolong Zhou 0001, Sixian Chan 0001, Shengyong Chen |
Comput. Vis. Media | 4 |
| 2017 | A niching evolutionary algorithm with adaptive negative correlation learning for neural network ensemble
Weiguo Sheng 0001, Pengxiao Shan, Shengyong Chen, Yurong Liu, Fuad E. Alsaadi |
Neurocomputing | 3 |
| 2017 | Discriminative Histogram Intersection Metric Learning and Its Applications
Pengyi Hao, Shengyong Chen |
J. Comput. Sci. Technol. | 5 |
| 2017 | Adaptive Compressive Tracking based on Locality Sensitive Histograms
Sixian Chan 0001, Xiaolong Zhou 0001, Shengyong Chen |
Pattern Recognit. | 4 |
| 2017 | Fireworks Algorithm with Enhanced Fireworks InteractionabstractAs a relatively new metaheuristic in swarm intelligence, fireworks algorithm (FWA) has exhibited promising performance on a wide range of optimization problems. This paper aims to improve FWA by enhancing fireworks interaction in three aspects: 1) Developing a new Gaussian mutation operator to make sparks learn from more exemplars; 2) Integrating the regular explosion operator of FWA with the migration operator of biogeography-based optimization (BBO) to increase information sharing; 3) Adopting a new population selection strategy that enables high-quality solutions to have high probabilities of entering the next generation without incurring high computational cost. The combination of the three strategies can significantly enhance fireworks interaction and thus improve solution diversity and suppress premature convergence. Numerical experiments on the CEC 2015 single-objective optimization test problems show the effectiveness of the proposed algorithm. The application to a high-speed train scheduling problem also demonstrates its feasibility in real-world optimization problems. Bei Zhang 0004, Yujun Zheng 0001, Min-Xia Zhang, Shengyong Chen |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2017 | A Spatio-Temporal CRF for Human Interaction UnderstandingabstractA better understanding of human interactions in videos can be achieved by simultaneously considering the coarse interactions between people, the action of each individual, and the activity of all people as a whole. We divide the recognition task into two stages. The first stage discriminates interactions and noninteractions, actions and activities based on local image information, while during the second stage, actions and activities are recognized in a global manner based on the local recognition results. A conditional random field (CRF) is designed to model human interactions in the spatio-temporal space. Different from most existing global models which cover either action or activity variables only, our model covers them both by considering the interactions between different types of variables. The graph structure of the CRF is predicted by a model learned from training data, which is different from traditional graph construction methods that typically rely on human heuristics. We learn the parameters of the CRF via structured support vector machine. We propose an efficient inference algorithm to tackle the estimation of labels in long videos containing many people. Our model admits both semantic-level understanding of human interactions in videos and competitive action and activity recognition performance. Zhenhua Wang 0003, Sheng Liu 0002, Jianhua Zhang 0002, Shengyong Chen, Qiu Guan |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2017 | A Pythagorean-Type Fuzzy Deep Denoising Autoencoder for Industrial Accident Early WarningabstractEarly warning is crucial for preventing industrial accidents and mitigating damage, but current methods are often time-consuming, error-prone, and incompetent to deal with uncertainty. This paper presents a fuzzy deep neural network for early warning of industrial accidents, which equips the classical deep denoising autoencoder (DDAE) model with Pythagorean-type fuzzy parameters in order to enhance the model's representation ability and robustness. To efficiently train the fuzzy deep model, we propose a hybrid algorithm combining Hessian-free optimization and biogeography-based optimization metaheuristic to balance global search and local search. Experiments on datasets from several industrial zones in China show that the proposed Pythagorean-type fuzzy DDAE (PFDDAE) can achieve much higher accuracy of accident risk classification than the classical DDAE and the fuzzy DDAE using regular fuzzy parameters, and the proposed hybrid learning algorithm exhibits significant performance advantage over some other learning algorithms in training PFDDAE. In particular, a test on the 2014 Kunshan aluminum dust explosion accident shows that the deep learning model would be very likely to prevent the accident if it was adopted in advance. Yujun Zheng 0001, Shengyong Chen, Yu Xue 0003, Jin-Yun Xue |
IEEE Trans. Fuzzy Syst. | 2 |
| 2017 | Iterative Re-Constrained Group Sparse Face Recognition With Adaptive Weights LearningabstractIn this paper, we consider the robust face recognition problem via iterative re-constrained group sparse classifier (IRGSC) with adaptive weights learning. Specifically, we propose a group sparse representation classification (GSRC) approach in which weighted features and groups are collaboratively adopted to encode more structure information and discriminative information than other regression based methods. In addition, we derive an efficient algorithm to optimize the proposed objective function, and theoretically prove the convergence. There are several appealing aspects associated with IRGSC. First, adaptively learned weights can be seamlessly incorporated into the GSRC framework. This integrates the locality structure of the data and validity information of the features into l2,p-norm regularization to form a unified formulation. Second, IRGSC is very flexible to different size of training set as well as feature dimension thanks to the l2,p-norm regularization. Third, the derived solution is proved to be a stationary point (globally optimal if p ≥ 1). Comprehensive experiments on representative data sets demonstrate that IRGSC is a robust discriminative classifier which significantly improves the performance and efficiency compared with the state-of-the-art methods in dealing with face occlusion, corruption, and illumination changes, and so on. Jianwei Zheng 0001, Shengyong Chen, Guojiang Shen, Wanliang Wang |
IEEE Trans. Image Process. | 3 |
| 2017 | Airline Passenger Profiling Based on Fuzzy Deep Machine LearningabstractPassenger profiling plays a vital part of commercial aviation security, but classical methods become very inefficient in handling the rapidly increasing amounts of electronic records. This paper proposes a deep learning approach to passenger profiling. The center of our approach is a Pythagorean fuzzy deep Boltzmann machine (PFDBM), whose parameters are expressed by Pythagorean fuzzy numbers such that each neuron can learn how a feature affects the production of the correct output from both the positive and negative sides. We propose a hybrid algorithm combining a gradient-based method and an evolutionary algorithm for training the PFDBM. Based on the novel learning model, we develop a deep neural network (DNN) for classifying normal passengers and potential attackers, and further develop an integrated DNN for identifying group attackers whose individual features are insufficient to reveal the abnormality. Experiments on data sets from Air China show that our approach provides much higher learning ability and classification accuracy than existing profilers. It is expected that the fuzzy deep learning approach can be adapted for a variety of complex pattern analysis tasks. Yujun Zheng 0001, Weiguo Sheng 0001, Xing-Ming Sun, Shengyong Chen |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2016 | A pipeline using multi-layer Tumors Automata for interactive multi-label image segmentationabstractIn this paper, we investigate a novel algorithm to the problem of interactive image segmentation. We propose an extension of the Growcut framework using the Tumors Automata (TA) formed from the superpixel. The proposed TA is similar to Cellular Automata but can directly deal with superpixel. The superpixels (image segments) can provide powerful boundary cues to guide segmentation, where superpixels can be collected easily by over-segmenting the image using any reasonable existing segmentation algorithms. Given a small number of user-labelled superpixels, the rest of the image is segmented automatically by a TA. When the automaton labels the image, the segmentation evolution is faster than Growcut because of the iterative process. Moreover, a level set method and multi-layer TA are employed to further improve the performance. Experiments conducted on the Berkeley Segmentation Database demonstrate the superior performance of our method over the state-of-the-art methods. Sixian Chan 0001, Xiaolong Zhou 0001, Zhuo Zhang 0012, Shengyong Chen |
HSI | 4 |
| 2016 | Distributed diffusion nonnegative LMS algorithm over sensor networksabstractSince most distributed estimation algorithms only try to achieve high estimation precision while ignoring the positive-negative problem of components in the true parameter, estimation using these methods may be physically absurd and uninterpretable. In order to avoid erroneous results, we need to add a nonnegative constraint on the parameter to be estimated. In this paper, we propose a novel distributed diffusion nonnegative LMS algorithm with regularization for estimating some specific parameter. The algorithm keeps the non-negativity of all components in the parameter in the adaptation process. Simulations results illustrate the advantage of our algorithm in the low steady MSD level and high convergence rate. Wei Huang 0015, Yuzhu Ji, Shengyong Chen |
INDIN | 4 |
| 2016 | Objectness ranking by uniform Bayesian model with multimodal and global cuesabstractCategory‐independent object detection and localisation plays an important role in many computer vision tasks. In this study, an efficient method is proposed for generic objectness ranking by fusing two dimension (2D) or 3D information. A novel Bayesian model is designed to integrate multimodal cues and global cues to estimate object location, scale and number. In the pure trichannel colour space, the authors employ global spatial information as new global cues. From the colour+depth (red, green and blue+D) aspect, the authors compute multimodal saliency and oversegments to find two new multimodal cues. Local and regional depth cues are also explored and combined with them together so that a reliable objectness ranking scheme can be implemented. The proposed method is evaluated on web‐public common 2D and RGB+D datasets. In RGB+D cases, the experimental results show that the proposed method achieves an average 5% improvement over state‐of‐the‐art methods. Furthermore, for achieving the similar recall rates, the authors’ method only needs 30% amounts of sampled windows with respect of other available methods. Jianhua Zhang 0002, Junhao Xiao 0001, Shengyong Chen, Jianwei Zhang 0001 |
IET Comput. Vis. | 4 |
| 2016 | Diffusion LMS with component-wise variable step-size over sensor networksabstractIn this study, the authors propose a novel component ‐ wise variable step‐size (CVSS) diffusion distributed algorithm for estimating a specific parameter over sensor networks. The novelty of the CVSS algorithm is that step‐sizes vary from each other on different components at each iteration. They derive the steady‐state value of global mean‐square deviation (MSD) and relative MSD (RMSD). In the numerical simulations, they compare the proposed CVSS algorithm with several other least mean square (LMS) algorithms. Results show that, when compared with these other algorithms, the CVSS algorithm can effectively reduce steady‐state value and speed up convergence rate of RMSD while not sacrificing the convergence rate of MSD. Results also reveal that the proposed CVSS algorithm can achieve reduced difference of steady‐state values of relative estimation error on various components. Wei Huang 0015, Xi Yang 0006, Duanyang Liu, Shengyong Chen |
IET Signal Process. | 4 |
| 2016 | Adaptive Multisubpopulation Competition and Multiniche Crowding-Based Memetic Algorithm for Automatic Data ClusteringabstractAutomatic data clustering, whose goal is to recover the proper number of clusters as well as appropriate partitioning of data sets, is a fundamental yet challenging problem in unsupervised learning. In this paper, adaptive multisubpopulation competition (AMC) and multiniche crowding are proposed and incorporated into a memetic algorithm to tackle the problem. The AMC mechanism is developed to ensure a diverse search over solution subspaces corresponding to different numbers of clusters while allowing more promising subspaces to be more intensively searched. In this mechanism, the amount of individuals to be migrated between subpopulations is adaptively controlled according to the performance of subpopulations as well as the diversity of cluster numbers in population. Further, the migration is restricted to occur between subpopulations with relatively similar performances. Additionally, subpopulations with different performances are devised to search their corresponding subspaces with different exploration powers. The adaptive multiniche crowding scheme is designed to promote a diverse search of the subspace while allowing an efficient convergence of the corresponding subpopulation. This is achieved by dynamically adjusting parameter values of a multiniche crowding method to form and maintain diverged niches of high fitness within the subpopulation. The performance of proposed algorithm has been demonstrated through a series of experiments on both artificial and real data, and compared with existing methods. The results reveal that our proposed algorithm can achieve superior clustering performance and outperform related methods. Weiguo Sheng 0001, Shengyong Chen, Mengmeng Sheng, Gang Xiao 0001, Jiafa Mao, Yujun Zheng 0001 |
IEEE Trans. Evol. Comput. | 2 |
| 2015 | Observer-based consensus tracking for second-order leader-following nonlinear multi-agent systems with adaptive coupling parameter design
Xiaole Xu, Shengyong Chen, Lixin Gao 0004 |
Neurocomputing | 2 |
| 2015 | A hybrid fireworks optimization method with differential evolution operators
Yujun Zheng 0001, Xinli Xu, Haifeng Ling, Shengyong Chen |
Neurocomputing | 4 |
| 2015 | A Hybrid Neuro-Fuzzy Network Based on Differential Biogeography-Based Optimization for Online Population Classification in EarthquakesabstractTimely and accurate identification and classification of victims in earthquakes is crucial for improving rescue efficiency, but available information about victims and their surrounding environment is often vague and imprecise. Rescue wings is a web-based intelligent system that monitors and analyzes the statuses of identified victims to support decision making in earthquake rescue operations. A key component of the system is a Takagi-Sugeno (T-S)-type neuro-fuzzy network for disaster-stricken population classification, and one important input of the network is the output of another T-S-type recurrent neuro-fuzzy network for recognizing the movement patterns from the users' temporal location data. A novel differential biogeography-based optimization (DBBO) algorithm is developed for parameter optimization of both the main network and the subnetwork. Experimental results have shown that the hybrid neuro-fuzzy network exhibits good classification performance in comparison with some other typical neuro-fuzzy networks, and the proposed DBBO outperforms some state-of-the-art evolutionary algorithms in network learning. The solution approach has also been successfully applied to the 2013 Ya'an Earthquake in Sichuan province, China. Yujun Zheng 0001, Haifeng Ling, Shengyong Chen, Jin-Yun Xue |
IEEE Trans. Fuzzy Syst. | 3 |
| 2015 | Emergency Railway Transportation Planning Using a Hyper-Heuristic ApproachabstractThe railway has played a significant role in disaster relief transportation in China. This paper presents an emergency railway transportation problem, which is to use the limited transport capability to meet the urgent relief transportation requirements. We have applied several state-of-the-art evolutionary algorithms to a variety of problem instances, but have found that none of them can obtain satisfactory solutions in all cases. To overcome this obstacle, we integrate a set of individual heuristic operators into a hyperheuristic framework, which performs a stochastic search on the low-level heuristics by using feedback on their performance in the process of problem solving, thus yielding a high overall performance on different instances. Computational experiments show that the hyperheuristic exhibits significant advantages over the individual heuristics. The problem model and the hyperheuristic solution approach have also been successfully applied to the emergency railway transportation during the 2013 Dingxi earthquake, in China. Yujun Zheng 0001, Min-Xia Zhang, Haifeng Ling, Shengyong Chen |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2015 | A Biometric Key Generation Method Based on Semisupervised Data ClusteringabstractStoring biometric templates and/or encryption keys, as adopted in traditional biometrics-based authentication methods, has raised a matter of serious concern. To address such a concern, biometric key generation, which derives encryption keys directly from statistical features of biometric data, has emerged to be a promising approach. Existing methods of this approach, however, are generally unable to appropriately model user variations, making them difficult to produce consistent and discriminative keys of high entropy for authentication purposes. This paper develops a semisupervised clustering scheme, which is optimized through a niching memetic algorithm, to effectively and simultaneously model both intra- and interuser variations. The developed scheme is employed to model the user variations on both single features and feature subsets with the purpose of recovering a large number of consistent and discriminative feature elements for key generation. Moreover, the scheme is designed to output a large number of clusters, thus further assisting in producing long while consistent and discriminative keys. Based on this scheme, a biometric key generation method is finally proposed. The performance of the proposed method has been evaluated on the biometric modality of handwritten signatures and compared with existing methods. The results show that our method can deliver consistent and discriminative keys of high entropy, outperforming-related methods. Weiguo Sheng 0001, Shengyong Chen, Gang Xiao 0001, Jiafa Mao, Yujun Zheng 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 2 |
| 2015 | Efficient kernel discriminative common vectors for classification
Jianwei Zheng 0001, Qiongfang Huang, Shengyong Chen, Wanliang Wang |
Vis. Comput. | 3 |
| 2014 | Incremental min-max projection analysis for classification
Jianwei Zheng 0001, Shengyong Chen, Wanliang Wang |
Neurocomputing | 3 |
| 2014 | Multilocal Search and Adaptive Niching Based Memetic Algorithm With a Consensus Criterion for Data ClusteringabstractClustering is deemed one of the most difficult and challenging problems in machine learning. In this paper, we propose a multilocal search and adaptive niching-based genetic algorithm with a consensus criterion for automatic data clustering. The proposed algorithm employs three local searches of different features in a sophisticated manner to efficiently exploit the decision space. Furthermore, we develop an adaptive niching method, which can dynamically adjust its parameter value depending on the problem instance as well as the search progress, and incorporate it into the proposed algorithm. The adaptation strategy is based on a newly devised population diversity index, which can be used to promote both genetic diversity and fitness. Consequently, diverged niches of high fitness can be formed and maintained in the population, making the approach well-suited to effective exploration of the complex decision space of clustering problems. The resulting algorithm has been used to optimize a consensus clustering criterion, which is suggested with the purpose of achieving reliable solutions. To evaluate the proposed algorithm, we have conducted a series of experiments on both synthetic and real data and compared it with other reported methods. The results show that our proposed algorithm can achieve superior performance, outperforming related methods. Weiguo Sheng 0001, Shengyong Chen, Michael C. Fairhurst, Gang Xiao 0001, Jiafa Mao |
IEEE Trans. Evol. Comput. | 2 |
| 2014 | Population Classification in Fire Evacuation: A Multiobjective Particle Swarm Optimization ApproachabstractIn an emergency evacuation operation, accurate classification of the evacuee population can provide important information to support the responders in decision making; and therefore, makes a great contribution in protecting the population from potential harm. However, real-world data of fire evacuation is often noisy, incomplete, and inconsistent, and the response time of population classification is very limited. In this paper, we propose an effective multiobjective particle swarm optimization method for population classification in fire evacuation operations, which simultaneously optimizes the precision and recall measures of the classification rules. We design an effective approach for encoding classification rules, and use a comprehensive learning strategy for evolving particles and maintaining diversity of the swarm. Comparative experiments show that the proposed method performs better than some state-of-the-art methods for classification rule mining, especially on the real-world fire evacuation dataset. This paper also reports a successful application of our method in a real-world fire evacuation operation that recently occurred in China. The method can be easily extended to many other multiobjective rule mining problems. Yujun Zheng 0001, Haifeng Ling, Jin-Yun Xue, Shengyong Chen |
IEEE Trans. Evol. Comput. | 4 |
| 2013 | Self-Adaptive Matching In Local Windows For Depth EstimationabstractThis paper proposes a novel local stereo matching approach based on self-adapting matching window. We improve the accuracy of stereo matching in 3 steps. First, we integrate shape and size information, and construct robust minimum matching windows by applying a self-adapting method. Then, two matching cost optimization strategies are employed for handling both occlusion regions and image borders. Last, we perform a refinement algorithm for obtaining more accurate depth map. Experiment results on the Middlebury stereo image pairs prove that the proposed matching method performs equally well in comparison with other state-of-the-art local approaches. Haiqiang Jin, Sheng Liu 0002, Xuhua Yang 0001, Shengyong Chen |
ECMS | 4 |
| 2013 | Image Super-Resolution Reconstruction Using Map EstimationabstractThis paper presents a promising super-resolution (SR) approach using maximum a posteriori (MAP) estimation. We consider the high resolution (HR) estimation as a Markov Random Field (MRF), using a transformed gradient field prior to repair the image fuzzy problem caused by MRF. An improved Normalized Convolution method is proposed to obtain a first good estimation. We build a reasonable energy function and minimize the posterior energy by gradient descent algorithm. Experimental results on realistic image sequence and comparisons with several other SR techniques show that our approach gives the best results both qualitative and quantitative. Xin-Long Lu, Shengyong Chen, Xin Wang 0204, Sheng Liu 0002, Chunyan Yao, Xianping Huan |
ECMS | 2 |
| 2013 | Isogeometric Analysis For Dynamic Model SimulationabstractThis paper proposes a method of constructing a dynamic model of a ventricle based on isogeometric simulation so as to diagnose cardiac disease more accurately. Isogeometric simulation is an accurate simulation technology based on NURBS, which has evolved into an essential tool for a semi-analytical representation of geometric entities. Especially, a new method of moving control points is used to achieve a dynamic model of the ventricle. This method promises the model to be very accurate, efficient, and successive, in comparison with traditional models. Furthermore, the paper also puts forward a new error estimation method, which adopts the vector norm to get an overall analysis of the error coefficient in each direction. The error estimation method avoids a complicated estimation for each knot. It can not only be used to evaluate the value of the error accurately, but also reduce the local error by adjusting the control points. Moreover, the proposed NURBS model can especially be useful to analyze the motion and dynamics of the heart, and it is important for doctors to find early cues of cardiac diseases. Huabin Yin, Qiu Guan, Shengyong Chen |
ECMS | 3 |
| 2013 | Cooperative particle swarm optimization for multiobjective transportation planning
Yujun Zheng 0001, Shengyong Chen |
Appl. Intell. | 2 |
| 2013 | Simultaneous image color correction and enhancement using particle swarm optimization
Ngai Ming Kwok, Haiyan Shi, Quang Phuc Ha, Gu Fang 0001, Shengyong Chen, Xiuping Jia |
Eng. Appl. Artif. Intell. | 5 |
| 2013 | Dimension estimation of image manifolds by minimal cover approximation
Mingyu Fan, Xiaoqin Zhang 0002, Shengyong Chen, Hujun Bao, Stephen J. Maybank |
Neurocomputing | 3 |
| 2013 | Leader-following consensus of discrete-time multi-agent systems with observer-based protocols
Xiaole Xu, Shengyong Chen, Wei Huang 0015, Lixin Gao 0004 |
Neurocomputing | 2 |
| 2013 | PABM-EDCF: parameter adaptive bi-directional mapping mechanism for video transmission over WSNs
Xin-Wei Yao 0001, Wanliang Wang, Shuang-Hua Yang, Shengyong Chen |
Multim. Tools Appl. | 4 |
| 2013 | Finding Optimal Focusing Distance and Edge Blur Distribution for Weakly Calibrated 3-D Visionabstract3-D Vision is now a common sensing method frequently used in industrial applications. With the convenience of an uncalibrated system, 3-D reconstruction by a self-calibration technique is possible, but always incomplete or unreliable. This paper presents a novel method to analyze the blur distribution in an image and find the optimal focusing distance so that additional constraints can be used to generate absolute measurement of the models. With the assumption of a Gaussian distribution model of the point spread function, this paper applies two theorems to efficiently compute the defocusing extent on stripe edges. Because the blurring diameter implies the distance from the sensor to the surface, we can upgrade the 3-D map obtained from self-calibration with the known scaling factor. Through theoretical and experimental analysis, we find that not only the technology is feasible, but also both the accuracy and the efficiency are satisfactory. Shengyong Chen, Youfu Li 0001 |
IEEE Trans. Ind. Informatics | 1 |
| 2013 | Intelligent Video Systems and Analytics: A SurveyabstractRecent technology and market trends have demanded the significant need for feasible solutions to video/camera systems and analytics. This paper provides a comprehensive account on theory and application of intelligent video systems and analytics. It highlights the video system architectures, tasks, and related analytic methods. It clearly demonstrates that the importance of the role that intelligent video systems and analytics play can be found in a variety of domains such as transportation and surveillance. Research directions are outlined with a focus on what is essential to achieve the goals of intelligent video systems and analytics. Honghai Liu 0001, Shengyong Chen, Naoyuki Kubota |
IEEE Trans. Ind. Informatics | 2 |
| 2013 | Discover Novel Visual Categories From Dynamic Hierarchies Using Multimodal AttributesabstractLearning novel visual categories from observations and experiences in unexplored environment is a vitally important cognitive ability for human beings. A dynamic category hierarchy that is an inherent structure in a human mind is a key component for this ability. This paper develops a framework to build dynamic category hierarchy based on object attributes and a topic model. Since humans trend to utilize multimodal information to learn novel categories, we also develop an algorithm to learn multimodal object attributes from multimodal data. The new multimodal attributes can describe objects efficiently and can generalize from learned categories to novel ones. By comparison with a state-of-the-art unimodal attribute, the multimodal attributes can achieve 4%-19% improvements on average. We also develop a constrained topic model, which can accurately construct category hierarchies for large-scale categories. Based on them, the novel framework can effectively detect novel categories and relate them with known categories for further category learning. Extensive experiments are conducted using a public multimodal dataset, i.e., color and point cloud data, to evaluate the multimodal attributes and the dynamic category hierarchy. The experimental results show the effectiveness of multimodal attributes to describe objects and the satisfactory performance of the dynamic category hierarchy to discover novel categories. By comparison with state-of-the-art methods, the dynamic category hierarchy achieves 7% improvements. Jianhua Zhang 0002, Jianwei Zhang 0001, Shengyong Chen |
IEEE Trans. Ind. Informatics | 3 |
| 2012 | Constructing dynamic category hierarchies for novel visual category discoveryabstractCategory hierarchies are commonly used to compactly represent large numbers of categories and reduce the complexity of the classification problem. In this paper we introduce a novel and extended application of category hierarchies which is a powerful novel framework developed to construct dynamic category hierarchies and automatically discover novel visual categories. The dynamic is a characteristic of category hierarchies which can facilitate an important cognitive ability, the discovering of novel categories. We develop a constrained hierarchical latent Dirichlet allocation to build accurate category hierarchies. We employ object attributes as features to describe objects, which can transfer knowledge across categories and can efficiently describe novel categories. By combining them in the novel framework, novel visual object categories can be efficiently discovered and described. Extensive experiments based on PASCAL VOC 2008 and the LabelMe image database show the satisfactory performance of the proposed framework. Jianhua Zhang 0002, Jianwei Zhang 0001, Shengyong Chen, Ying Hu 0001, Haojun Guan |
IROS | 3 |
| 2012 | Reliable and secure encryption key generation from fingerprintsabstractPurpose Biometric authentication, which requires storage of biometric templates and/or encryption keys, raises a matter of serious concern, since the compromise of templates or keys necessarily compromises the information secured by those keys. To address such concerns, efforts based on dynamic key generation directly from the biometrics have recently emerged. However, previous methods often have quite unacceptable authentication performance and/or small key spaces and therefore are not viable in practice. The purpose of this paper is to propose a novel method which can reliably generate long keys while requires storage of neither biometric templates nor encryption keys. Design/methodology/approach This proposition is achieved by devising the use of fingerprint orientation fields for key generation. Additionally, the keys produced are not permanently linked to the orientation fields, hence, allowing them to be replaced in the event of key compromise. Findings The evaluation demonstrates that the proposed method for dynamic key generation can offer both good reliability and security in practice, and outperforms other related methods. Originality/value In this paper, the authors propose a novel method which can reliably generate long keys while requires storage of neither biometric templates nor encryption keys. This is achieved by devising the use of fingerprint orientation fields for key generation. Additionally, the keys produced are not permanently linked to the orientation fields, hence, allowing them to be replaced in the event of key compromise. Weiguo Sheng 0001, Gareth Howells 0001, Michael C. Fairhurst, Farzin Deravi, Shengyong Chen |
Inf. Manag. Comput. Secur. | 5 |
| 2012 | Super-resolution in practice: the complete pipeline from image capture to super-resolved subimage creation using a novel frame selection method
Maria Petrou, Mohamed Hisham Jaward, Shengyong Chen, Mark Briers |
Mach. Vis. Appl. | 3 |
| 2012 | Acceleration Strategies in Generalized Belief PropagationabstractGeneralized belief propagation is a popular algorithm to perform inference on large-scale Markov random fields (MRFs) networks. This paper proposes the method of accelerated generalized belief propagation with three strategies to reduce the computational effort. First, a min-sum messaging scheme and a caching technique are used to improve the accessibility. Second, a direction set method is used to reduce the complexity of computing clique messages from quartic to cubic. Finally, a coarse-to-fine hierarchical state-space reduction method is presented to decrease redundant states. The results show that a combination of these strategies can greatly accelerate the inference process in large-scale MRFs. For common stereo matching, it results in a speed-up of about 200 times. Shengyong Chen, Zhongjie Wang 0001 |
IEEE Trans. Ind. Informatics | 1 |
| 2012 | A Hierarchical Model Incorporating Segmented Regions and Pixel Descriptors for Video Background SubtractionabstractBackground subtraction is important for detecting moving objects in videos. Currently, there are many approaches to performing background subtraction. However, they usually neglect the fact that the background images consist of different objects whose conditions may change frequently. In this paper, a novel hierarchical background model is proposed based on segmented background images. It first segments the background images into several regions by the mean-shift algorithm. Then, a hierarchical model, which consists of the region models and pixel models, is created. The region model is a kind of approximate Gaussian mixture model extracted from the histogram of a specific region. The pixel model is based on the cooccurrence of image variations described by histograms of oriented gradients of pixels in each region. Benefiting from the background segmentation, the region models and pixel models corresponding to different regions can be set to different parameters. The pixel descriptors are calculated only from neighboring pixels belonging to the same object. The experimental results are carried out with a video database to demonstrate the effectiveness, which is applied to both static and dynamic scenes by comparing it with some well-known background subtraction methods. Shengyong Chen, Jianhua Zhang 0002, Youfu Li 0001, Jianwei Zhang 0001 |
IEEE Trans. Ind. Informatics | 1 |
| 2012 | Guest Editorial Special Section on Intelligent Video Systems and AnalyticsabstractThe 11 papers in this special section focus on intelligent video systems and analytics. Honghai Liu 0001, Shengyong Chen, Naoyuki Kubota |
IEEE Trans. Ind. Informatics | 2 |
| 2012 | Elman Fuzzy Adaptive Control for Obstacle Avoidance of Mobile Robots Using Hybrid Force/Position IncorporationabstractThis paper addresses a virtual force field between mobile robots and obstacles to keep them away with a desired distance. An online learning method of hybrid force/position control is proposed for obstacle avoidance in a robot environment. An Elman neural network is proposed to compensate the effect of uncertainties between the dynamic robot model and the obstacles. Moreover, this paper uses an Elman fuzzy adaptive controller to adjust the exact distance between the robot and the obstacles. The effectiveness of the proposed method is demonstrated by simulation examples. Shuhuan Wen, Wei Zheng 0005, Jinghai Zhu, Xiaoli Li 0002, Shengyong Chen |
IEEE Trans. Syst. Man Cybern. Part C | 5 |
| 2011 | Integrate multi-modal cues for category-independent object detection and localizationabstractTo detect and localize objects is an indispensable step for many computer vision tasks. Most of the state-of-the-art methods of object detection and localization are category-dependent. These methods can achieve a significant performance. However, they are useless for detecting and localizing objects belonging to an unknown category when applying them to an unknown environment. In this paper, a method is proposed for detecting and localizing generic objects without specifying their categories. The proposed method combines diverse cues, including multi-scale saliency, superpixels straddling, intensity, depth and global information, into a uniform Bayesian framework to obtain accurate detection and localization. By comparison to state-of-the-art methods, our experiments show the promising performance of the proposed method based on the PASCAL VOC 08 dataset and our indoor scene dataset. Jianhua Zhang 0002, Junhao Xiao 0001, Jianwei Zhang 0001, Houxiang Zhang, Shengyong Chen |
IROS | 5 |
| 2008 | Vision Processing for Realtime 3-D Data Acquisition Based on Coded Structured LightabstractStructured light vision systems have been successfully used for accurate measurement of 3-D surfaces in computer vision. However, their applications are mainly limited to scanning stationary objects so far since tens of images have to be captured for recovering one 3-D scene. This paper presents an idea for real-time acquisition of 3-D surface data by a specially coded vision system. To achieve 3-D measurement for a dynamic scene, the data acquisition must be performed with only a single image. A principle of uniquely color-encoded pattern projection is proposed to design a color matrix for improving the reconstruction efficiency. The matrix is produced by a special code sequence and a number of state transitions. A color projector is controlled by a computer to generate the desired color patterns in the scene. The unique indexing of the light codes is crucial here for color projection since it is essential that each light grid be uniquely identified by incorporating local neighborhoods so that 3-D reconstruction can be performed with only local analysis of a single image. A scheme is presented to describe such a vision processing method for fast 3-D data acquisition. Practical experimental performance is provided to analyze the efficiency of the proposed methods. Shengyong Chen, Youfu Li 0001, Jianwei Zhang 0001 |
IEEE Trans. Image Process. | 1 |
| 2007 | Realtime Structured Light Vision with the Principle of Unique Color CodesabstractTo date, several successful structured light vision systems for accurate 3D measurement in machine vision have been set up. However, these are usually limited to scanning stationary objects or static environments since tens of images have to be captured for recovering one 3D scene, which results in the industry largely avoiding this technology. This paper presents a method of grid-pattern design based on the principles of uniquely color-encoded structured light, to improve the reconstruction efficiency for real-time processing. For a live scene, the 3D measurement is desired to only capture a single image. To realize this, an important problem for the color-encoded projection is the unique indexing of the color codes in the image. It is essential that each light grid be uniquely identified by incorporating the local neighborhoods in the pattern so that 3D reconstruction can be performed with only local analysis of a single image. This paper describes such a method in the design of the special grid patterns and its corresponding 3D reconstruction method for fast vision perception. Shengyong Chen, Youfu Li 0001, Jianwei Zhang 0001 |
ICRA | 1 |
| 2007 | Active Illumination for Robot VisionabstractA vision sensor is the robot's eye to perceive its environment, but the perception performance can be significantly affected by illumination conditions. This paper presents strategies of adaptive illumination control for robot vision to achieve the best scene interpretation. It investigates how to obtain the most comfortable illumination conditions for a vision sensor. In a "comfort" condition the image reflects the natural properties of the concerned object. "Discomfort" may occur if some scene information is lost. Strategies are proposed to optimize the pose and optical parameters of the luminaire and the sensor, with emphasis on controlling the intensity and avoiding glare. Shengyong Chen, Jianwei Zhang 0001, Houxiang Zhang, Wanliang Wang, Youfu Li 0001 |
ICRA | 1 |
| 2007 | Runtime reconfiguration of a modular mobile robot with serial and parallel mechanismsabstractThis paper presents a novel field robot JL-I based on a reconfigurable concept for urban search and rescue applications. The robot consists of three identical modules; each module is an entire robotic system that can perform distributed activities. It features three-degrees-of-freedom (DOF) active joints actuated by serial and parallel mechanisms for changing shape and flexible docking mechanism. The docking mechanism enables adjacent modules to connect or disconnect flexibly and automatically. DOF analysis, working space analysis and the kinematics of the 3D active joint between connected modules are studied thoroughly. In the end a series of successful tests confirm the principles and the robot's capabilities. Houxiang Zhang, Shengyong Chen, Wanliang Wang, Jianwei Zhang 0001, Guanghua Zong |
IROS | 2 |
| 2006 | A Focal Cue for Metric Measurement of 3D SurfacesabstractThis paper finds a method for computing the best-focused location from an image and using it as a dimensional cue for acquisition of a 3D scene surface. In some situations in 3D vision, an object cannot be reconstructed into a 3D model with metric dimensions. Rather, it can only be reconstructed into a 3D structure up to a similarity transformation. To upgrade the 3D model from a similarity transformation to a Euclidean transformation, we propose a method based on the best-focused locations. By analyzing the blur distribution in an image, this method finds the best-focused locations from an image, which provides an additional cue for upgrading the reconstructed 3D structure. Hence, we can obtain not only the object's shape, but also the dimensions and sizes of surface features Shengyong Chen, Youfu Li 0001, Jianwei Zhang 0001 |
IROS | 1 |
| 2005 | Transient Chaotic Discrete Neural Network for Flexible Job-Shop Scheduling
Xinli Xu, Qiu Guan, Wanliang Wang, Shengyong Chen |
ISNN (1) | 4 |
| 2005 | A Visual Automatic Incident Detection Method on Freeway Based on RBF and SOFM Neural Networks
Xuhua Yang 0001, Qiu Guan, Wanliang Wang, Shengyong Chen |
ISNN (3) | 4 |
| 2005 | Vision sensor planning for 3-D model acquisitionabstractA novel method is proposed in this paper for automatic acquisition of three-dimensional (3-D) models of unknown objects by an active vision system, in which the vision sensor is to be moved from one viewpoint to the next around the target to obtain its complete model. In each step, sensing parameters are determined automatically for incrementally building the 3-D target models. The method is developed by analyzing the target's trend surface, which is the regional feature of a surface for describing the global tendency of change. While previous approaches to trend analysis are usually focused on generating polynomial equations for interpreting regression surfaces in three dimensions, this paper proposes a new mathematical model for predicting the unknown area of the object surface. A uniform surface model is established by analyzing the surface curvatures. Furthermore, a criterion is defined to determine the exploration direction, and an algorithm is developed for determining the parameters of the next view. Implementation of the method is carried out to validate the proposed method. Shengyong Chen, Youfu Li 0001 |
IEEE Trans. Syst. Man Cybern. Part B | 1 |
| 2004 | Active Viewpoint Planning for Model ConstructionabstractThis paper presents a novel method of viewpoint planning for incrementally building the models of unknown objects or environments by an active vision system. The proposed method is based on the model of trend surface, which is the regional feature of a surface for describing the global tendency of change. A new mathematical model is developed for predicting the unknown area of the object surface. A unique surface model is established by analyzing the surface curvature. Furthermore, a criterion is defined to determine the exploration direction. The algorithm is developed for determining the next view pose, which satisfies the placement constraints such as resolution, focus, and field of view. Finally, implementation of the method is carried out to verify the proposed method. Shengyong Chen, Youfu Li 0001 |
ICRA | 1 |
| 2004 | Automatic sensor placement for model-based robot visionabstractThis paper presents a method for automatic sensor placement for model-based robot vision. In such a vision system, the sensor often needs to be moved from one pose to another around the object to observe all features of interest. This allows multiple three-dimensional (3-D) images to be taken from different vantage viewpoints. The task involves determination of the optimal sensor placements and a shortest path through these viewpoints. During the sensor planning, object features are resampled as individual points attached with surface normals. The optimal sensor placement graph is achieved by a genetic algorithm in which a min-max criterion is used for the evaluation. A shortest path is determined by Christofides algorithm. A Viewpoint Planner is developed to generate the sensor placement plan. It includes many functions, such as 3-D animation of the object geometry, sensor specification, initialization of the viewpoint number and their distribution, viewpoint evolution, shortest path computation, scene simulation of a specific viewpoint, parameter amendment. Experiments are also carried out on a real robot vision system to demonstrate the effectiveness of the proposed method. Shengyong Chen, Y. F. Li |
IEEE Trans. Syst. Man Cybern. Part B | 1 |
| 2003 | Dynamically reconfigurable visual sensing for 3D perceptionabstractIn many applications, a vision sensor often needs to move from one place to another and change its configuration for perception of different object features. A dynamic reconfigurable vision sensor is useful in such a case to gaze at the features. This paper introduces this concept and investigates the issues in self-recalibrating a 6-DOF structured light system under changing sensing configuration. The relative pose between the projector and camera of the system is calibrated by taking a single view of the scene, so that the 3D measurements and reconstruction can be performed immediately when and if the configuration of the system is changed. Experiments were carried out to demonstrate the implementation of the proposed method. Shengyong Chen, Youfu Li 0001 |
ICRA | 1 |
| 2003 | Automatic recalibration of an active structured light vision systemabstractA structured light vision system using pattern projection is useful for robust reconstruction of three-dimensional objects. One of the major tasks in using such a system is the calibration of the sensing system. This paper presents a new method by which a two-degree-of-freedom structured light system can be automatically recalibrated, if and when the relative pose between the camera and the projector is changed. A distinct advantage of this method is that neither an accurately designed calibration device nor the prior knowledge of the motion of the camera or the scene is required. Several important cues for self-recalibration are explored. The sensitivity analysis shows that high accuracy in-depth value can be achieved with this calibration method. Some experimental results are presented to demonstrate the calibration technique. Youfu Li 0001, Shengyong Chen |
IEEE Trans. Robotics Autom. | 2 |
| 2002 | Optimum viewpoint planning for model-based robot visionabstractIn some model-based vision tasks, such as automatic inspection of industrial parts, a set of viewpoints must be planned for sampling all features of interest around the object. This paper presents the techniques of deciding the optimal viewpoint distribution and a shortest path through these viewpoints, which are achieved by the genetic algorithm and Christofides algorithm respectively. Shengyong Chen, Youfu Li 0001 |
IEEE Congress on Evolutionary Computation | 1 |
| 2002 | Self Recalibration of a Structured Light Vision System from a Single ViewabstractStructured-light system is widely used for reconstructing 3D objects in machine vision. One of the major tasks in establishing such a system is the laborious and tedious calibration of the sensors. This paper presents a new method which dynamically calibrates the system automatically, if and when the relative pose between the camera and the projector is changed. A distinct advantage of this method is that neither the design of a calibration pattern/device nor the pre-knowledge of the movement of camera or scene is required. Several important cues for self-recalibration, including geometrical cue and focus cue, are explored in this paper Finally, some experimental observations are presented to illustrate the implementation of this new method. Shengyong Chen, Youfu Li 0001 |
ICRA | 1 |
| 2002 | A Method of Automatic Sensor Placement for Robot Vision in Inspection TasksabstractThis paper presents an automatic sensor placement technique for robot vision in inspection tasks. In such vision systems, a sensor often needs to be moved from one pose to another around the object to sample all features of interest. Multiple 3D images are taken from different vantage points. The technique involves deciding the optimal sensor placements and a shortest path through these viewpoints for automatic generation of an inspection plan. A viewpoint is expressed by N parameters and a topology of viewpoints is achieved by genetic algorithm. The inspection plan is evaluated using a min-max criterion and the shortest path is determined by Christofides algorithm. In addition, a computation example is presented to illustrate the techniques and algorithms. Shengyong Chen, Youfu Li 0001 |
ICRA | 1 |