EDBT 2026 Demo / reviewers in the wild / expert
Fan Shi 0001
dblp:96/8708-1
· DBLP profile ↗
44ranked-venue papers
0as first author
43since 2021 · last 2026
0000-0003-2074-0228ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 16 · 15 since 2021Artificial intelligence and machine learning · 11 · 10 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 9 since 2021Systems, architecture and hardware · 5 · 5 since 2021Computer networks · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A2P-Net: Asymmetric Domain-Adaptive Prototype Network for Cross-Domain Multimodal Sensor RetrievalabstractMultimodal sensor data from inertial measurement units (IMUs), including accelerometers, gyroscopes, and inclinometers, encode environmental conditions that are difficult to capture through images or text. In maritime settings, classifying sea states from such sensor streams is important for autonomous navigation but faces two interacting difficulties. First, training relies heavily on synthetic simulation data whose idealized physics diverge from real ocean measurements, creating a domain gap. Second, extreme sea states are rare in both domains, and the resulting class imbalance is compounded by the much larger volume of synthetic samples, which together skew gradient updates away from the scarce but operationally important real-world tail classes. We propose A2P-Net, an end-to-end framework that tackles both problems jointly. An adaptive heterogeneous encoder with decoupled channel–temporal attention maps variable-dimension sensor inputs into a shared latent space. Domain-adversarial training aligns synthetic and real feature distributions in that space, and prototype-based metric learning builds per-class retrieval anchors while an asymmetric weighting scheme up-weights real-domain samples to correct the optimization bias. On two custom sea state datasets that mix real and simulated ship motion recordings, A2P-Net reaches 98.7% and 98.9% F1 on real-only evaluation, outperforming the strongest baseline by 1.1–1.3 percentage points. It also ranks first on 15 of 30 UEA multivariate time-series benchmarks with an average accuracy of 74.2%. Mengna Liu, Xu Cheng 0003, Fan Shi 0001, Shengyong Chen |
ICMR | 5 |
| 2026 | Temporal-channel decoupled learning for sea state estimation from ship motion data under class imbalance
Mengna Liu, Xu Cheng 0003, Fan Shi 0001, Shengyong Chen |
Eng. Appl. Artif. Intell. | 5 |
| 2026 | Learning invariant representation for light field adversarial salient object detection
Mianzhao Wang, Fan Shi 0001, Xu Cheng 0003, Shengyong Chen |
Eng. Appl. Artif. Intell. | 2 |
| 2026 | Time-Filtering Graph Learning With Spatial-Temporal Diffusion for Robust Blade Icing DetectionabstractAccurate detection of wind turbine blade icing is essential for ensuring the safety and efficiency of wind farm operations. Although current machine learning and deep learning approaches are capable of identifying icing states, they still suffer from three major limitations: heavy reliance on manual feature engineering limits the capture of meaningful spatiotemporal dependencies; highly imbalanced data leads to increased false negatives and alarm delays; and sensitivity to data noise often introduces false positives. To address these challenges, this paper proposes an imbalance- and noise-resistant TG-Diff network, which achieves accurate and robust icing state classification by effectively integrating spatiotemporal information from multiple sensors. The TG-Diff comprises three core modules: a Temporal Filtering-based Graph Learning Module (TF-GLM), a Temporal-Spatial Diffusion Graph Convolutional Network (TSD-GCN), and a Distance-Based Classifier (DBC). Specifically, the TF-GLM dynamically infers node relationships to construct robust graph topologies; the TSD-GCN enhances feature representation and suppresses noise through diffusion mechanisms; and the DBC effectively mitigates class imbalance by leveraging distance-based decision boundaries. Experimental results demonstrate that the proposed modules work synergistically to collectively enhance the accuracy and robustness of icing detection in real-world complex environments. Lingzhu Hu, Xiufeng Liu 0001, Fan Shi 0001, Xu Cheng 0003 |
IEEE Internet Things J. | 5 |
| 2026 | Boundary-aware and multi-angle modeling-based object tracking in polarimetric images
Qiaohui Wang, Fan Shi 0001, Mianzhao Wang, Xinbo Geng, Meng Zhao 0001 |
Knowl. Based Syst. | 2 |
| 2026 | Light field collaborative perception for visual object tracking
Mianzhao Wang, Fan Shi 0001, Xu Cheng 0003, Meng Zhao 0001 |
Pattern Recognit. | 2 |
| 2026 | Pioneering Video Semantic Segmentation With Light Field Imaging and Spatial-Angular-Temporal Fusion
Fan Shi 0001, Xu Cheng 0003, Shengyong Chen |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2026 | Hierarchical Spatial-Angular Representation Learning for Point-Supervised Salient Object Detection in Light FieldsabstractLight Field Salient Object Detection (LFSOD) aims to identify visually distinctive regions by leveraging the complementary spatial–angular information inherent in 4D light field imagery. A major challenge lies in modeling angular dependencies and maintaining spatial coherence under sparse supervision. In this article, we propose a weakly supervised network that consists of three interdependent modules. First, the Light Field Division (LFD) module utilizes epipolar geometry to extract direction-aware boundary features, enhancing the encoding of angular disparities. Second, the Light Field Spatial Association (LFSA) module anchors cross-view feature alignment using central-viewpoint annotations, thereby enforcing spatial consistency and mitigating redundant representations. Third, the Light Field Saliency Local Clustering (LFLC) module introduces a joint boundary-appearance modeling strategy that integrates adaptive clustering with error-aware regularization to refine structural predictions. Experiments on three benchmark datasets show that our method consistently outperforms mainstream weakly supervised approaches. It also achieves superior performance compared to several fully supervised methods. Xinbo Geng, Fan Shi 0001, Xu Cheng 0003, Shengyong Chen |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2025 | SCSegamba: Lightweight Structure-Aware Vision Mamba for Crack Segmentation in StructuresabstractPixel-level segmentation of structural cracks across various scenarios remains a considerable challenge. Current methods encounter challenges in effectively modeling crack morphology and texture, facing challenges in balancing segmentation quality with low computational resource usage. To overcome these limitations, we propose a lightweight Structure-Aware Vision Mamba Network (SCSegamba), capable of generating high-quality pixel-level segmentation maps by leveraging both the morphological information and texture cues of crack pixels with minimal computational cost. Specifically, we developed a StructureAware Visual State Space module (SAVSS), which incorporates a lightweight Gated Bottleneck Convolution (GBC) and a Structure-Aware Scanning Strategy (SASS). The key insight of GBC lies in its effectiveness in modeling the morphological information of cracks, while the SASS enhances the perception of crack topology and texture by strengthening the continuity of semantic information between crack pixels. Experiments on crack benchmark datasets demonstrate that our method outperforms other state-of-the-art (SOTA) methods, achieving the highest performance with only 2.8M parameters. On the multi-scenario dataset, our method reached 0.8390 in F1 score and 0.8479 in mIoU. The code is available at https://github.com/Karl1109/SCSegamba. Fan Shi 0001, Xu Cheng 0003, Shengyong Chen |
CVPR | 3 |
| 2025 | Harnessing Light Field Angular Cues and Spatial Geometries for Semantic Segmentationabstract4D light field imaging captures rich spatial-angular information, providing essential geometric cues for semantic segmentation tasks. In this paper, we introduce a novel backbone network called the Light Field Extraction Interaction Network (LFEI-Net). LFEI-Net excels in extracting global structures and multi-scale spatial-angular features, capturing feature dependencies through channel modeling and diverse feature interactions. Unlike traditional methods that depend on pyramid and dilated feature extraction, LFEI-Net pioneers an efficient method by integrating large-scale horizontal depth-wise convolution (HDWC) and vertical depth-wise convolution (VDWC) with interactive operations for comprehensive spatial multi-scale feature extraction. Furthermore, we present the Multi-Angular Modeling (MAM) module, which effectively captures scene angle variations from multiple perspectives and precisely delineates object boundaries, thereby improving model adaptability. Our experimental evaluations on two datasets demonstrate that LFEI-Net significantly outperforms state-ofthe-art (SOTA) 2D and 4D light field semantic segmentation methods, achieving mean Intersection over Union (mIoU) of 83.72% and 86.88%, respectively. Fan Shi 0001, Xu Cheng 0003 |
ICASSP | 2 |
| 2025 | Serial Local Patterns and Irregular Dependencies Extract and Cascaded Fusion Network for Structural Crack SegmentationabstractAchieving pixel-level crack segmentation in complex scenarios is a major challenge, as current methods have difficulty effectively integrating both local features and irregular pixel dependencies. In this paper, we introduce a Cascaded Fusion Network (LICFN) specifically designed for crack segmentation, which extracts fine local details and pixel dependencies using a hybrid feature extractor and effectively enhances and fuses them through a cascaded fusion module. To comprehensively evaluate the network, we also created a benchmark dataset, TUT, which includes various scenarios. Experimental results show that our method surpasses others, achieving F1 and mIoU scores of 0.8439 and 0.8509, respectively. The dataset is available at https://github.com/Karl1109/TUT. Xu Cheng 0003, Xiufeng Liu 0001, Fan Shi 0001 |
ICASSP | 5 |
| 2025 | Collaborative Association Network for Multi-view Multi-Human Association and Tracking using Constraint Optimization and Object SearchabstractMulti-view multi-human association and tracking (MvMHAT) enhances scene perception using multiple cameras, crucial for applications such as surveillance and crowd analysis. Inherent feature disparities between views complicate similarity calculations. Recent works combine representation and motion information to address this issue. However, existing methods neglect parallax-induced angular issues and inconsistent object counts across views. To address these challenges, we introduce a collaborative association network combining temporal and spatial clues. Our method incorporates multi-scale adaptive alignment, cross-view and cross-frame feature fusion, to obtain comprehensive global feature representations for each object. We also formulate data association as a mixed-constraint optimization problem to enhance the scalability of our method. Additionally, we propose a novel object search loss to improve cross-view and cross-frame data association. Experiments on benchmarks demonstrate the efficiency of our method in MvMHAT task, significantly outperforming state-of-the-art methods. Fan Shi 0001, Meng Zhao 0001, Xu Cheng 0003 |
ICASSP | 2 |
| 2025 | Multi-Scale Convolutional Networks with Class-Normalized Logit Clipping for Robust Sea State Estimation from Noisy Ship Motion DataabstractAutonomous ships utilize automation systems to achieve unmanned navigation, driving innovation in maritime transportation. However, sea conditions, influenced by dynamic factors such as wave height, wind speed, and ocean currents, present a challenge in accurately assessing these conditions. Traditional classification models often assume accurate labels, but noisy labels are prevalent in real-world applications. Existing methods, such as noise sample filtering or loss function adjustment, have limited applicability and poor generalization when dealing with complex sea condition data. To address this issue, this study proposes an end-to-end neural network model. The model's feature extraction module uses deep representation learning to capture latent patterns in the data, and a loss function is designed to mitigate the impact of outliers. The integration of these components allows the model to perform accurate classification even in the presence of noisy labels. Extensive experiments on public and sea condition datasets validate the effectiveness of this approach, demonstrating that the model exhibits strong generalization capabilities and holds great promise for practical applications. Mengna Liu, Xu Cheng 0003, Xiufeng Liu 0001, Fan Shi 0001, Jianhua Zhang 0002, Shengyong Chen |
ICRA | 5 |
| 2025 | FedFAS: Federated Few-shot Abdominal Organs Segmentation across Heterogeneous ClientsabstractFederated Learning (FL) provides a solution for learning a global model without transferring data, which helps protect privacy in clinical applications. However, existing FL methods often assume that clients have sufficient training samples to generalize the model, and therefore perform poorly in the case of small samples. Furthermore, existing works pay little attention to more challenging medical image segmentation tasks, especially in the case of class-heterogeneous FL. Therefore, in this paper, we construct a framework for federated few-shot medical image segmentation. Specifically, each client obtains local prototypes with limited training samples, which are then uploaded to the server to form a global class prototype library. The clients then select and utilize global class prototypes to calculate global-to-local prototype comparisons to correct local training. In addition, we propose a personalized aggregation strategy for local tasks to enhance the client’s generalization capability for unseen classes and enable the client to learn a discriminative feature space. We establish FL settings using two widely-used datasets and conduct experiments to demonstrate the effectiveness and superiority of our approach. Yi Zhang 0111, Junpeng Wu, Meng Zhao 0001, Xu Cheng 0003, Yao Zhang 0021, Fan Shi 0001 |
IJCNN | 6 |
| 2025 | Dual-Path Contrastive Learning For Wind Turbine Icing DetectionabstractWind energy, characterized by its clean and replenishable nature, is increasingly used worldwide due to its environmental friendliness and wide distribution of resources. However, ice accretion on turbine blades in cold regions, often resulting from cold weather conditions, significantly impacts both the operational performance and security of wind energy production, which significantly increases maintenance costs, resulting in a significant reduction in the energy output performance of wind turbines. Specifically, blade icing alters the aerodynamic characteristics of the blade surface, increases wind resistance, and reduces wind energy conversion efficiency. Furthermore, ice accretion may result in a non-uniform mass allocation across the blade surfaces. This imbalance can induce vibrations within the turbine system, thereby compromising its operational stability and structural integrity. The key challenges are complex sensor parameter variations, high labeling costs, and data imbalance, making accurate icing prediction difficult. In order to tackle these difficulties, this study introduces a technique based on dual-path contrastive learning. This method balances the dataset using a sliding window technique and utilizes both icing loss features and expert features for dual-path processing to fully exploit the feature sets. Furthermore, ice accretion may result in a non-uniform mass allocation across the blade surfaces. This imbalance can induce vibrations within the turbine system, thereby compromising its operational stability and structural integrity, particularly exhibiting excellent performance in handling data imbalance. Aili Xu, Jiamei Zhou, Xu Cheng 0003, Fan Shi 0001, Yongming Han, Guoqian Jiang |
INDIN | 4 |
| 2025 | LFMamba: Focal Stack-aware State Space Modeling for Light Field Salient Object DetectionabstractSalient object detection (SOD) in light field data presents unique challenges due to dynamic semantic inconsistencies across focal slices and representation heterogeneity between focal slices and the all-focus image. Existing methods often treat focal slices uniformly or rely on simple fusion strategies, which fail to address focus-induced semantic drift and cross-modal feature misalignment. To tackle these issues, we propose LFMamba, a unified network that jointly models dynamic semantic consistency and adaptive cross-modal fusion. We design the Focal-aware State Space Module (FSSM), which generates focal-aware semantic prompts through low-rank decomposition and adaptively routes them according to focal plane indices, thereby enabling bidirectional semantic propagation across slices through non-causal state transitions. Furthermore, we introduce the Focal-guided Cross-modal Fusion Module (FCFM), which mitigates cross-modal heterogeneity by a two-stage hierarchical strategy, combining structure-aware low-level alignment and gated high-level semantic fusion. Extensive experiments on four public light field SOD benchmarks demonstrate that LFMamba achieves superior performance compared to state-of-the-art methods, with improved robustness and consistency under complex focal variation scenarios. Xinbo Geng, Fan Shi 0001, Xu Cheng 0003, Meng Zhao 0001, Shengyong Chen |
ACM Multimedia | 2 |
| 2025 | LIDAR: Lightweight Adaptive Cue-Aware Fusion Vision Mamba for Multimodal Segmentation of Structural CracksabstractAchieving pixel-level segmentation with low computational cost using multimodal data remains a key challenge in crack segmentation tasks. Existing methods lack the capability for adaptive perception and efficient interactive fusion of cross-modal features. To address these challenges, we propose a Lightweight Adaptive Cue-Aware Vision Mamba network (LIDAR), which efficiently perceives and integrates morphological and textural cues from different modalities under multimodal crack scenarios, generating clear pixel-level crack segmentation maps. Specifically, LIDAR is composed of a Lightweight Adaptive Cue-Aware Visual State Space module (LacaVSS) and a Lightweight Dual Domain Dynamic Collaborative Fusion module (LD3CF). LacaVSS adaptively models crack cues through the proposed mask-guided Efficient Dynamic Guided Scanning Strategy (EDG-SS), while LD3CF leverages an Adaptive Frequency Domain Perceptron (AFDP) and a dual-pooling fusion strategy to effectively capture spatial and frequency-domain cues across modalities. Moreover, we design a Lightweight Dynamically Modulated Multi-Kernel convolution (LDMK) to perceive complex morphological structures with minimal computational overhead, replacing most convolutional operations in LIDAR. Experiments on three datasets demonstrate that our method outperforms other state-of-the-art (SOTA) methods. On the light-field depth dataset, our method achieves 0.8204 in F1 and 0.8465 in mIoU with only 5.35M parameters. Code and datasets are available at https://github.com/Karl1109/LIDAR-Mamba. Fan Shi 0001, Xu Cheng 0003, Mengfei Shi, Xia Xie 0003, Shengyong Chen |
ACM Multimedia | 3 |
| 2025 | Prior Knowledge-Driven Hybrid Prompter Learning for RGB-Event TrackingabstractEvent data can asynchronously capture variations in light intensity, thereby implicitly providing valuable complementary cues for RGB-Event tracking. Existing methods typically employ a direct interaction mechanism to fuse RGB and event data. However, due to differences in imaging mechanisms, the representational disparity between these two data types is not fixed, which can lead to tracking failures in certain challenging scenarios. To address this issue, we propose a novel prior knowledge-driven hybrid prompter learning framework for RGB-Event tracking. Specifically, we develop a frame-event hybrid prompter that leverages prior tracking knowledge from the foundation model as intermediate modal support to mitigate the heterogeneity between RGB and event data. By leveraging its rich prior tracking knowledge, the intermediate modal reduces the gap between the dense RGB and sparse event data interactions, effectively guiding complementary learning between modalities. Meanwhile, to mitigate the internal learning disparities between the lightweight hybrid prompter and the deep transformer model, we introduce a pseudo-prompt learning strategy that lies between full fine-tuning and partial fine-tuning. This strategy adopts a divide-and-conquer approach to assign different learning rates to modules with distinct functions, effectively reducing the dominant influence of RGB information in complex scenarios. Extensive experiments conducted on two public RGB-Event tracking datasets show that the proposed HPL outperforms state-of-the-art tracking methods, achieving exceptional performance. Mianzhao Wang, Fan Shi 0001, Xu Cheng 0003, Shengyong Chen |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | Zero-shot object visual navigation using relation of historical objects with target transfer
Jiangpeng Zheng, Fan Shi 0001, Meng Zhao 0001, Shengyong Chen |
J. Supercomput. | 2 |
| 2025 | FRAME: Feature Rectification for Class Imbalance LearningabstractClass imbalance learning is a challenging task in machine learning applications. To balance training data, traditional class imbalance learning approaches, such as class resampling or reweighting, are commonly applied in the literature. However, these methods can have significant limitations, particularly in the presence of noisy data, missing values, or when applied to advanced learning paradigms like semi-supervised or federated learning. To address these limitations, this paper proposes a novel and theoretically-ensured latentFeatureRectification method for clAss iMbalance lEarning (FRAME). The proposed FRAME can automatically learn multiple centroids for each class in the latent space and then perform class balancing. Unlike data-level methods, FRAME balances feature in the latent space rather than the original space. Compared to algorithm-level methods, FRAME can distinguish different classes based on distance without the need to adjust the learning algorithms. Through latent feature rectification, FRAME can effectively mitigate contaminated noises/missing values without worrying about structural variations in the data. In order to accommodate a wider range of applications, this paper extends FRAME to the following three main learning paradigms: fully-supervised learning, semi-supervised learning, and federated learning. Extensive experiments on 10 binary-class datasets demonstrate that our FRAME can achieve competitive performance than the state-of-the-art methods and its robustness to noises/missing values. Xu Cheng 0003, Fan Shi 0001, Yao Zhang 0021, Huan Li 0003, Xiufeng Liu 0001, Shengyong Chen |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2025 | TDSF-Net: Tensor Decomposition-Based Subspace Fusion Network for Multimodal Medical Image ClassificationabstractData from multimodalities bring complementary information for deep learning-based medical image classification models. However, data fusion methods simply concatenating features or images barely consider the correlations or complementarities among different modalities and easily suffer from exponential growth in dimensions and computational complexity when the modality increases. Consequently, this article proposes a subspace fusion network with tensor decomposition (TD) to heighten multimodal medical image classification. We first introduce a Tucker low-rank TD module to map the high-level dimensional tensor to the low-rank subspace, reducing the redundancy caused by multimodal data and high-dimensional features. Then, a cross-tensor attention mechanism is utilized to fuse features from the subspace into a high-dimension tensor, enhancing the representation ability of extracted features and constructing the interaction information among components in the subspace. Extensive comparison experiments with state-of-the-art (SOTA) methods are conducted on one self-established and three public multimodal medical image datasets, verifying the effectiveness and generalization ability of the proposed method. The code is available at https://github.com/1zhang-yi/TDSFNet. Yi Zhang 0111, Guoxia Xu, Meng Zhao 0001, Hao Wang 0003, Fan Shi 0001, Shengyong Chen |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2025 | A Novel Robustness-Enhancing Adversarial Defense Approach to AI-Powered Sea State Estimation for Autonomous Marine VesselsabstractSea state information is significant for the guide of maritime activities of autonomous vessels. The sea state estimation (SSE) model, powered by artificial intelligence (AI), has shown great effectiveness but is susceptible to malicious data attacks. These attacks can lead to significant declines in the system’s performance and result in incorrect predictions about the sea state. This study introduces SecureSSE, a strategy for protecting SSE models in autonomous marine vessels from adversarial attacks. This approach incorporates three main components: 1) the multiscale feature extraction learning (MFEL) module; 2) the feature convolution aggregation learning (FCAL) module; and 3) the perturbation examples training (PET) module. The PET module is specifically crafted to create perturbation examples that are in line with unaltered data, leveraging the capabilities of both the MFEL and FCAL modules to efficiently extract and integrate detailed features from ship motion data. Our proposed SecureSSE approach is shown to significantly improve the resilience of deep learning models against potential attacks. Through experimental testing, we have validated the effectiveness of this method in enhancing SSE. Additional ablation studies highlight the critical role of each module within the SecureSSE framework. To our knowledge, this is the first study to address adversarial attacks in this context and to propose a comprehensive defense mechanism for SSE systems in autonomous marine vessels. Xu Cheng 0003, Fan Shi 0001, Hanwei Zhang 0001, Hongning Dai, Houxiang Zhang, Shengyong Chen |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2024 | Robust Classification of Incomplete Time Series with Noisy LabelsabstractMissing data and noisy labeling are common problems in time series analysis. The traditional approach to deal with missing data is to separate interpolation and classification, which is not interactive and provides unsatisfactory performance. While advanced methods can learn features from missing information, feature representation is limited due to the accumulation of interpolation errors. For noisy label interference, a robust loss function is a simpler and more general solution for robust learning. This study proposes an end-to-end neural network that unifies data interpolation and feature learning within a single framework. The focus is placed on extracting useful information from incomplete time series data, and for the computation of classification loss, a robustness loss function is used which effectively reduces the impact of noisy labels. The model is evaluated on 20 univariate time series from the UCR archive after noise processing. The results show that the model outperforms state-of-the-art methods in classifying incomplete time series under noisy labels, especially at high missing rates with high noise rates. Pengshuai Yao, Mengna Liu, Xu Cheng 0003, Fan Shi 0001, Lili Guo 0001 |
CSCWD | 5 |
| 2024 | Overcoming the Pitfalls of Vision-Language Model for Image-Text Retrieval
Feifei Zhang 0001, Sijia Qu, Fan Shi 0001, Changsheng Xu |
ACM Multimedia | 3 |
| 2024 | DSNet: A dynamic squeeze network for real-time weld seam image segmentation
Fan Shi 0001, Mounir Kaaniche, Meng Zhao 0001, Yan Jing, Shengyong Chen |
Eng. Appl. Artif. Intell. | 3 |
| 2024 | Multilevel Signal Decomposition Layer-Specific Residual Network for Blade Icing PredictionabstractWind energy is crucial for sustainable systems but faces reduced productivity due to blade icing. Current detection methods are either costly or heavily reliant on domain-specific knowledge. Data-driven methods show promising performance but encounter challenges such as extracting multi-scale features for blade icing detection from noisy sensor data and addressing the imbalance between icing and non-icing states. To overcome these challenges, we propose an innovative data-driven approach named Multilevel Signal Decomposition Layer-Specific Residual Network (MSD-LRN) for blade icing detection. Our model first employs wavelet decomposition to extract multi-scale features from noisy sensor data and then uses heterogeneous structures to learn the hidden knowledge at each scale. This heterogeneous structure is designed to address the issue of inconsistent information across different scales in wavelet transforms, where information successively decreases. We address data imbalance using resampling techniques. Our approach is validated on three blade icing datasets with varying imbalance ratios, achieving F1 scores of 89.60%, 83.02%, and 78.18%, surpassing existing baselines. Additionally, we introduce random Gaussian noise to test the model's ability to learn robust features from noisy data through wavelet decomposition. Sizhuo Chen, Mengna Liu, Fan Shi 0001, Xu Cheng 0003 |
IEEE Signal Process. Lett. | 4 |
| 2024 | High-Order Spatial Interactions Enhanced Lightweight Model for Optical Remote Sensing Image-Based Small Ship DetectionabstractAccurate and reliable optical remote sensing image-based small-ship detection is crucial for maritime surveillance systems, but existing methods often struggle with balancing detection performance and computational complexity. In this article, we propose a novel lightweight framework called HSI-ShipDetectionNet that is based on high-order spatial interactions (HSIs) and is suitable for deployment on resource-limited platforms, such as satellites and unmanned aerial vehicles. HSI-ShipDetectionNet includes a prediction branch specifically for tiny ships and a lightweight hybrid attention block (LHAB) for reduced complexity. In addition, the use of an HSI module improves advanced feature understanding and modeling ability. Our model is evaluated using the public Kaggle and FAIR1M marine ship detection datasets and compared with multiple state-of-the-art models including small object detection models, lightweight detection models, and ship detection models. The results show that HSI-ShipDetectionNet outperforms the other models in terms of detection performance while being lightweight and suitable for deployment on resource-limited platforms. Xu Cheng 0003, Fan Shi 0001, Xiufeng Liu 0001, Huan Huo, Shengyong Chen |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | Selective Feature Fusion and Irregular-Aware Network for Pavement Crack DetectionabstractRoad cracks on highways and main roads are among the most prominent defects. Given the inherent inaccuracy, time-consuming nature, and labor intensiveness of manual road crack detection, there’s a compelling need for automated solutions. The irregular shape of cracks, along with complex background conditions encompassing varying lighting, tree shadows, and dark stains, poses a significant challenge for computer vision-based approaches. Most cracks exhibit irregular edge patterns, which are pivotal features for accurate detection. In response to recent advancements in deep learning within the realm of computer vision, this paper introduces an innovative neural network architecture termed the ‘Selective Feature Fusion and Irregular-Aware Network (SFIAN)’ designed specifically for crack detection on pavements. The proposed network selectively integrates features from multiple levels, enhancing and controlling the flow of valuable information at each stage while effectively modeling irregular crack objects. In an extensive evaluation, this paper conducts experiments on five distinct crack datasets and compares the results with twelve state-of-the-art crack detection methods, including the latest edge detection and semantic segmentation techniques. The experimental findings demonstrate the superior performance of the proposed method, surpassing baseline methods by a notable margin, with an increase of approximately 13.3% in the F1-score, all without introducing additional time complexity. Furthermore, the model achieves real-time processing, achieving a remarkable speed of 35 frames per second (FPS) on images at 320$\times$480 pixels, facilitated by NVIDIA 3090 hardware. Xu Cheng 0003, Fan Shi 0001, Meng Zhao 0001, Xiufeng Liu 0001, Shengyong Chen |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2024 | A Prototype-Empowered Kernel-Varying Convolutional Model for Imbalanced Sea State Estimation in IoT-Enabled Autonomous ShipabstractSea State Estimation (SSE) is essential for Internet of Things (IoT)-enabled autonomous ships, which rely on favorable sea conditions for safe and efficient navigation. Traditional methods, such as wave buoys and radars, are costly, less accurate, and lack real-time capability. Model-driven methods, based on physical models of ship dynamics, are impractical due to wave randomness. Data-driven methods are limited by the data imbalance problem, as some sea states are more frequent and observable than others. To overcome these challenges, we propose a novel data-driven approach for SSE based on ship motion data. Our approach consists of three main components: a data preprocessing module, a parallel convolution feature extractor, and a theoretical-ensured distance-based classifier. The data preprocessing module aims to enhance the data quality and reduce sensor noise. The parallel convolution feature extractor uses a kernel-varying convolutional structure to capture distinctive features. The distance-based classifier learns representative prototypes for each sea state and assigns a sample to the nearest prototype based on a distance metric. The efficiency of our model is validated through experiments on two SSE datasets and the UEA archive, encompassing thirty multivariate time series classification tasks. The results reveal the generalizability and robustness of our approach. Mengna Liu, Xu Cheng 0003, Fan Shi 0001, Xiufeng Liu 0001, Hongning Dai, Shengyong Chen |
IEEE Trans. Sustain. Comput. | 3 |
| 2023 | Enhancing Ocean Scene Video Captioning with Multimodal Pre-Training and Video-Swin-TransformerabstractWith the success of multimodal pre-training models in the video-language field and various downstream tasks, previous multimodal models used 3DCNN networks as video feature extractors, which have limitations in interacting and fusing with text features. This paper proposes a multimodal pre-training model that utilizes a Video-Swin-Transformer-based network to encode both video and text data, to achieve better performance in video understanding. The model consists of four modules: video encoder, text encoder, interact encoder, and caption decoder to accomplish the task of ocean scene video captioning. A dataset of ocean scene videos, including various content types such as sea surfaces and shores, is also constructed. The training process is divided into two stages: pre-training and fine-tuning. Pre-training is performed on the Howto100m dataset to allow the model to learn video captions in natural scenes and complete video-language matching tasks. The fine-tuning stage is then performed on the ocean1000 dataset to better understand the events and content in ocean scene videos and generate captions that conform to ocean scene video descriptions. The model achieves satisfying results on both the public dataset YouCook2 and the proprietary dataset Ocean1000, demonstrating its ability in video-text information fusion and interaction. Meng Zhao 0001, Fan Shi 0001, Meng'en Zhang, Yu He 0001, Shengyong Chen |
IECON | 3 |
| 2023 | Encoder Activation Diffusion and Decoder Transformer Fusion Network for Medical Image Segmentation
Xueru Li, Guoxia Xu, Meng Zhao 0001, Fan Shi 0001, Hao Wang 0003 |
PRCV (13) | 4 |
| 2023 | SANet: A novel segmented attention mechanism and multi-level information fusion network for 6D object pose estimation
Xinbo Geng, Fan Shi 0001, Xu Cheng 0003, Mianzhao Wang, Shengyong Chen, Hongning Dai |
Comput. Commun. | 2 |
| 2023 | Learning intra-inter-modality complementary for brain tumor segmentation
Jiangpeng Zheng, Fan Shi 0001, Meng Zhao 0001 |
Multim. Syst. | 2 |
| 2023 | Visual Object Tracking Based on Light-Field Imaging in the Presence of Similar DistractorsabstractVisual object tracking is of great importance in the field of computer vision. One of the main challenges is the difficulty of identifying moving targets from nearby similar distractors with a single-view image of the scene. To overcome this challenge, in this article, we acquire multiview images of the scenes by using a light-field camera. The multiview images are able to capture the 4-D structure instead of the 2-D plane of the objects but are more difficult to process. Therefore, we propose a novel representation for multiview images, i.e., the macro-epipolar plane image (macro-EPI), which highlights both spatial topological and angular information of the target and distractors. It is obtained by slicing the original multiview images into pieces and properly restacking these pieces in an ordinal manner. The resulting macro-EPI is mapped into the 2-D space; therefore, we adapt a modified autoencoder network to train a macro-EPI feature extractor. Thereafter, we design a composite framework of two-pattern convolution filters based on a discriminative correlation filter for object tracking, which successfully discriminates the target from the distractors by merging the macro-EPI features and the single-view image features. The experiments also show that our method outperforms the state-of-the-art methods in the presence of similar distractors. Mianzhao Wang, Fan Shi 0001, Xu Cheng 0003, Meng Zhao 0001, Yao Zhang 0021, Shengyong Chen |
IEEE Trans. Ind. Informatics | 2 |
| 2023 | A Novel Class-Imbalanced Ship Motion Data-Based Cross-Scale Model for Sea State EstimationabstractSea state estimation (SSE) is significant to the development of autonomous ships, which can enhance the sustainable development of maritime transportation. Traditional model-based methods are limited by their drawbacks, such as high costs and inaccurate estimations. The deep learning model shows superior performance, but it requires that the sample quantity for each sea state should be almost the same. Since the occurrence probability of each state is different, the ships mainly work in low sea states, and the collected ship motion data for different sea states are highly imbalanced. This work proposes a novel class-imbalanced ship motion data-based cross-scale model for SSE. The model consists of three major components: a multi-scale feature learning module, a cross-scale feature learning module, and a prototype classifier module. The multi-scale and cross-scale feature learning modules are designed to learn abundant coarse and fine-level features from the ship motion data. The prototype classifier is utilized to overcome the limitation of the conventional softmax classifier to produce better estimates. Our research highlights our model’s remarkable scalability and versatility with 30 publicly available datasets in time series classification, demonstrating superior performance over baseline methods in 21 cases. Notably, it outperformed ShapeNet by 5.72% and EDI by 26.3%. We further validated our model’s proficiency using ship motion datasets, consistently surpassing eight state-of-the-art baselines and five class-imbalanced learning methods. Ablation and sensitivity studies, emphasize the critical role of each model component. Our findings underscore the model’s robustness and its potential to advance time series classification in diverse domains. Xu Cheng 0003, Xiufeng Liu 0001, Fan Shi 0001, Zhengru Ren, Shengyong Chen |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2022 | Object Detection based on Light Field ImagingabstractImproving the object confidence score of an image using the machine intelligence method is one of the critical objectives for object detection. In this paper, we propose an end-to-end framework to object detection based on light field (LF) imaging. First, we apply refocusing technology to enhance the visual feature expression between different objects in LF images. then, we combine the LF refocusing technology with the efficient darknet53 feature extraction network and multi-feature fusion method, which ensures effectiveness in improving the confidence score of the object in the image. To evaluate the framework, we use a light field camera (LFC) to construct a new real LF image detection dataset, which consists of as many as 20 kinds of common objects (person, car, etc.). Compared with the popular methods, the confidence score achieved by our method shows an improvement for the detection of any single object. Specifically, for the detection of cars and people, the confidence score is increased by 0.03 and 0.02, respectively. For the detection of bicycles, the confidence score significantly improves by 0.27. To some extent our method can also solve the problem of object missed and false detection. Fan Shi 0001, Meng Zhao 0001 |
CSCWD | 2 |
| 2022 | LFBCNet: Light Field Boundary-aware and Cascaded Interaction Network for Salient Object DetectionabstractIn light field imaging techniques, the abundance of stereo spatial information aids in improving the performance of salient object detection. In some complex scenes, however, applying the 4D light field boundary structure to discriminate salient objects from background regions is still under-explored. In this paper, we propose a light field boundary-aware and cascaded interaction network based on light field macro-EPI, named LFBCNet. Firstly, we propose a well-designed light field multi-epipolar-aware learning (LFML) module to learn rich salient boundary cues by perceiving the continuous angle changes from light field macro-EPI. Secondly, to fully excavate the correlation between salient objects and boundaries at different scales, we design multiple light field boundary interactive (LFBI) modules and cascade them to form a light field multi-scale cascade interaction decoder network. Each LFBI is assigned to predict exquisite salient objects and boundaries by interactively transmitting the salient object and boundary features. Meanwhile, the salient boundary features are forced to gradually refine the salient object features during the multi-scale cascade encoding. Furthermore, a light field multi-scale-fusion prediction (LFMP) module is developed to automatically select and integrate multi-scale salient object features for final saliency prediction. The proposed LFBCNet can accurately distinguish tiny differences between salient objects and background regions. Comprehensive experiments on large benchmark datasets prove that the proposed method achieves competitive performance over 2-D, 3-D, and 4-D salient object detection methods. Mianzhao Wang, Fan Shi 0001, Xu Cheng 0003, Meng Zhao 0001, Yao Zhang 0021, Shengyong Chen |
ACM Multimedia | 2 |
| 2022 | Synthetic-to-real: instance segmentation of clinical cluster cells with unlabeled synthetic trainingabstractMOTIVATION: The presence of tumor cell clusters in pleural effusion may be a signal of cancer metastasis. The instance segmentation of single cell from cell clusters plays a pivotal role in cluster cell analysis. However, current cell segmentation methods perform poorly for cluster cells due to the overlapping/touching characters of clusters, multiple instance properties of cells, and the poor generalization ability of the models. RESULTS: In this article, we propose a contour constraint instance segmentation framework (CC framework) for cluster cells based on a cluster cell combination enhancement module. The framework can accurately locate each instance from cluster cells and realize high-precision contour segmentation under a few samples. Specifically, we propose the contour attention constraint module to alleviate over- and under-segmentation among individual cell-instance boundaries. In addition, to evaluate the framework, we construct a pleural effusion cluster cell dataset including 197 high-quality samples. The quantitative results show that the numeric result of APmask is > 90%, a more than 10% increase compared with state-of-the-art semantic segmentation algorithms. From the qualitative results, we can observe that our method rarely has segmentation errors. Meng Zhao 0001, Fan Shi 0001, Xuguo Sun, Shengyong Chen |
Bioinform. | 3 |
| 2022 | Light field imaging for computer vision: a surveyabstractLight field (LF) imaging has attracted attention because of its ability to solve computer vision problems. In this paper we briefly review the research progress in computer vision in recent years. For most factors that affect computer vision development, the richness and accuracy of visual information acquisition are decisive. LF imaging technology has made great contributions to computer vision because it uses cameras or microlens arrays to record the position and direction information of light rays, acquiring complete three-dimensional (3D) scene information. LF imaging technology improves the accuracy of depth estimation, image segmentation, blending, fusion, and 3D reconstruction. LF has also been innovatively applied to iris and face recognition, identification of materials and fake pedestrians, acquisition of epipolar plane images, shape recovery, and LF microscopy. Here, we further summarize the existing problems and the development trends of LF imaging in computer vision, including the establishment and evaluation of the LF dataset, applications under high dynamic range (HDR) conditions, LF image enhancement, virtual reality, 3D display, and 3D movies, military optical camouflage technology, image recognition at micro-scale, image processing method based on HDR, and the optimal relationship between spatial resolution and four-dimensional (4D) LF information acquisition. LF imaging has achieved great success in various studies. Over the past 25 years, more than 180 publications have reported the capability of LF imaging in solving computer vision problems. We summarize these reports to make it easier for researchers to search the detailed methods for specific solutions. Fan Shi 0001, Meng Zhao 0001, Shengyong Chen |
Frontiers Inf. Technol. Electron. Eng. | 2 |
| 2022 | A Class-Imbalanced Heterogeneous Federated Learning Model for Detecting Icing on Wind Turbine BladesabstractWind farms are typically located at high latitudes, resulting in a high risk of blade icing. Data-driven approaches offer promising solutions for blade icing detection, but they rely on a considerable amount of data. Data exchange between multiple wind farms would improve the performance of detection models, due to the spatio-temporal dependencies capable of reflecting different meteorological conditions. The traditional centralized approach for icing detection faces many challenges, including the requirement of high storage and computational capacity of the server, vulnerability to cyberattacks, and operators’ reluctance of sharing data for commercial reasons. To address these challenges, this article proposes a heterogeneous federated learning (FL) model for wind turbine blade icing detection. The structures of the server and client models in the presented method are different, in contrast to the traditional FL of sharing the same structure. In addition, this article addresses the class imbalance problem in the training data. Last, this article conducts comprehensive experiments to evaluate the proposed method using real-world data from 20 turbines in two wind farms, and compares it with two state-of-the-art FL models and five well-known class imbalance methods. The experimental results verify the effectiveness and superiority of the proposed method. Xu Cheng 0003, Fan Shi 0001, Yongping Liu, Jiehan Zhou, Xiufeng Liu 0001 |
IEEE Trans. Ind. Informatics | 2 |
| 2022 | A Blockchain-Empowered Cluster-Based Federated Learning Model for Blade Icing Estimation on IoT-Enabled Wind TurbineabstractWind energy is a fast-growing renewable energy but faces blade icing. Data-driven methods provide talented solutions for blade icing detection, but a considerable amount of Internet of Things data needs to be collected to a central server, which may lead to the leakage of sensitive business data. To address this limitation, this article proposesBLADE, a Blockchain-empowered imbalanced federated learning (FL) model for blade icing detection. With the help of the Blockchain, the conventional FL is improved without worrying about the failure of the single centralized server and boosts the privacy preserving. A validation mechanism is introduced into the Blockchain to enhance the defense against poisoning attacks. In addition, a novel imbalanced learning algorithm is integrated into BLADE to solve the class imbalance problem in the sensor data. BLADE is evaluated on ten wind turbines from two wind farms. The experimental results verify the effectiveness, superiority, and feasibility of the proposed BLADE. Xu Cheng 0003, Fan Shi 0001, Meng Zhao 0001, Shengyong Chen, Hao Wang 0003 |
IEEE Trans. Ind. Informatics | 3 |
| 2022 | An Online Multiobject Tracking Network for Autonomous Driving in Areas Facing EpidemicabstractMulti-object tracking is of great importance in autonomous driving. However, with the outbreak of COVID-19, multi-object tracking faces new challenges in areas gripped by epidemics because of complex motion blur, frequent occlusions, and appearance deformations. To reliably improve object trajectory association in epidemic-plagued areas, we propose a temporal-spatial aggregation embedding network (TSAEN) for multi-object tracking. Our embedding network contains a temporal-aware correlation module (TACM) and spatial-aggregate embedding module (SAEM) that can fully obtain and aggregate appearance clues related to moving objects in previous frames. The TACM learns the temporal homogeneity features of the current and previous frames to perceive features with correlated appearance cues. Then, the SAEM adjusts the spatial deformation for each perceived temporal homogeneity feature and aggregates them for re-ID embedding learning. The experimental results demonstrate that our proposed method is able to achieve excellent overall performance. Mianzhao Wang, Fan Shi 0001, Meng Zhao 0001, Xu Cheng 0003 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2022 | A Novel Deep Class-Imbalanced Semisupervised Model for Wind Turbine Blade Icing DetectionabstractWind energy is of great importance for future energy development. In order to fully exploit wind energy, wind farms are often located at high latitudes, a practice that is accompanied by a high risk of icing. Traditional blade icing detection methods are usually based on manual inspection or external sensors/tools, but these techniques are limited by human expertise and additional costs. Model-based methods are highly dependent on prior domain knowledge and prone to misinterpretation. Data-driven approaches can offer promising solutions but require a massive amount of labeled training data, which are not generally available. In addition, the data collected for icing detection tend to be imbalanced because, most of the time, wind turbines operate under normal conditions. To address these challenges, this article presents a novel deep class-imbalanced semisupervised (DCISS) model for estimating blade icing conditions. DCISS integrates class-imbalanced and semisupervised learning (SSL) using a prototypical network that can rebalance features and measure the similarities between labeled and unlabeled samples. In addition, a channel calibration attention module is proposed to improve the ability to extract features from raw data. The proposed model has been evaluated using the blade icing datasets of three wind turbines. Compared to the classical anomaly detection and state-of-the-art SSL algorithms, DCISS shows significant advantages in terms of accuracy. Compared to five different class-imbalanced loss functions, the proposed DCISS is competitive. The generalization and practicability of the proposed model are further verified in the use case of online estimation. Xu Cheng 0003, Fan Shi 0001, Xiufeng Liu 0001, Meng Zhao 0001, Shengyong Chen |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2020 | Fused 3-Stage Image Segmentation for Pleural Effusion Cell ClustersabstractThe appearance of tumor cell clusters in pleural effusion is usually a vital sign of cancer metastasis. Segmentation, as an indispensable basis, is of crucial importance for diagnosing, chemical treatment, and prognosis in patients. However, accurate segmentation of unstained cell clusters containing more detailed features than the fluorescent staining images remains to be a challenging problem due to the complex background and the unclear boundary. Therefore, in this paper, we propose a fused 3-stage image segmentation algorithm, namely Coarse segmentation-Mapping-Fine segmentation (CMF) to achieve unstained cell clusters from whole slide images. Firstly, we establish a tumor cell cluster dataset consisting of 107 sets of images, with each set containing one unstained image, one stained image, and one ground-truth image. Then, according to the features of the unstained and stained cell clusters, we propose a three-stage segmentation method: 1) Coarse segmentation on stained images to extract suspicious cell regions-Region of Interest (ROI); 2) Mapping this ROI to the corresponding unstained image to get the ROI of the unstained image (UI-ROI); 3) Fine Segmentation using improved automatic fuzzy clustering framework (AFCF) on the UI-ROI to get precise cell cluster boundaries. Experimental results on 107 sets of images demonstrate that the proposed algorithm can achieve better performance on unstained cell clusters with an F1 score of 90.40%. Sike Ma, Meng Zhao 0001, Hao Wang 0003, Fan Shi 0001, Xuguo Sun, Shengyong Chen, Hongning Dai |
ICPR | 4 |