Chengliang Wang 0002

dblp:77/1962-2 · DBLP profile ↗
← Back
73ranked-venue papers
11as first author
58since 2021 · last 2026
0000-0003-0877-1064ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 25 · 4 first-author · 21 since 2021Graphics, computer vision, multimedia, augmented reality and games · 18 · 4 first-author · 15 since 2021Systems, architecture and hardware · 13 · 2 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 1 first-author · 9 since 2021Databases, data management, data science and information retrieval · 6 · 3 since 2021Human-computer interaction and ubiquitous computing · 6 · 5 since 2021Computer networks · 2 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 GeoCon-Diff: Geometrically-Consistent Diffusion for High-Fidelity 3D Reconstruction from a Single Occluded Image
Chengliang Wang 0002
ICIC (20)2
2026 Voxel-to-pillar: A height-aware pillar feature encoding method for efficient 3D object detection
Chengliang Wang 0002, Shu-Mao Wang, Ji Liu 0006, Yonggang Luo
Inf. Sci.2
2026 GFPP-MAE: gradient-guided frequency reconstruction and position predictions advance MAE for 3D CT image segmentation
Yuping Peng, Xing Xiao, Chengliang Wang 0002, Hongqian Wang
Multim. Syst.4
2025 SAM-OCTA2: Layer Sequence OCTA Segmentation with Fine-tuned Segment Anything Model 2
abstract
Segmentation of indicated targets aids in the precise analysis of optical coherence tomography angiography (OCTA) samples. Existing segmentation methods typically perform on 2D projection targets, making it challenging to capture the variance of segmented objects through the 3D volume. To address this limitation, the low-rank adaptation technique is adopted to fine-tune the Segment Anything Model (SAM) version 2, enabling the tracking and segmentation of specified objects across the OCTA scanning layer sequence. To further this work, a prompt point generation strategy in frame sequence and a sparse annotation method to acquire retinal vessel (RV) layer masks are proposed. This method is named SAM-OCTA2 and has been experimented on the OCTA-500 dataset. It achieves state-of-the-art performance in segmenting the foveal avascular zone (FAZ) on regular 2D en-face and effectively tracks local vessels across scanning layer sequences. The code is available at: https://github.com/ShellRedia/SAM-OCTA2.
Xinrun Chen, Chengliang Wang 0002, Haojian Ning, Mengzhan Zhang, Mei Shen
ICASSP2
2025 LLGS: Illuminating Gaussian Splatting via absorptance Modulation
abstract
Low-light images are typically characterized by low pixel intensity and color distortion, presenting a significant challenge for accurate 3D reconstruction with 3D Gaussian Splatting (3DGS). Traditional 2D enhancement methods fail to maintain consistent illumination, affecting reconstruction quality. We propose Low-Light Gaussian (LLGS), which can directly leverage low-light images for 3D reconstruction and synthesizing normal-light novel views. LLGS incorporates absorptance to simulate light behavior in low-light conditions, assuming objects maintain normal illumination while reflected light intensity attenuates due to absorptance during rendering. This approach enables the capture of authentic color information in dimly lit scenes. LLGS outperforms current enhancement algorithms and Neural Radiance Fields (NeRF) in image quality and processing efficiency, making it highly effective for low-light 3D scene reconstruction and novel view synthesis.
Jianwen Gan, Bo Zheng 0007, Chengliang Wang 0002, Yingbo Wu
ICASSP4
2025 SAM Adaptation with Refocused Attention and Diverse Prompts for Medical Image Segmentation
abstract
The adaptation research of SAM in the field of medical image mainly adopts two types of methods: parameter fine-tuning and prompt engineering, but these methods face two main issues: (1)Parameter fine-tuning methods have limitations in focusing the model encoder’s attention on the foreground of medical image; (2)Prompt engineering methods have not fully utilized SAM’s ability to handle diverse types of prompts. In response to these issues, we propose the SAM-RD model, which includes a Refocused Attention module (RA), a Sparse Prompt generator (SP), and a Dense Prompt generator (DP). RA refocuses the attention within the SAM encoder to optimize its ability to capture segmentation object features. The SP and DP automatically generate diverse prompts, including positive/negative point prompts, box prompts, and mask prompts. Experiments on multiple medical datasets of different modalities show that SAM-RD outperforms the current state-of-the-art model based on SAM in medical image segmentation.
Liangshan Zhu, Chengliang Wang 0002
ICASSP3
2025 SA-SAM: Stronger Adaptation for SAM in Camouflaged Object Detection
Zhengqiang Jia, Chengliang Wang 0002, Zhongshi He
ICIC (3)4
2025 SMBA-MIL: SAM-Enhanced Multi-branch Attention Multi-instance Learning for Whole Slide Image Classification
Biyun Zhou, Chengliang Wang 0002, Chao Liao, Hongqian Wang
ICIC (5)2
2025 SAM-GA: SAM-Guided Grouped Aggregation Network for Weakly Supervised cardiac MRI Segmentation
abstract
Scribble supervision is increasingly vital in cardiac MRI segmentation due to its low annotation cost. However, challenges still exist due to limited supervision signals and the complexity of cardiac MRI. We propose SAM-GA, a SAM-guided segmentation training framework, to address these challenges. SAM generates pseudo-labels to supervise segmentation networks, addressing limited supervision signals challenge. Meanwhile, we propose three strategies, entropy-based point prompting, multi-class fusion, and cross-pseudo-label fine-tune, to optimize these pseudo-labels and adapt SAM for cardiac MRI. Additionally, we design a dual-branch network, Grouped Aggregation Network(GANet), to handle the complexity of cardiac MRI challenge. GANet improves segmentation consistency and accuracy by guiding each branch to focus on different granularity information and combining their outputs. Our comparative experiments on two cardiac benchmark datasets demonstrate that our method achieves state-of-the-art performance, validating its effectiveness in addressing the challenges of weakly supervised cardiac MRI segmentation. The code is available on GitHub.
Chengliang Wang 0002, Yonggang Luo
ICME2
2025 Retinal OCT Anomaly Detection Based on Suspicious Strategy and Relational Learning
abstract
Retinal OCT anomaly detection is an important research field of ocular disease diagnostics. Reconstruction-based methods utilizing only normal samples are the dominant approach in the field. These studies have yielded significant results. However, two issues remain: (1) Due to the strong generalization of the model, abnormal regions can also be reconstructed correctly, leading to small reconstruction error that affect classification accuracy. (2) The reconstruction-based method only focuses on the internal features of the individual samples, unable to model the large differences among normal samples. In this paper, we propose a model named SR-AD, including a Difference-Awareness Module (DAM) and a Relational Consistency Learning module (RCL). DAM designs a masking strategy that reduces the exposure of abnormal information to avoid good reconstruction of abnormal regions, thus enlarging the abnormal image reconstruction error. RCL reconstructs the relationships between normal samples to model the characteristics of normal samples comprehensively from a global perspective. Experiments on the SpectralisOCT dataset validate that our model achieves advanced performance, reaching 99.07% in AUC.
Minghui Zhai, Liangshan Zhu, Chengliang Wang 0002, Yonggang Luo
ICME4
2025 Nucleus-SAM:Point-Supervised SAM for Nucleus Segmentation
abstract
Point-supervised nucleus segmentation in pathology images is crucial and efficient for medical applications. However, point-supervised methods face two limitations: 1) They generate pseudo-labels using image prior knowledge, but the diversity of nuclear morphology leads to shape noise in the pseudo-labels. 2) Current point-supervised models are trained on limited data, lacking generalization. In recent years, the vision foundation model SAM has garnered attention due to its strong generalization ability. Existing studies fine-tune SAM for nucleus segmentation but still require pixel-level labels. In this paper, We introduce Nucleus-SAM, consisting of two modules: NDM and GAM. 1) NDM utilizes heat maps, solving the issue of pseudo-labels noise and better learning the morphological features of the nucleus. 2) GAM combines weakly supervised fine-tuning and automatic point generator, enabling SAM to adapt to nucleus segmentation, enhancing the model’s overall segmentation performance and generalization without pixel-level labels. Experiments show that Nucleus-SAM outperforms SOTA point-supervised and SAM point prompt models by an average of 3.26% in IoU and 2.33% in Dice scores across multiple datasets.
Liangshan Zhu, Chengliang Wang 0002, Zailin Yang
ICME4
2025 Assisted-MobileSAM: Unleashing the Potential of MobileSAM through Teacher Assistant Distillation
abstract
The Segment Anything Model (SAM) has demonstrated remarkable potential in the field of image segmentation, but its high computational cost limits its application on edge devices. Therefore, effectively compressing SAM has become a current research focus. Existing studies have attempted to compress SAM using knowledge distillation techniques, but these methods encounter two critical issues: 1) They overlook the capacity gap between the teacher and student models, making it difficult for the student model to effectively learn and mimic the knowledge of the teacher model; 2) The allocation of the student’s learning capacity is suboptimal, as the student lacks focus in mimicking the teacher’s features, leading to an excessive emphasis on redundant features. To address these issues, we propose a distillation framework, Assisted-MobileSAM, built upon the MobileSAM architecture. By introducing a coupled-optimization teacher assistant model as an intermediary for knowledge transfer, our framework alleviates the model capacity discrepancy through a staged distillation strategy. Additionally, we dynamically adjust the student’s focus on mimicking the teacher assistant’s image embeddings through targeted learning knowledge distillation, thereby optimizing the knowledge transfer process. Experimental results show that our method achieves similar parameter size and inference speed to TinySAM, with a 3.52%/0.81% improvement in mIoU on the SA-1B/COCO dataset.
Dejian Fang, Chengliang Wang 0002, Lingqiu Zeng
IJCNN4
2025 mimicSAM: Guided Knowledge Transfer for SAM via Key Regions and Boundary Information
abstract
Segment Anything Model (SAM) has received wide attention for its excellent segmentation capability and generalization performance, but its vast image encoder limits its deployment in real-time applications. To solve this problem, many researchers adopt knowledge distillation to lighten the image encoder of SAM, but the current work on knowledge distillation for SAM suffers from two problems: one is ignoring the learning of inter-sample relations, and another is ignoring the transfer of boundary information. Therefore, we propose a distillation method, mimcSAM, which is based on MobileSAM and introduces two modules in the feature map extraction phase: the Regional Attention Distillation Module (RADM) and the Boundary Gradient Distillation Module (BGDM). The RADM explicitly models relationships between samples by filtering key regions and constructing a queue of regional features, enabling the student model to align its relational understanding of key regions across samples. BGDM employs the Sobel operator to compare gradients on feature maps of teacher and student models, guiding the student to learn the teacher model’s sensitivity to boundaries. Experiments on the SA-1B, COCO, and LVIS datasets show that our method is 1.8% higher on mAP and 1.5% higher on mIoU than the baseline. Compared with the current relational distillation method, our method has the most significant improvement, which proves its superiority in improving the segmentation accuracy.
Yixing Ma, Chengliang Wang 0002, Lingqiu Zeng, Hongqian Wang
IJCNN2
2025 LaneGS:Novel View Synthesis for Lane Change Scenarios in Autonomous Driving
abstract
In autonomous driving, lane changing plays a crucial role in data augmentation. By synthesizing new viewpoints during lane changing, we can effectively increase the diversity of training data, thereby improving the model’s generalization capability in different driving scenarios. However, autonomous driving data is often collected from a single lane, leading to sparse viewpoints, which makes it challenging to synthesize lane changes using existing methods like NeRF and 3D Gaussian Splatting, as they suffer from significant rendering quality degradation, including road texture blurring and distortion, especially with large viewpoint shifts. To address this issue, the LaneGS method is proposed, aiming to optimize novel view synthesis for road regions. Specifically, two key modules are designed: the point cloud normal constraint module and the virtual view regularization module. The point cloud normal constraint module leverages prior knowledge of point cloud normals to guide the orientation and shape of Gaussian ellipsoids, ensuring alignment with the road’s geometry, reducing overfitting, and enhancing multi-view consistency. Furthermore, to further improve the rendering quality of new views after lane changing, the virtual view regularization module incorporates information from the new viewpoint, effectively guiding the optimization process. Qualitative and quantitative comparisons demonstrate the effectiveness of our method in synthesizing new views after lane changing. Our method achieves high-quality lane change view synthesis with a 3-meter horizontal displacement on the Waymo, KITTI, and Virtual KITTI 2 datasets. Additionally, our optimization improves the FID score by 13.40% while maintaining the reconstruction quality of the original views.
Youtao Tang, Ji Liu 0006, Chengliang Wang 0002, Yonggang Luo
IJCNN3
2025 SBR-GS:Enhancing Scene Editability based on Static Background Reconstruction using Gaussian Splatting
abstract
Removing dynamic objects and reconstructing high-quality background are crucial for enhancing model editability and scene diversity in autonomous driving. Existing methods face challenges with the holes due to the persistent occlusion of the street background by dynamic objects, limiting flexible scene editing. To address this, we propose a framework, SBR-GS, focusing on removing dynamic objects and improving the reconstruction performance of the static background. Specifically, Inpaint Anything (IA) is integrated to resolve hole issues from continuous occlusions. A texture consistency module is designed to optimize the texture continuity between occluded and visible regions. To mitigate artifacts from the newly generated background lacking depth supervision, a depth estimation network is incorporated to enhance background depth in occluded areas. Additionally, a pixel-aware gradient density control method is introduced to improve overall reconstruction quality, addressing clarity issues due to insufficient initial point clouds. Experimental results indicate that this method achieves a 22.05% improvement in the FID metric on the Waymo dataset and a 26.67% improvement on the KITTI dataset compared to state-of-the-art methods. This approach not only enhances static background reconstruction quality but also supports the development of highly editable models in autonomous driving contexts.
Chengliang Wang 0002, Ji Liu 0006, Yonggang Luo
IJCNN2
2025 SAM-AEKD: SAM-Based Knowledge Distillation with Adaptive Fusion and Edge Features for Enhancing Medical Image Segmentation
abstract
Semantic segmentation models have made significant progress by optimizing for the characteristics of medical images, but their accuracy and generalization ability still need improvement. The vision foundation model SAM excels in segmentation accuracy and generalization, but its performance drops significantly when directly applied to medical image segmentation due to its training on natural images. To address this, we propose a distillation framework with SAM as the teacher and medical segmentation model as the student, distilling SAM’s powerful feature extraction capabilities into medical segmentation model to enhance its performance. Directly using existing knowledge distillation methods introduces two challenges: 1) a large semantic gap between SAM and medical segmentation model’s feature layers; 2) SAM’s strong edge feature extraction ability is underutilized. In this paper, we present the SAM-Based Knowledge Distillation with Adaptive Fusion and Edge Features (SAM-AEKD), consisting of two modules: FAFM and EFEM. FAFM dynamically weights and convolves multi-layer features from the teacher or student encoders, avoiding semantic mismatches during distillation. EFEM extracts edge features from the fused features using multi-scale convolution and channel aggregation, enabling knowledge distillation in both edge and fused feature spaces, enhancing the student’s edge feature extraction ability. Experiments on three medical modality datasets show that SAM-AEKD significantly improves students’ segmentation accuracy, outperforming other distillation methods.
Shirong Zhou, Chengliang Wang 0002, Lingqiu Zeng, Zhongshi He
IJCNN3
2025 SCN-Pillar: Construct a Pillar-based Fully Sparse Lightweight 3D Detector via Sparse ConvNeXt
abstract
Since autonomous driving requires high-precision object detection in real-time, and multi-line LiDAR generates huge point clouds, developing a lightweight 3D detector is crucial. The high sparsity and unstructured nature of point clouds require transforming raw data into a structured format for effective feature extraction. Nevertheless, despite the decrease in computational complexity achieved through the transformation, the resulting structure exhibits high sparsity. Consequently, using conventional neural networks for detectors necessitates substantial additional computational resources. Voxel-based detectors densely partition the point cloud in the height space and must use 3D convolutions. Therefore, compared with pillar-based 3D detectors, voxel-based 3D detectors generally achieve higher object detection accuracy, but their detection speed is much slower. Given these challenges, we propose a fully sparse ConvNeXt block for more efficient pillar feature extraction that selectively extracts features from effective data positions. We have developed SCN-Pillar, a pillar-based, fully sparse, lightweight 3D detector that adopts the sparse ConvNeXt. The SCN-Pillar has been validated on the Waymo open dataset, showcasing enhancements in accuracy across a range of object detection tasks. The APH improvement in pedestrian detection has reached more than 1.2. It only requires the computational cost of the pillar-based solution, yet its object detection accuracy exceeds that of the voxel-based solution. The object detection speed reaches 18.28 FPS. The code is available at https://github.com/kaikailab/SCN-Pillar.
Chengliang Wang 0002, Yonggang Luo, Bo Zheng 0007
ICMR2
2025 POGS: Position-Optimized Gaussian Splatting for Reconstruction in Unevenly Illuminated Scenarios
abstract
The advent of 3D Gaussian Splatting (3DGS) has achieved a major breakthrough in the field of 3D reconstruction and realized real-time high-quality novel view synthesis. However, 3DGS relies on the color backpropagation gradients, which are greatly affected by the illumination intensity, to optimize the position parameters of gaussian spheres such as the average center coordinates and opacity, resulting in issues like blurring and artifacts when handling scenes with uneven illumination. This is particularly important in 3D reconstruction of autonomous driving scenarios, especially in night scenes with streetlights and vehicle lights. To address this issue, we propose the POGS framework, which integrates depth and edge constraints to optimize the position parameters of gaussian spheres and reduce the interference of illumination intensity. Our approach first incorporates a pre-trained monocular depth estimation network to generate depth maps, which are used to constrain the positions of gaussian spheres. In addition, we introduce an Edge Optimization Loss and incorporate the Segment Anything Model (SAM) to generate contour maps for constraining the weights of gaussian spheres, enabling the model to pay more attention to the gaussian spheres of the objects composing the scene. Our method significantly enhances the reconstruction performance of 3DGS on datasets with low texture, high depth of field, and low illumination, improving the SSIM by 3.1%, and outperforms the state-of-the-art method RawNeRF (2% SSIM) in NeRF while maintaining real-time rendering speed.
Chaoyu Gao, Chengliang Wang 0002, Ji Liu 0006, Yonggang Luo, Bo Zheng 0007
SMC2
2025 Automated dual CNN-based feature extraction with SMOTE for imbalanced diabetic retinopathy classification
Danyal Badar Soomro, Chengliang Wang 0002, Mahmood Ashraf, Dina Abdulaziz Alhammadi, Shtwai Alsubai, Carlo Maria Medaglia, Nisreen Innab, Muhammad Umer 0001
Image Vis. Comput.2
2024 HMGS: Hybrid Model of Gaussian Splatting for Enhancing 3D Reconstruction with Reflections
Hengbin Zhang, Chengliang Wang 0002, Ji Liu 0006, Yonggang Luo, Lecheng Xie
ACCV (10)2
2024 CAVE-OCTA: A Vessel Expressions Controlled Diffusion Framework for Synthetic OCTA Image Generation
abstract
Optical coherence tomography angiography (OCTA) is a key technique for diagnosing retinal diseases. It clearly scans the detailed retinal vessel distribution, density, and other key biomarkers onto the images. Collecting and annotating high-quality OCTA data is challenging and requires significant manual costs. This paper proposes a fine-tuning framework called CAVE-OCTA (Controlled According to the Vessel Expressions) for generating synthetic en-face OCTA samples. It fine-tunes the stable diffusion model and controls the vascular distribution of the generated images through diversified retinal vessel expressions. We quantitatively evaluated the synthetic image quality and proved it to improve segmentation results on existing models. Additionally, we provided methods for creating several control prompt images. The code is available at https://github.com/ShellRedia/CAVE-OCTA.
Xinrun Chen, Mei Shen, Haojian Ning, Mengzhan Zhang, Chengliang Wang 0002
BIBM5
2024 Texture-Unet: A Texture-Aware Network for Bone Marrow Smear Whole-Slide Image Region of Interest Segmentation
abstract
Bone marrow smear cytology involves observing and analyzing the morphological features of bone marrow cells, and identifying regions of interest (ROI) where the cells are morphologically clear and evenly distributed is a crucial part of this process. However, existing deep learning methods for selecting ROI in whole-slide images (WSI) of bone marrow smears have overlooked the unique characteristics of the smears, particularly the texture information, resulting in inadequate performance. To overcome this issue, this paper proposes a texture-aware region of interest segmentation model (Texture-Unet) for WSI of bone marrow smears. To enhance the network’s ability to extract and perceive texture information, we specifically designed two modules, the Texture Extraction Module (TEM) and the Texture Deep Supervision Module (TDSM). We evaluated our method using a self-constructed dataset. The experimental results indicated a 4.64% improvement in terms of Intersection over Union (IoU), and a 3.04% improvement in terms of the Dice coefficient, compared to the baseline. The experimental results confirm that Texture-Unet performs well in terms of accuracy and efficiency for ROI segmentation in bone marrow smear WSI.
Chengliang Wang 0002, Zailin Yang, Xuelian Wu, Longrong Ran
ICASSP3
2024 An Accurate and Efficient Neural Network for OCTA Vessel Segmentation and a New Dataset
abstract
Optical coherence tomography angiography (OCTA) is a noninvasive imaging technique that can reveal high-resolution retinal vessels. In this work, we propose an accurate and efficient neural network for retinal vessel segmentation in OCTA images. The proposed network achieves accuracy comparable to other SOTA methods, while having fewer parameters and faster inference speed (e.g. 110x lighter and 1.3x faster than U-Net), which is very friendly for industrial applications. This is achieved by applying the modified Recurrent ConvNeXt Block to a full resolution convolutional network. In addition, we create a new dataset containing 918 OCTA images and their corresponding vessel annotations. The data set is semi-automatically annotated with the help of Segment Anything Model (SAM), which greatly improves the annotation speed. For the benefit of the community, our code and dataset can be obtained from https://github.com/nhjydywd/OCTA-FRNet.
Haojian Ning, Chengliang Wang 0002, Xinrun Chen
ICASSP2
2024 SAM-OCTA: A Fine-Tuning Strategy for Applying Foundation Model OCTA Image Segmentation Tasks
abstract
In the analysis of optical coherence tomography angiography (OCTA) images, the operation of segmenting specific targets is necessary. Existing methods typically train on supervised datasets with limited samples (approximately a few hundred), which can lead to overfitting. To address this, the low-rank adaptation technique is adopted for foundation model fine-tuning and proposed corresponding prompt point generation strategies to process various segmentation tasks on OCTA datasets. This method is named SAM-OCTA and has been experimented on the publicly available OCTA-500 dataset. While achieving state-of-the-art performance metrics, this method accomplishes local vessel segmentation as well as effective artery-vein segmentation, which was not well-solved in previous works. The code is available at: https://github.com/ShellRedia/SAM-OCTA.
Chengliang Wang 0002, Xinrun Chen, Haojian Ning
ICASSP1
2024 AIM-MIL: Adversarial Instance Mining for Robust Multi-instance Learning in Whole Slide Image Classification
Biyun Zhou, Chengliang Wang 0002, Hongqian Wang
ICONIP (9)2
2024 Weakly Supervised Segmentation of Plasma Cells in Bone Marrow via Scribble Annotations
abstract
Multiple Myeloma (MM) is a type of white blood cell cancer involving the abnormal multiplication of plasma cells. Performing nucleus and cytoplasm segmentation is a critical step in the analysis of MM. Compared to traditional semantic segmentation methods, weakly supervised semantic segmentation models based on scribble annotations require less labeling information, which can reduce the workload of manual annotation. However, due to the lack of explicit supervisory information for contours, these methods perform poorly at object boundaries, particularly in cell segmentation. Therefore, we propose a tri-branch network that incorporates a contour pseudo label supervision module for multitask learning and a contour attention mechanism to enhance the capability of segmenting cell contours. We constructed the Multiple Myeloma Plasma Cells scribble-annotated dataset based on SegPC-2021 and validated the effectiveness of our method on this dataset. Experimental results demonstrate that our method surpasses other weakly supervised approaches based on scribble annotations.
Chengliang Wang 0002, Longrong Ran, Zailin Yang
IJCNN2
2024 NFE-Net: Detection and Segmentation of Thyroid Nodules in Ultrasound Images Based on Nodule Feature Enhanced
abstract
Deep learning-based methods are commonly used for thyroid nodules detection and segmentation in ultrasound images, but the shape and size of the nodules vary greatly, there are also solid nodules that closely resemble the background and device-induced artifacts, making it difficult to accurately localize and segment the nodules. In this paper, NFE-Net is proposed to address the above difficulties, which designs receptive field enhancement path (RFEP) and texture-boundary guidance path (TBGP) in the neck of the Mask RCNN for nodule feature enhancement, so as to improve the localization and segmentation performance of network. RFEP enhances the learning ability of the network for nodule’s scale by introducing multi-scale receptive field enhancement module (RFEM); TBGP computes the channel correlation to extract the texture and boundary features of the nodule in information extraction module (IEM), further, uses true texture and boundary mask for deep supervised learning, finally fuses the upper and lower layers of the features by feature fusion block (FFB). Experimental results on the public thyroid datasets TN3K, DDTI show that our approach outperforms six state-of-the-art methods.
Zhaoxin Long, Supeng Yin, Chengliang Wang 0002, Hongqian Wang
IJCNN4
2024 Predict EGFR Mutation Status on CT Images Using Texture and Contour Enhanced Masked Autoencoders
abstract
The EGFR mutation status significantly influences targeted therapy for non-small cell lung cancer. In recent years, there has been significant progress in non-invasive EGFR mutation status prediction studies based on chest CT images. However, these studies commonly rely on extensive private data for training rather than small-scale publicly available dataset, thereby failing to overcome the dependency on high-cost large-scale annotated data. Additionally, these studies generally neglect the texture and contour features with strong discriminative power, leading to insufficient performance. This paper proposes a two-stage framework for EGFR mutation status prediction. Initially, we utilize self-supervised Masked Autoencoders (MAE) to pre-train the encoder on in-domain chest CT images, overcoming the problem of insufficient annotated data to reduce the dependency on high-cost annotated data. Subsequently, fine-tune the encoder to make it suitable for downstream EGFR prediction. Simultaneously, we propose texture and contour enhanced MAE (TCMAE), designing Multi-layer Features Aggregation Module (MFAM) to fully exploit multi-layer semantic features, introducing Texture and Contour Prediction Module (TCPM) to enhance the model’s capability in extracting texture and contour features through multitask learning, utilizing Modified Spectral Block (MSB) to adjust the weighting between high and low frequency features. Experiments demonstrate that despite using only a small-scale public dataset, NSCLC-Radiogenomics, the proposed method still achieves high accuracy.
Yuping Peng, Zhongshi He, Chengliang Wang 0002, Hongqian Wang
IJCNN4
2024 CIS-Net: An End-to-end Chromosome Instance Segmentation Method Based on Disturbance Features Deactivation and Strengthened Disparity
abstract
Chromosomes are the carriers of human genetic information. Abnormal numbers of chromosomes can cause a variety of diseases. Karyotype analysis is an important means to assist in the diagnosis of chromosome abnormalities. Current deep learning-based karyotype analysis methods first segment single chromosomes on metaphase images and then classify each single chromosome. There are two problems with these methods. First, large intra-class differences and small interclass differences of chromosomes can easily lead to classification errors. Second, chromosome segmentation and classification are independent of each other, making overall optimization impossible. Therefore, we proposed CIS-Net. Disturbance Features Deactivation Module(DFDM) is first designed to improve the model’s key feature selection capabilities; then Strengthened Disparity Loss(SD Loss) is designed to enhance the ability to distinguish easily confused chromosome classes. Meanwhile, we define chromosome karyotype analysis as a 24-class instance segmentation problem, realizing end-to-end chromosome segmentation and classification on the whole metaphase image. Experiments on the AutoKary2022 dataset demonstrate that CISNet performs well compared to existing models, achieving the best mAP of 90.7%.
Yulin Tan, Chengliang Wang 0002, Zailin Yang, Longrong Ran
IJCNN3
2024 Anomaly Detection in Chest X-ray Images with Adversarial Masked Autoencoder
abstract
Chest X-ray is the most commonly used detection method for lung diseases, but manual screening often has omissions, so computer-aided diagnosis of chest X-ray abnormalities is necessary. However, since abnormal data relies on expert annotation and is difficult to obtain, unsupervised anomaly detection (UAD) using only normal data has become the focus of attention in the field of medical images. Current UAD methods based on reconstruction often use reconstruction error of original image and reconstructed image as the anomaly score, but the strong reconstruction ability of the autoencoder results in small abnormal image reconstruction error, which is similar to the normal image reconstruction error, making the detection results not satisfactory. Therefore, we proposed CGMAE. In the training stage, masked autoencoder is used as the generator to reconstruct the image, and a discriminator with the reconstructed image and the original image as input is added at the end. Through adversarial learning, the discriminator learns the distribution of normal data. In the testing stage, Gaussian distribution is used to construct the anomaly score, increasing the gap between normal and abnormal data, which is conducive to separating abnormal data from normal data. At the same time, it was found that the current reconstruction models for chest X-ray anomaly detection did not consider the different importance of chest X-ray foreground and background when reconstructing images better. Therefore, we designed a regional weighted loss to enable the generator to reconstruct high-resolution chest X-ray images and enhance the discriminator’s ability to learn data distribution. Experiments on two public datasets, Zhanglab dataset and Chexpert dataset, show that CGMAE exceeds SOTA by 1.41% and 0.98% in AUC metrics respectively.
Yehong Tong, Zhongshi He, Chengliang Wang 0002
IJCNN4
2024 Dual-Branch Retinal OCT Anomaly Detection Based on Knowledge Distillation and Reconstruction
abstract
Currently, anomaly detection methods for retinal OCT images can be mainly divided into two types: reconstruction-based and knowledge distillation-based. The former has weak inter-class descriptiveness and unclear feature boundaries as it is trained with normal images only, and without comparing with abnormal images. The latter is insensitive to rare diseases and lacks a pretrained model with high-performance. In previous works, these two methods are used independently, thus unable to address the innate limitations of the model effectively. Therefore, we proposed a dual-branch model named DRTNet_AD. The two branches are connected through a student network (Encoder) serving as an intermediary bridge. The knowledge distillation network branch learns descriptive features through a pretrained model, while the reconstruction network branch enhances the model’s sensitivity to rare diseases by learning the distribution of normal images. Furthermore, in order to obtain a pretrained model with good feature extraction and descriptive abilities, we proposed a strategy to construct pseudo-anomaly and a layer structure context-aware module (LCAM). DRTNet_AD achieves the SOTA performance on public dataset SpectralisOCT with AUC scores of 98.25%
Minghui Zhai, Zhongshi He, Chengliang Wang 0002
IJCNN4
2024 An FPGA-based kNN Seach Accelerator for point cloud registration
abstract
Point cloud registration assumes a crucial role in several fields, including 3D reconstruction and pose estimation. The prevalent technique for point cloud registration is Iterative Closest Point (ICP). Nonetheless, the k-nearest neighbor (kNN) search, an essential component of the ICP process, often falls short in meeting real-time requirements due to its substantial time consumption. Consequently, extensive efforts have been dedicated to accelerating the kNN search within the ICP framework. This study proposes an FPGA-based kNN Search Accelerator, leveraging an improved LSH approach to expedite point cloud access and search. Experimental results underscore its superiority with a 120x and 15x speed-up compared to CPU and GPU implementations of the kNN process, respectively. Remarkably, the kNN search completes in a mere 0.64 ms, surpassing the performance of prior works.
Chengliang Wang 0002, Zhetong Huang, Ao Ren
ISCAS1
2024 Visual Navigation by Fusing Object Semantic Feature
abstract
The key of object goal visual navigation is to learn the spatial relationships between environmental objects and assess their semantic correlations with the target object. We propose an end-to-end visual navigation model based on deep reinforcement learning, called G2SNet, which consists of two feature maps and a specialized fusion feature network: GloVe Feature Map (GFM), Sbbox Feature Map (SFM), and GloVe fusion Network (GNet). GFM represents the position and the semantic information of the objects contained in the observation image, which addresses the issue of the interference in the target recognition caused by the complex background information in the observed image. SFM provides the object sizes in the field of view to assist the distance judgment. GNet relies entirely on network learning to compute the semantic correlations and spatial positional relationships among objects in the environment, enabling the agent to possess better generalization capabilities. This allows learning of spatial relationships between objects in GFM. Experiments on AI2-THOR demonstrate the effectiveness of our proposed three new structures, and the average SPL of the four known scenarios is increased by 16.6%.
Chengliang Wang 0002, Zhongshi He, Hongqian Wang
SMC3
2024 Enhancing Autofocus Performance through Predictive Motion-Targeting and Self-Attention in a Deep Reinforcement Learning Framework
abstract
In focusing tasks on moving targets, traditional methods that rely on maximizing contrast struggle to capture moving objects due to insufficient focusing speed. Deep learning-based methods have attempted to directly predict the optimal focal length for the target; however, due to low prediction accuracy, they often lead to out-of-focus situations when capturing moving objects. In recent years, some approaches have utilized reinforcement learning to automatically explore focal length adjustment patterns, thus achieving better results than traditional methods. However, these approaches have not considered the motion characteristics of the targets, leading to a need for further improvement in focusing performance. To overcome these limitations, we introduce a motion-based feature and deep reinforcement learning-driven autofocus algorithm named MF-DRLAF (Motion Features based Deep Reinforcement Learning Autofocus Model) for moving targets. This novel method tracks the object, predicts its motion state through feature extraction, and uses deep reinforcement learning to dynamically adjust the focus. We utilize a self-attention mechanism to adaptively learn various motion patterns and employ a feature pool structure to enhance processing efficiency. Experiments and real-world testing on a Google Pixel3 demonstrate that our approach significantly enhances autofocus performance on moving objects, highlighting its potential for broader imaging applications. This approach offers a promising direction for future development in autofocus technology.
Xiaolin Wei, Ruilong Yang, Chengliang Wang 0002, Hongqian Wang
SMC4
2024 Improve Deep Learning Autofocus with Depth Information Supervision and Current Focal Distance Cues
abstract
Traditional autofocus methods search for the optimal focal distance (FD) by evaluating image quality from focal stacks, resulting in time-consuming focusing processes. Recently, deep learning has being adopted for single-shot autofocus methods, which can predict the optimal FD directly from a single input image. However, these methods often suffer from low prediction accuracy due to the lack of global features and structured global supervisory information, as they rely solely on the image's region of interest (ROI) as input and a single value for supervision. We propose a deep learning network named MPFS (Multi-Head Network with Per-Pixel Focal Distance Supervision), which takes a full-frame photograph as input and uses the optimal focal distance per pixel for supervision, this method effectively addresses the issues of missing global features and insufficient supervisory information by leveraging these enhancements. Additionally, the network integrates current camera focal distance information to mitigate the scale ambiguity caused by the lack of absolute scale information. To validate the effectiveness of the proposed method, we designed an experiment using a dataset annotated with optimal FD per pixel. Experimental results on this dataset indicate that our approach achieves a 0.22 decrease in the Mean Absolute Error (Mae) metric compared to the state-of-the-art models, with improvements of 0.02 and 0.004 in$\boldsymbol{d}_{\mathbf{1}}$and$\boldsymbol{d}_{\mathbf{2}}$metrics.
Xiaolin Wei, Ruilong Yang, Chengliang Wang 0002, Hongqian Wang
SMC4
2024 SMDNet: A Pulmonary Nodule Classification Model Based on Positional Self-Supervision and Multi-Direction Attention
abstract
Accurate classification of pulmonary nodules holds importance in the early diagnosis of lung cancer. Unlike 2D models, 3D models can simultaneously utilize multiple slices as input to capture features. However, 3D models face challenges in capturing nodule features in different directions and discerning feature differences in various positions of computed tomography (CT). We introduce a pulmonary nodule classification model, SMDNet. Firstly, a multi-direction attention is proposed to capture nodule features from sagittal, coronal, and axial axes. Secondly, distinct labels are assigned to the cubes at different cropping positions from CT for binary classification to capture local differences. Besides, gradient boosting decision tree (GBDT) is employed to combine shallow features with deep features to improve accuracy. Comparative experimental results on the largest publicly available dataset of pulmonary nodules, LIDC-IDRI, showed that SMDNet achieves a 4.81% improvement in accuracy under identical data processing.
Chengliang Wang 0002, Hongqian Wang
SMC2
2024 WS-SSD: Achieving faster 3D object detection for autonomous driving via weighted point cloud sampling
Chengliang Wang 0002, Zhuo Zeng
Expert Syst. Appl.2
2024 De-redundancy in wireless capsule endoscopy video sequences using correspondence matching and motion analysis
Libin Lan, Chunxiao Ye, Chao Liao, Chengliang Wang 0002
Multim. Tools Appl.4
2024 Selective transfer subspace learning for small-footprint end-to-end cross-domain keyword spotting
Chengliang Wang 0002, Zhuo Zeng
Speech Commun.2
2023 DB-UNet: MLP Based Dual Branch UNet for Accurate Vessel Segmentation in OCTA Images
abstract
Optical coherence tomography angiography (OCTA) is a new non-invasive imaging technology that has been widely used in clinical practice. Automatic segmentation of retina vessels in OCTA images helps to improve the efficiency of disease diagnosis. However, due to the slender and tiny structure of retina vessels, classical deep learning segmentation methods, such as UNet and some of its variants, cannot handle it very accurately. In this work, we propose a dual branch UNet (DB-UNet), which has a pure-convolutional branch to extract detailed features such as microvessels, and a UNet branch to extract high-level features. The final output of the network is synthesized from the outputs of the two branches. We also design and apply a multilayer perceptron (MLP) block to further improve the performance of our model. Experiments on two OCTA vessel segmentation datasets show that the proposed method has better segmentation performance than existing methods.
Chengliang Wang 0002, Haojian Ning, Xinrun Chen
ICASSP1
2023 Adaptive Focal Inverse Distance Transform Maps for Cell Recognition
Chengliang Wang 0002, Zailin Yang, Longrong Ran
ICONIP (6)3
2023 TCNet: Texture and Contour-Aware Model for Bone Marrow Smear Region of Interest Selection
Chengliang Wang 0002, Zailin Yang, Longrong Ran
ICONIP (10)1
2023 Distributed Cooperative Search Algorithm with Information Screening
abstract
This paper focuses on distributed cooperative search using multiple Unmanned Ground Vehicles (UGVs) for a dynamic target. To enhance search efficiency and optimize resource usage, we propose improvements in the search algorithm framework, including advancements in search map creation and updates, distributed information fusion, and collaborative decision-making. In contrast to traditional approaches, we introduce environmental uncertainty, which correlates with cell detection intervals, reflecting information importance. In the information fusion stage, our innovative mechanism screens information based on decision horizon and environmental uncertainty, while considering communication constraints to minimize redundant data transmission. Additionally, in the collaborative decision stage, the utility function was optimized by considering the cumulative environmental uncertainties for generated path covering cells, thereby enhancing search area coverage and cell revisit probability. Experimental results with four UGVs demonstrate the superiority of our algorithm, achieving 11.7% search efficiency improvement, 11.8% trajectory length reduction and 45.6% computation time reduction compared with data fusion algorithm, validating our approach’s effectiveness in practical scenarios.
Chengliang Wang 0002
MSN1
2023 An Authentication Algorithm for Sets of Spatial Data Objects
Chengliang Wang 0002, Xiaobing Hu, Hongwen Zhou, Yanai Wang
SecureComm (1)2
2023 SVM-based subspace optimization domain transfer method for unsupervised cross-domain time series classification
Chengliang Wang 0002, Zhuo Zeng
Knowl. Inf. Syst.2
2023 FedMDS: An Efficient Model Discrepancy-Aware Semi-Asynchronous Clustered Federated Learning Framework
abstract
Federated learning (FL) is an emerging distributed machine learning paradigm that protects privacy and tackles the problem of isolated data islands. At present, there are two main communication strategies of FL: synchronous FL and asynchronous FL. The advantages of synchronous FL are the high precision and easy convergence of the model. However, this synchronous communication strategy has the risk of the straggler effect. Asynchronous FL has a natural advantage in mitigating the straggler effect, but there are threats of model quality degradation and server crash. In this paper, we propose a model discrepancy-aware semi-asynchronous clustered FL framework,FedMDS, which alleviates the straggler effect by 1) a clustered strategy based on the delay and direction of the model update and 2) a synchronous trigger mechanism that limits the model staleness.FedMDSleverages the clustered algorithm to reschedule the clients. Each group of clients performs asynchronous updates until the synchronous update mechanism based on the model discrepancy is triggered. We evaluateFedMDSbased on four typical federated datasets in a non-IID setting and compareFedMDSto the baselines. The experimental results show thatFedMDSsignificantly improves average test accuracy by more than$+9.2\%$on the four datasets compared toTA-FedAvg. In particular,FedMDSimproves absolute Top-1 test accuracy by$+37.6\%$on FEMNIST compared toTA-FedAvg. The frequency of the average synchronization waiting time ofFedMDSis significantly lower than that ofTA-FedAvgon all datasets. Moreover,FedMDScan improve the accuracy and alleviate the straggler effect.
Yu Zhang 0184, Duo Liu 0002, Moming Duan, Xianzhang Chen, Ao Ren, Yujuan Tan, Chengliang Wang 0002
IEEE Trans. Parallel Distributed Syst.8
2022 VEA: An FPGA-Based Voxel Encoding Accelerator for 3D Object Detection with LiDAR
abstract
Voxel-based 3D object detection methods have been applied in various applications such as autonomous driving, robot navigation, and Augmented Reality. However, the sparse and unstructured characteristics of the point cloud and voxels prevent high-performance voxel encoding and usually require generalized platforms, such as CPUs. In this paper, an FPGAbased Voxel Encoding Accelerator (VEA) is proposed, which contains a generalized voxel generator and a feature extender. The generalized voxel generator decouples the point storage and voxel information storage, leading to high-speed voxelization and low memory consumption. The feature extender can efficiently extract the geometric information of the voxels and extend the features of the points. Based on the proposed VEA, an FPGA-based 3D object detection accelerator is implemented, and experimental results show that the proposed VEA can outperform prior studies by 19× faster in voxelization and 1.3×~ 9.6× faster in object detection.
Ao Ren, Yujuan Tan, Zhetong Huang, Chengliang Wang 0002, Xianzhang Chen, Duo Liu 0002
ICCD6
2022 Triplet Confidence for Robust Out-of-vocabulary Keyword Spotting
abstract
Keyword Spotting (KWS) is a task that detects predefined keywords in a stream of audio. Although state-of-the-art deep neural networks perform well on KWS, they are not robust against out-of-vocabulary (OOV) samples because of their over-reliance on labeled data and the great imbalance between in-vocabulary (IV) and OOV data resulting from the infinity of OOV data. Besides, some end-to-end(E2E) KWS try to treat it as a multi-classification task, rejecting most OOV samples through posterior processing, but they cannot ensure the classifier can always be sure that its judgment is correct and keywords are usually too short to extract many representative features. To address these issues, we introduce a KWS model that can keep robustness on OOV samples by learning confidence estimates of the model. Confidence estimation is output by the self-attentional confidence branch, which can focus on single keyword in context. And we propose a loss function, named Triplet Correction Loss, which learning confidence estimation to improve the reliability of the model without relying on labeled data. Compared with the state-of-the-art methods, our proposed network increases 1.80%(V2-25) and 7.63%(Librispeech) accuracy on OOV samples, while keeping the high accuracy on IV dataset. We also provide ablation study to prove that our method is effective.
Chengliang Wang 0002, Yujie Hao, Chao Liao
ISCAS1
2022 Federated learning with workload-aware client scheduling in heterogeneous systems
Duo Liu 0002, Moming Duan, Yu Zhang 0184, Ao Ren, Xianzhang Chen, Yujuan Tan, Chengliang Wang 0002
Neural Networks8
2022 FRL: Fast and Reconfigurable Accelerator for Distributed Sound Source Localization
abstract
Sound source localization (SSL) has been widely applied in industrial and civil fields. And with the development of wearable devices and the Internet of Things (IoT), it is attractive to deploy the SSL system onto embedded and portable devices. However, the software-based SSL system causes excessive response delay and is often affected by environmental noise. To overcome this obstacle, we propose the fast and precise localization (FPL) algorithm for distributed SSL systems. It combines the benefits of both time difference of arrival (TDOA) and steered response power (SRP) methods, and thus it is able to localize sound sources fast and precisely. To further improve the localization speed, we propose the fast and reconfigurable localization (FRL) accelerator, which is an algorithm-hardware co-designed SSL accelerator. It adopts multiple distributed localization nodes for higher localization precision and higher robustness to environmental interference, and it can be configured into either the fast or precise mode to adapt to various environments. Experimental evaluations show that our proposed FPL algorithm can achieve high localization speed and precision, and the field-programmable gate array (FPGA)-based FRL accelerator outperforms the software implementation by$48.6\times $and outperforms the prior FPGA-based SSL accelerators by$20\times \sim 838.2\times $.
Chengliang Wang 0002, Heping Liu, Zhihai Zhang, Xianzhang Chen, Yujuan Tan, Duo Liu 0002, Ao Ren
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2022 SENTunnel: Fast Path for Sensor Data Access on Automotive Embedded Systems
Rongwei Zheng, Xianzhang Chen, Duo Liu 0002, Jiapin Wang, Ao Ren, Chengliang Wang 0002, Yujuan Tan
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.7
2021 Forseti: An Efficient Basic-block-level Sensitivity Analysis Framework Towards Multi-bit Faults
abstract
The per-instruction sensitivity analysis framework is developed to evaluate the resiliency of a program and identify the segments of the program needing protection. However, for multi-bit hardware faults, the per-instruction sensitivity analysis frameworks can cause large overhead for redundant analyses. In this paper, we propose a basic-block-level sensitivity analysis framework, Forseti, to reduce the analysis overhead in analyzing impacts of modern microprocessors' multi-bit faults on programs. We implement Forseti in LLVM and evaluate it with five typical workloads. Extensive experimental results show that Forseti can achieve more than 90% sensitivity classification accuracy and 6.16× speedup over instruction-level analysis.
Jinting Ren, Xianzhang Chen, Duo Liu 0002, Moming Duan, Renping Liu 0002, Chengliang Wang 0002
DATE6
2021 Detection of Retinal Vascular Bifurcation and Crossover Points in Optical Coherence Tomography Angiography Images Based on CenterNet
Chengliang Wang 0002, Shitong Xiao, Chao Liao
ICONIP (6)1
2021 FedSAE: A Novel Self-Adaptive Federated Learning Framework in Heterogeneous Systems
abstract
Federated Learning (FL) is a novel distributed machine learning which allows thousands of edge devices to train model locally without uploading data concentrically to the server. But since real federated settings are resource-constrained, FL is encountered with systems heterogeneity which causes a lot of stragglers directly and then leads to significantly accuracy reduction indirectly. To solve the problems caused by systems heterogeneity, we introduce a novel self-adaptive federated framework FedSAE which adjusts the training task of devices automatically and selects participants actively to alleviate the performance degradation. In this work, we 1) propose FedSAE which leverages the complete information of devices' historical training tasks to predict the affordable training workloads for each device. In this way, FedSAE can estimate the reliability of each device and self-adaptively adjust the amount of training load per client in each round. 2)combine our framework with Active Learning to self-adaptively select participants. Then the framework accelerates the convergence of the global model. In our framework, the server evaluates devices' value of training based on their training loss. Then the server selects those clients with bigger value for the global model to reduce communication overhead. The experimental result indicates that in a highly heterogeneous system, FedSAE converges faster than FedAvg, the vanilla FL framework. Furthermore, FedSAE outperforms than FedAvg on several federated datasets - FedSAE improves test accuracy by 26.7% and reduces stragglers by 90.3% on average.
Moming Duan, Duo Liu 0002, Yu Zhang 0184, Ao Ren, Xianzhang Chen, Yujuan Tan, Chengliang Wang 0002
IJCNN8
2021 CSAFL: A Clustered Semi-Asynchronous Federated Learning Framework
abstract
Federated learning (FL) is an emerging distributed machine learning paradigm that protects privacy and tackles the problem of isolated data islands. At present, there are two main communication strategies of FL: synchronous FL and asynchronous FL. The advantages of synchronous FL are that the model has high precision and fast convergence speed. However, this synchronous communication strategy has the risk that the central server waits too long for the devices, namely, the straggler effect which has a negative impact on some time-critical applications. Asynchronous FL has a natural advantage in mitigating the straggler effect, but there are threats of model quality degradation and server crash. Therefore, we combine the advantages of these two strategies to propose a clustered semi-asynchronous federated learning (CSAFL) framework. We evaluate CSAFL based on four imbalanced federated datasets in a non-IID setting and compare CSAFL to the baseline methods. The experimental results show that CSAFL significantly improves test accuracy by more than +5% on the four datasets compared to TA-FedAvg. In particular, CSAFL improves absolute test accuracy by +34.4% on non-IID FEMNIST compared to TA-FedAvg.
Yu Zhang 0184, Moming Duan, Duo Liu 0002, Ao Ren, Xianzhang Chen, Yujuan Tan, Chengliang Wang 0002
IJCNN8
2021 Improving Efficiency and Lifetime of Logic-in-Memory by Combining IMPLY and MAGIC Families
Minhui Zou, Junlong Zhou, Jin Sun 0001, Chengliang Wang 0002, Shahar Kvatinsky
J. Syst. Archit.4
2021 Bridging Mismatched Granularity Between Embedded File Systems and Flash Memory
abstract
The mismatch between logical and physical I/O granularity inhibits the deployment of embedded file systems. Most existing embedded file systems manage logical space with a small unit, which is no longer the case of the flash operation granularity. Manually enlarging the logical I/O granularity of file systems requires enormous transplanting efforts. Moreover, large logical pages signify the write amplification problem, which turns to severe space consumption and performance collapse. This article designs a novel storage middleware, NV-middle, for legacy-embedded file systems with large-capacity flash memories. Legacy-embedded storage schemes can be smoothly transplanted into new platforms with different hardware read/write granularity. Moreover, the legacy optimization schemes can be maximally reserved, without inducing write amplification problems. We implement NV-middle with the state-of-the-art embedded file system, YAFFS2. Comprehensive evaluations show that NV-middle can achieve times of performance improvement over manually transplanted YAFFS2 with various workloads.
Runyu Zhang 0002, Duo Liu 0002, Zhaoyan Shen, Xiongxiong She, Chaoshu Yang, Xianzhang Chen, Yujuan Tan, Chengliang Wang 0002
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.8
2021 Multilevel Privacy Controlling Scheme to Protect Behavior Pattern in Smart IoT Environment
abstract
Traditional approaches generally focus on the privacy of user’s identity in a smart IoT environment. Privacy of user’s behavior pattern is an important research issue to address smart technology towards improving user’s life. User’s behavior pattern consists of daily living activities in smart IoT environment. Sensor nodes directly interact with activities of user and forward sensing data to service provider server (SPS). While availing the services provided by a server, users may lose privacy since the untrusted devices have information about user’s behavior pattern and it may share data with adversary. In order to resolve this problem, we propose a multilevel privacy controlling scheme (MPCS) which is different from traditional approaches. MPCS is divided into two parts: (i) behavior pattern privacy degree (BehaviorPrivacyDeg), which works as follows: firstly, frequent pattern mining‐based time‐duration algorithm (FPMTA) finds the normal pattern of activity by adopting unsupervised learning. Secondly, patterns compact algorithm (PCA) is proposed to store and compact the mined pattern in each sensor device. Then, abnormal activity detection time‐duration algorithm (AADTA) is used by current triggered sensors, in order to compare the current activity with normal activity by computing similarity among them; (ii) multilevel privacy design model: we have divided privacy of users into four levels in smart IoT environment, and by using these levels, the server can configure privacy level for users according to their concern. Multilevel privacy design model consists of privacy‐level configuration protocol (PLCP) and activity design model. PLCP provides fine privacy controls to users while enabling users to set privacy level. In PLCP, we introduce level concern privacy algorithm (LCPA) and location privacy algorithm (LPA), so that adversary could not damage the data of user’s behavior pattern. Experiments are performed to evaluate the accuracy and feasibility of MPCS in both simulation and real‐case studies. Results show that our proposed scheme can significantly protect the user’s behavior pattern by detecting abnormality in real time.
Muhammad Mehran Arshad Khan, Muhammad Awais Javeed, Muhammad Umar Farooq 0002, Adeel Akram, Chengliang Wang 0002
Wirel. Commun. Mob. Comput.6
2020 Security Enhancement for RRAM Computing System through Obfuscating Crossbar Row Connections
abstract
Neural networks (NN) have gained great success in visual object recognition and natural language processing, but this kind of data-intensive applications requires huge data movements between computing units and memory. Emerging resistive random-access memory (RRAM) computing systems have demonstrated great potential in avoiding the huge data movements by performing matrix-vector-multiplications in memory. However, the nonvolatility of the RRAM devices may lead to potential stealing of the NN weights stored in crossbars and the adversary could extract the NN models from the stolen weights. This paper proposes an effective security enhancing method for RRAM computing systems to thwart this sort of piracy attack. We first analyze the theft methods of the NN weights. Then we propose an efficient security enhancing technique based on obfuscating the row connections between positive crossbars and their pairing negative crossbars. Two heuristic techniques are also presented to optimize the hardware overhead of the obfuscation module. Compared with existing NN security work, our method eliminates the additional RRAM writing operations used for encryption/decryption, without shortening the lifetime of RRAM computing systems. The experiment results show that the proposed methods ensure the trial times of brute-force attack are more than (16!)17and the classification accuracy of the incorrectly extracted NN models is less than 20%, with minimal area overhead.
Minhui Zou, Zhenhua Zhu 0002, Yi Cai 0003, Junlong Zhou, Chengliang Wang 0002, Yu Wang 0002
DATE5
2020 Unified-TP: A Unified TLB and Page Table Cache Structure for Efficient Address Translation
abstract
To improve the performance of address translation in applications with large memory footprints, techniques, such as hugepages and HW coalescing, are proposed to increase the coverage of limited hardware translation entries by exploiting the contiguous memory allocation to lower Tanslation Lookaside Buffer (TLB) miss rate. Furthermore, Page Table Caches (PTCs) are proposed to store the upper-level page table entries to reduce the TLB miss handling latency. Both increasing TLB coverage and reducing TLB miss handling latency have proved to be effective in speeding up address translation, to a certain extent. Nevertheless, our preliminary studies suggest that the structural separation between TLBs and PTCs in existing computer systems makes these two methods less effective because they are exclusively used in TLBs and PTCs respectively. In particular, the separate structures cannot dynamically adjust their sizes according to the workloads, resulting in low resource utilization and inefficient address translation. To address these issues, we propose a unified structure, called Unified - Tp,which stores PTC and TLB entries together. Besides, Our modified LRU algorithm helps identify the cold TLB and PTC entries and dynamically adjust the numbers of TLB and PTC entries to adapt to different workloads. Furthermore, we introduce a scheme of parallel search when receiving memory access requests. Our experimental results show that Unified-TP can reduce the numbers of TLB misses by an average of 35.69 % and improve the performance by an average of 11.12% compared with separately structured TLBs and PTCs.
Zhulin Ma, Yujuan Tan, Hong Jiang 0001, Zhichao Yan 0001, Duo Liu 0002, Xianzhang Chen, Qingfeng Zhuge, Edwin H.-M. Sha, Chengliang Wang 0002
ICCD9
2020 Making Inconsistent Components More Efficient For Hybrid B+Trees
abstract
The emergence of non-volatile memories (NVMs) provides opportunities for efficiently manipulating and storing tree-based indexes. Traditional tree structures fail to take full advantage of NVMs due to the considerable shifts of entries. Those redundant write activities induce severe performance collapse and power consumption, which is unacceptable for embedded systems. Advanced schemes, such as NV-Tree, remain only leaf nodes in NVM to lighten the burden of writes. However, NV-Tree suffers from frequency and in-DRAM structures. In this paper, we develop a novel tree structure, Marionette-Tree, to address these issues. We follow the hybrid layout of non-leaf/leaf nodes and adopt bitmap-based leaf nodes for minimum NVM writes. We aggregate the pointers of internal nodes into a shadow array, allowing recording more keys in an internal node. We also design a delayed-split-migration scheme to minimize the management overhead of the shadow array. Extensive evaluations demonstrate that Marionette-Tree can achieve 1.38 × and 5.73 × insertion speedup over two state-of-the-arts.
Xiongxiong She, Chengliang Wang 0002, Fenghua Tu
ICPADS2
2020 An Adaptive Fusion Model Based on Kalman Filtering and LSTM for Fast Tracking of Road Signs
abstract
The detection and tracking of road signs plays a critical role in various autopilot application. Utilizing convolutional neural networks(CNN) mostly incurs a big run-time overhead in feature extraction and object localization. Although Klaman filter(KF) is a commonly-used tracker, it is likely to be impacted by omitted objects in the detection step. In this paper, we designed a high-efficient detector that combines ThunderNet and Region Growing Detector(RGD) to detect road signs, and built a fusion model of long short term memory network (LSTM) and KF in the state estimation and the color histogram. The experimental results demonstrate that the proposed method improved the state estimation accuracy by 6.4% and enhanced the Frames Per Second(FPS) to 41.
Chengliang Wang 0002, Chao Liao
ICPR1
2020 Helena: Real-time Contact-free Monitoring of Sleep Activities and Events around the Bed
abstract
In this paper, we introduce a novel real-time and contact-free sensor system, Helena, that can be mounted on a bed frame to continuously monitor sleep activities (entry/exit of bed, movement, and posture changes), vital signs (heart rate and respiration rate), and falls from bed in a real-time and pervasive computing manner. The smart sensor senses bed vibrations generated by body movements to characterize sleep activities and vital signs based on advanced signal processing and machine learning methods. The device can provide information about sleep patterns, generate real-time results, and support continuous sleep assessment and health tracking. The novel method for detecting falls from bed has not been attempted before and represents a life-changing for high-risk communities, such as seniors. Comprehensive tests and validations were conducted to evaluate system performances using FDA approved and wearable devices. Our system has an accuracy of 99.5% detecting on-bed (entries), 99.73% detecting off-bed (exits), 97.92% detecting movements on the bed, 92.08% detecting posture changes, and 97% detecting falls from bed. The system estimation of heart rate (HR) ranged ±2.41 beats-per-minute compared to Apple Watch Series 4, while the respiration rate (RR) ranged ±0.89 respiration-per-minute compared to an FDA oximeter and a metronome.
José Clemente, Maria Valero, Fangyu Li 0002, Chengliang Wang 0002, Wen-Zhan Song 0001
PerCom4
2020 Fine-tuning Pre-trained Convolutional Neural Networks for Gastric Precancerous Disease Classification on Magnification Narrow-band Imaging Images
Chengliang Wang 0002, Jianying Bai, Guobin Liao
Neurocomputing2
2018 A New Asymmetric User Similarity Model Based on Rational Inference for Collaborative Filtering to Alleviate Cold Start Problem
Chengliang Wang 0002
ICIC (1)2
2018 Transfer Learning with Convolutional Neural Network for Early Gastric Cancer Classification on Magnifiying Narrow-Band Imaging Images
abstract
Immediate diagnosis and treatment of early gastric cancer (EGC) can efficiently improve the survival of gastric cancer (GC). Magnification endoscopy with narrow-band imaging (M-NBI) as a kind of main tool is widely applied in hospital to detect EGC by illustrating abnormal vascellum morphologies. In this paper, transfer learning by fine-tuning deep convolutional neural networks (CNNs) is applied to automatically classify M-NBI images into two groups: normal gastric images and EGC images. Moreover, this paper explores how transfer learning affects the classification performance from four aspects: training dataset, basic architectures of the deep CNN, the number of fine-tuned network layers and size of the network input image; this paper also gives some guidances for later researches in this area. VGG16, InceptionV3 and InceptionResNetV2 are selected to help us accomplish the M - NBI image classification task. Experimental results show that transfer learning of deep CNN features performs better than traditional handcraft methods. And the top accuracy, sensitivity and specificity are 0.985, 0.981 and 0.989, respectively, obtained by fine-tuning Inception V3.
Chengliang Wang 0002, Zhuo Zeng, Jianying Bai, Guobin Liao
ICIP2
2016 DPHK: real-time distributed predicted data collecting based on activity pattern knowledge mined from trajectories in smart environments
Chengliang Wang 0002, Ya-Yun Peng, Debraj De, Wen-Zhan Song 0001
Frontiers Comput. Sci.1
2013 Trajectory mining from anonymous binary motion sensors in Smart Environment
Chengliang Wang 0002, Debraj De, Wen-Zhan Song 0001
Knowl. Based Syst.1
2013 Object tracking under low signal-to-noise-ratio with the instantaneous-possible-moving-position model
Chengliang Wang 0002, Xiaoming Huo
Signal Process.1
2012 FindingHuMo: Real-Time Tracking of Motion Trajectories from Anonymous Binary Sensing in Smart Environments
abstract
In this paper we have proposed and designed FindingHuMo (Finding Human Motion), a real-time user tracking system for Smart Environments. FindingHuMo can perform device-free tracking of multiple (unknown and variable number of) users in the Hallway Environments, just from non-invasive and anonymous (not user specific) binary motion sensor data stream. The significance of our designed system are as follows: (a) fast tracking of individual targets from binary motion data stream from a static wireless sensor network in the infrastructure. This needs to resolve unreliable node sequences, system noise and path ambiguity, (b) Scaling for multi-user tracking where user motion trajectories may crossover with each other in all possible ways. This needs to resolve path ambiguity to isolate overlapping trajectories, FindingHumo applies the following techniques on the collected motion data stream: (i) a proposed motion data driven adaptive order Hidden Markov Model with Viterbi decoding (called Adaptive-HMM), and then (ii) an innovative path disambiguation algorithm (called CPDA). Using this methodology the system accurately detects and isolates motion trajectories of individual users. The system performance is illustrated with results from real-time system deployment experience in a Smart Environment.
Debraj De, Wen-Zhan Song 0001, Mingsen Xu, Chengliang Wang 0002, Diane J. Cook, Xiaoming Huo
ICDCS4
2011 Robust stability of impulsive Takagi-Sugeno fuzzy systems with parametric uncertainties
Xiaohong Zhang 0002, Chengliang Wang 0002, Dong Li 0009, Dan Yang 0001
Inf. Sci.2
2010 Exploratory Factor Analysis Approach for Understanding Consumer Behavior toward Using Chongqing City Card
Chengliang Wang 0002
ADMA (2)2
2005 An On-line Intelligent Recommendation System for Digital Products Using Fuzzy Logic
Yukun Cao, Chengliang Wang 0002
WISE3