Qilei Chen

dblp:190/9032 · DBLP profile ↗
← Back
17ranked-venue papers
2as first author
11since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 2 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021
YearPublicationVenuePosition
2025 One-stage Framework for Thyroid Nodule Detection with Mixup and Negative Sample Utilization
abstract
Ultrasound imaging plays a significant role in the early diagnosis of thyroid nodules. However, ambiguous nodule boundaries and their similarity to surrounding tissue cause challenges. These cause high false positive rates, reducing the diagnostic efficiency of ultrasound imaging. Although recent advancements in deep learning such as Faster R-CNN and YOLO, have shown progress, they still have limitations in handling imbalanced datasets and noise during real-time detection. To solve these challenges, our team proposes a novel deep learning framework integrating mixup data augmentation and negative sample utilization. This approach reduces false positives while maintaining a stable true positive detection rate. The method employs a one-stage-based object detection model as the baseline. The proposed framework based on that can improve the model’s robustness against background noise and optimize mixup parameters to enhance generalization. The experimental results demonstrate our approach significantly reduces false positive rates by 40% (from 50% to below 10%). In the meantime, it can stabilize high true positive rates above 80% in thyroid nodule detection during testing. This improvement highlights the framework can balance diagnostic sensitivity and specificity. The framework provides an efficient solution for detecting nodules in thyroid ultrasound imaging and has great potential to enhance diagnostic accuracy.
Qilei Chen, Zinan Xiong, Yu Cao 0002, Benyuan Liu
ICIP2
2025 A Multi-Stage Machine Learning Pipeline for Automated Bowel Preparation Scale Assessment in Colonoscopy Videos
abstract
Accurate assessment of intestinal cleanliness is essential for effective colonoscopy, but manual scoring with the Boston Bowel Preparation Scale (BBPS) remains subjective and labor-intensive. In this study, we propose a novel, fully automated multi-stage pipeline for objective BBPS-based bowel preparation assessment in colonoscopy videos. Our approach utilizes deep image classifiers, enhanced by a hierarchical sequence of binary classification tasks, to assign precise frame-level BBPS scores. From the resulting sequence of scores, we extract statistical and temporal features to train machine learning regression models for video-level BBPS prediction. Extensive experiments demonstrate that our pipeline achieves robust, reproducible, and granular video-level bowel cleanliness assessment, outperforming standard multi-class models at the frame level and offering a scalable tool for large-scale endoscopic quality analysis.
Yiqin He, Qilei Chen, Alimire Nabijiang, Benyuan Liu
ICMLA2
2025 Accurate Polyp Sizing via Attention-Guided Parallel CNN-ViT Depth and Learned Scale
abstract
Accurate sizing of colorectal polyps is essential for malignancy risk assessment and surveillance planning. However, both visual estimation and current instrument-based aids often suffer from variability or workflow disruption. Recent AI-based approaches typically rely on reference objects, camera calibration, or depth networks with suboptimal accuracy in endoscopic imagery. We present a calibration-free pipeline that segments the polyp and then couples ViTCAN-Depth, a monocular metric-depth model that fuses parallel CNN and Vision-Transformer encoders via channel-attention gating, with Depth–Pixel Linear (DPL), a lightweight module that converts pixel diameter and depth into real-world diameter using a learned scalar, eliminating camera calibration and reference objects. ViTCAN-Depth outperforms state-of-the-art CNN- or ViT-based metric-depth networks on the C3VD synthetic benchmark, and yields higher-fidelity depth maps on real colonoscopy images. Across 78 polyps from five independent cohorts, the system achieves a mean absolute error of 0.67mm (RMSE=0.88mm), suggesting its potential for integration into clinical workflows.
Alimire Nabijiang, Jiao Feng, Qilei Chen, Benyuan Liu
ICMLA4
2023 A CNN-Based Disease Detection Framework for Wireless Capsule Endoscopy Videos
abstract
In colonoscopy, wireless capsule endoscopy (WCE) is widely used since it is more physically friendly for patients than standard endoscopy. WCE is a low-risk and effective clinical operation for the small intestine endoscopy, which is less accessible to standard endoscopy. However, reviewing WCE videos can be challenging as it is time-consuming and requires considerable expertise. WCE videos are usually captured at a low resolution and partial frames are filtered out due to hardware limitations. Additional challenges arise from the diversity and complexity of gastrointestinal (GI) diseases. Moreover, inadequate clinic attention can cause clinical errors. Consequently, physicians are often overburdened with work. This paper presents a convolutional neutral networks (CNN) based framework to assist clinicians to review the WCE videos. The framework combines both image classification and object detection results to provide comprehensive results for disease detection in WCE videos. For disease classification, ResNet-50 was selected when experiments were conducted on the dataset. For disease detection, YOLO-X is employed on the images labelled with bonding boxes. Enhanced with an offline hard example mining (offline-HEM) procedure and fine-tuning on hyper-parameters, this framework can achieve high sensitivity for disease instances while maintaining acceptable specificity for false positives in video tests.
Qilei Chen, Yani Yin, Guanghui Lian, Shuijiao Chen, Yu Cao 0002, Benyuan Liu
HealthCom3
2023 A Greedy Algorithm-Based Self-Training Pipeline for Expansion of Dental Caries Dataset
abstract
Dental caries, the most prevalent oral disease, poses a significant healthcare challenge. Deep Neural Network (DNN)-based object detection techniques offer promising solutions to improve the efficiency of dental caries diagnosis. It is widely acknowledged that the performance of DNN models heavily relies on the availability of sufficient and accurately labeled data. The collection and annotation of dental X-ray images encounter obstacles due to privacy concerns and the requirement for specialized expertise. Consequently, the limited access to labeled dental image datasets restricts the potential of DNNs in supporting oral and dental healthcare. Self-Training (ST) is a semi-supervised machine leaning approach that addresses this problem to a large extent. It repeats the procedures of training a model on the labeled dataset, and then applying it to generate pseudo labels on the unlabeled dataset, and further using the combined data with the original and pseudo labels to train new models. However, the latent errors of the pseudo labels can arise and even be amplified throughout the ST pipeline, which leads to a significant performance decline for DNN models. In this paper, we propose a Greedy algorithm-based Self-Training (Greedy-ST) pipeline to address this problem. At each iteration, the Greedy-ST selects an optimal confidence threshold to generate predictions as pseudo labels, and uses static fine-tuning (SFT) and dynamic fine-tuning (DFT) to refine them. Experimental results demonstrate that by utilizing the pseudo labels generated by the Greedy-ST pipeline, the selected baseline model achieves improved performance compared to using the pseudo labels generated by the vanilla ST approach.
Qilei Chen, Yu Cao 0002, Xinwen Fu, Benyuan Liu
HealthCom5
2023 MLMSA: Multi-Level and Multi-Scale Attention for Lesion Detection in Endoscopy
abstract
The advancement of deep learning techniques has significantly improved abnormality detection in gastrointestinal (GI) endoscopy. However, this imaging process comes with challenges due to the complex nature of GI abnormalities. The wide variety of abnormalities in terms of type, color, texture, shape, and scale of lesions makes it difficult to accurately detect them in different scenarios. Furthermore, the presence of multiple types of lesions within the same region create complex scenarios that complicate abnormality detection. Additionally, differentiating early-stage cancers from non-cancerous lesions is a significant challenge even for experienced professionals. The simultaneous identification of cancers, particularly early-stage ones, and non-cancerous lesions within the same region remains a challenging issue in GI endoscopy imaging. In this study, we discover that multiple types of lesions exhibit a scale-sensitive characteristic that can be leveraged by multi-level feature-based deep learning models. Hence, we propose the use of a multi-level and multi-scale attention (MLMSA) neck module in a deep learning network. The MLMSA module utilizes multiple levels of features extracted from the backbone network to generate processed multi-level features that assist the detection head. By integrating the MLMSA module into the deep learning framework, our goal is to enhance the detection and differentiation of lesions, particularly the early-stage cancer, thereby advancing the capabilities of GI endoscopy imaging. Our experiment results show that integrating the MLMSA module leads to a significant improvement in the detection of GI abnormalities, providing compelling evidence for the enhanced performance achieved through the utilization of the MLMSA module in our approach.
Shuijiao Chen, Qilei Chen, Yizhe Zhang 0001, Yu Cao 0002, Benyuan Liu
HealthCom5
2023 A Deep Learning Framework with Pruning RoI Proposal for Dental Caries Detection in Panoramic X-ray Images
Qilei Chen, Yu Cao 0002, Xinwen Fu, Benyuan Liu
ICONIP (3)4
2022 Deep Learning Assisted Mouth-Esophagus Passage Time Estimation During Gastroscopy
abstract
A gastroscopy involves examining the upper digestive system using a flexible tube equipped with a small camera. Generally, it is performed to determine the cause of digestive symptoms, such as vomiting blood, stomach pains, and difficulty swallowing. Though this procedure has been performed since the mid-19th century, and various measures have been implemented to make it easier and less invasive, it is still not risk-free. One of the major complications is esophagus perforation, and most of them happen during the insertion of the gastroscopy. Therefore, it is necessary to develop an effective method for evaluating the performance of the operator. One appropriate metric is the time interval between the mouth and esophagus during the intubation. In this paper, we propose a gastroscopy video processing system based on deep learning to automatically evaluate the mouth-esophagus passage time. In this system, a Convolutional Neural Network (CNN) based model is adopted to detect the mouth and esophagus, track the timestamps of the last appearance of the mouth and the first appearance of the esophagus, and calculate the interval between those appearances. Our system is capable of dealing with abnormal circumstances that can occur during a procedure, as well as reporting accurate results. Experiment results show that our best model achieves an accuracy of 88.92% on image dataset, and an accuracy of 99.86% on videos for the mouth-esophagus passage time.
Zinan Xiong, Qilei Chen, Yu Cao 0002, Benyuan Liu
ICTAI2
2022 Enhance Chest X-ray Classification with Multi-image Fusion and Pseudo-3D Reconstruction
abstract
Chest radiography (X-ray) is a critical imaging modality for the diagnosis of thorax diseases. Automated classification of chest X-rays has enormous potential to benefit diagnostic decision making. Most existing models are designed to consider a single-view chest X-ray image, mostly the frontal view, failing to utilize the complementary information provided by the lateral view when available. In clinical practice, radiologists often examine both frontal and lateral views to gather more information for diagnosis. Thus, a well-designed model should take advantage of the paired views to enhance the classification results. To this end, we present a novel dual-view deep learning framework to enhance the classification that uses an intermediate bi-directional fusion architecture to exploit the intrinsic spatial correlation registered between the two views. In particular, we design two different modules, a vertical alignment fusion and a 3D reconstruction fusion, that can effectively extract and fuse important features from both input streams to generate a pair of recalibrating signals, which are then used to amplify the pertinent parts while suppressing the irrelevant segments of the feature map in each stream. We evaluate our proposed methods extensively on two large public chest X-ray datasets, CheXpert and MIMIC-CXR-JPG. Experiment results show that our methods can effectively improve the classification results over fine-tuned baselines and other state-of-the-art methods, demonstrating the benefit of exploiting the spatial correlation between the paired views in the model. We also find that our fusion methods significantly outperform the early fusion and late fusion models, highlighting the need to capture the spatial correlation across different intermediate feature layers.
Jing Ni, Zubin Bhuyan, Qilei Chen, Xinzi Sun, Dechun Wang, Yu Cao 0002, Benyuan Liu
IJCNN3
2022 AFP-Mask: Anchor-Free Polyp Instance Segmentation in Colonoscopy
abstract
Colorectal cancer (CRC) is a common and lethal disease. Globally, CRC is the third most commonly diagnosed cancer in males and the second in females. The most effective way to prevent CRC is through using colonoscopy to identify and remove precancerous growths at an early stage. The detection and removal of colorectal polyps have been found to be associated with a reduction in mortality from colorectal cancer. However, the false negative rate of polyp detection during colonoscopy is often high even for experienced physicians. With recent advances in deep learning based object detection techniques, automated polyp detection shows great potential in helping physicians reduce false positive rate during colonoscopy. In this paper, we propose a novel anchor-free instance segmentation framework that can localize polyps and produce the corresponding instance level masks without using predefined anchor boxes. Our framework consists of two branches: (a) an object detection branch that performs classification and localization, (b) a mask generation branch that produces instance level masks. Instead of predicting a two-dimensional mask directly, we encode it into a compact representation vector, which allows us to incorporate instance segmentation with one-stage bounding-box detectors in a simple yet effective way. Moreover, our proposed encoding method can be trained jointly with object detector. Our experiment results show that our framework achieves a precision of 99.36% and a recall of 96.44% on public datasets, outperforming existing anchor-free instance segmentation methods by at least 2.8% in mIoU on our private dataset.
Dechun Wang, Shuijiao Chen, Xinzi Sun, Qilei Chen, Yu Cao 0002, Benyuan Liu
IEEE J. Biomed. Health Informatics4
2021 Gastric Location Classification During Esophagogastroduodenoscopy Using Deep Neural Networks
abstract
Esophagogastroduodenoscopy (EGD) is a common procedure that visualizes the esophagus, stomach, and the duodenum by inserting a camera, attached to a long flexible tube, into the patient's mouth and down the stomach. A comprehensive EGD needs to examine all gastric locations, but since the camera is controlled manually, it is easy to miss some surface area and create diagnostic blind spots, which often result in life-costing oversights of early gastric cancer and other serious illnesses. In order to address this problem, we train a convolutional neural network to classify gastric locations based on the camera feed during an EGD, and based on the classifier and a triggering algorithm we propose, we build a video processing system that checks off each location as visited, allowing human operators to keep track of which locations they have visited and which they have not. Based on collected clinical patient reports, we consider six gastric locations, and we add a background class to our classifier to accomodate for the frames in EGD videos that do not resemble the six defined classes (including when the camera is outside of the patient body). Our best classifier achieves 98 % accuracy within the six gastric locations and 88 % accuracy including the background class, and our video processing system clearly checks off gastric locations in an expected order when being tested on recorded EGD videos. Lastly, we use class activation mapping to provide human-readable insight into how our trained classifier works.
Alexander Ding, Ying Li 0133, Qilei Chen, Yu Cao 0002, Benyuan Liu, Shuijiao Chen
BIBE3
2020 Pseudo-Labeling for Small Lesion Detection on Diabetic Retinopathy Images
abstract
Diabetic retinopathy (DR) is a primary cause of blindness in working-age people worldwide. About 3 to 4 million people with diabetes become blind because of DR every year. Diagnosis of DR through color fundus images is a common approach to mitigate such problem. However, DR diagnosis is a difficult and time consuming task, which requires experienced clinicians to identify the presence and significance of many small features on high resolution images. Convolutional Neural Network (CNN) has proved to be a promising approach for automatic biomedical image analysis recently. In this work, we investigate lesion detection on DR fundus images with CNN-based object detection methods. Lesion detection on fundus images faces two unique challenges. The first one is that our dataset is not fully labeled, i.e., only a subset of all lesion instances are marked. Not only will these unlabeled lesion instances not contribute to the training of the model, but also they will be mistakenly counted as false negatives, leading the model move to the opposite direction. The second challenge is that the lesion instances are usually very small, making them difficult to be found by normal object detectors. To address the first challenge, we introduce an iterative training algorithm for the semi-supervised method of pseudo-labeling, in which a considerable number of unlabeled lesion instances can be discovered to boost the performance of the lesion detector. For the small size targets problem, we extend both the input size and the depth of feature pyramid network (FPN) to produce a large CNN feature map, which can preserve the detail of small lesions and thus enhance the effectiveness of the lesion detector. The experimental results show that our proposed methods significantly outperform the baselines.
Qilei Chen, Jing Ni, Yu Cao 0002, Benyuan Liu, Honggang Zhang 0003
IJCNN1
2020 Retinopathy of Prematurity Stage Diagnosis Using Object Segmentation and Convolutional Neural Networks
abstract
Retinopathy of Prematurity (ROP) is an eye disorder primarily affecting premature infants with lower weights. It causes proliferation of vessels in the retina and could result in vision loss and, eventually, retinal detachment, leading to blindness. While human experts can easily identify severe stages of ROP, the diagnosis of earlier stages, which are the most relevant to determining treatment choice, are much more affected by variability in subjective interpretations of human experts. In recent years, there has been a significant effort to automate the diagnosis using deep learning. This paper builds upon the success of previous models and develops a novel architecture, which combines object segmentation and convolutional neural networks (CNN) to construct an effective classifier of ROP stages 1-3 based on neonatal retinal images. Motivated by the fact that the formation and shape of a demarcation line in the retina is the distinguishing feature between earlier ROP stages, our proposed system first trains an object segmentation model to identify the demarcation line at a pixel level and adds the resulting mask as an additional "color" channel in the original image. Then, the system trains a CNN classifier based on the processed images to leverage information from both the original image and the mask, which helps direct the model's attention to the demarcation line. In a number of careful experiments comparing its performance to previous object segmentation systems and CNN-only systems trained on our dataset, our novel architecture significantly outperforms previous systems in accuracy, demonstrating the effectiveness of our proposed pipeline.
Alexander Ding, Qilei Chen, Yu Cao 0002, Benyuan Liu
IJCNN2
2019 People Re-Identification by Multi-Branch CNN with Multi-Scale Features
abstract
People re-identification is a retrieval problem to find a person of interest among a gallery of person images from different cameras in various poses or view angles. How to get a strong feature representation for a person image plays an important role in performing people re-identification. In this paper, we present a novel end-to-end framework that extracts both global and local features with multiple scales to generate more discriminative representations. The model we design is a multi-branch network consisting of one global branch to obtain features of the whole input image from different convolutional layers and several local branches to obtain features from horizontal partitions in different granularities. Our method achieves state-of-the-art results on three challenging datasets (Market-1501, CUHK03 and DukeMTMC-reid).
Xinzi Sun, Qilei Chen, Yu Cao 0002, Benyuan Liu
ICIP3
2019 An Effective CNN Approach for Diabetic Retinopathy Stage Classification with Dual Inputs and Selective Data Sampling
abstract
Diabetic retinopathy (DR) is a vision-threatening complication among the diabetic population and a leading cause of blindness for working-age adults. Early detection and timely treatment can reduce the occurrence of blindness due to DR. Computer-aided diagnosis have great potential to significantly improve the accuracy and speed in DR detection over the traditional manual diagnosis process. In this paper, we present a deep convolutional neural network for DR stage classification, trained and evaluated on a large dataset. Our model uses high-resolution retinal fundus images of both the left and right eyes as inputs to take advantage of more detailed retinal lesion information in images and strong correlation between both eyes. Selective data sampling (SeS) is applied in the training process to mitigate the data imbalance problem. Experiments show that our model outperforms the fine-tuned Inception-v3 model by every measure, achieving an accuracy of 87.2% and a Kappa score of 0.806 on the Kaggle dataset.
Jing Ni, Qilei Chen, Chang Liu 0033, Yu Cao 0002, Benyuan Liu
ICMLA2
2019 Mini Lesions Detection on Diabetic Retinopathy Images via Large Scale CNN Features
abstract
Diabetic retinopathy (DR) is a diabetes complication that affects eyes. DR is a primary cause of blindness in working-age people and it is estimated that 3 to 4 million people with diabetes are blinded by DR every year worldwide. Early diagnosis have been considered an effective way to mitigate such problem. The ultimate goal of our research is to develop novel machine learning techniques to analyze the DR images generated by the fundus camera for automatically DR diagnosis. In this paper, we focus on identifying small lesions on DR fundus images. The results from our analysis, which include the lesion category and their exact locations in the image, can be used to facilitate the determination of DR severity (indicated by DR stages). Different from traditional object detection for natural images, lesion detection for fundus images have unique challenges. Specifically, the size of a lesion instance is usually very small, compared with the original resolution of the fundus images, making them diffcult to be detected. We analyze the lesion-vs-image scale carefully and propose a large-size feature pyramid network (LFPN) to preserve more image details for mini lesion instance detection. Our method includes an effective region proposal strategy to increase the sensitivity. The experimental results show that our proposed method is superior to the original feature pyramid network (FPN) method and Faster RCNN.
Qilei Chen, Xinzi Sun, Yu Cao 0002, Benyuan Liu
ICTAI1
2016 A Multi-Scale Cascaded Hierarchical Model for Image Labeling
abstract
Image labeling is an important and challenging task in the area of graphics and visual computing, where datasets with high quality labeling are critically needed. In this paper, based on the commonly accepted observation that the same semantic object in images with different resolutions may have different representations, we propose a novel multi-scale cascaded hierarchical model (MCHM) to enhance general image labeling methods. Our proposed approach first creates multi-resolution images from the original one to form an image pyramid and labels each image at different scale individually. Next, it constructs a cascaded hierarchical model and a feedback circle between image pyramid and labeling methods. The original image labeling result is used to adjust labeling parameters of those scaled images. Labeling results from the scaled images are then fed back to enhance the original image labeling results. These naturally form a global optimization problem under scale-space condition. We further propose a desirable iterative algorithm in order to run the model. The global convergence of the algorithm is proven through iterative approximation with latent optimization constraints. We have conducted extensive experiments with five widely used labeling methods on five popular image datasets. Experimental results indicate that MCHM improves labeling accuracy of the state-of-the-art image labeling approaches impressively.
Degui Xiao, Qilei Chen
Int. J. Pattern Recognit. Artif. Intell.2