EDBT 2026 Demo / reviewers in the wild / expert
Leiting Chen
dblp:48/1935 · also Lei-Ting Chen
· DBLP profile ↗
52ranked-venue papers
2as first author
28since 2021 · last 2026
0000-0003-2045-7369ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 23 · 13 since 2021Artificial intelligence and machine learning · 19 · 2 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 16 · 7 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Escaping the CAM Shadow: Uncertainty-Guided Reliable Learning for Weakly Supervised Semantic SegmentationabstractWeakly supervised semantic segmentation (WSSS) suffers from an inherent mismatch between coarse image-level annotations and dense pixel-level predictions. To bridge this gap, existing methods primarily focus on generating refined class activation maps (CAM) as pseudo-labels. However, we argue that this focus is insufficient as it overlooks a critical component: the segmentation decoder. The decoder is typically trained through superficial alignment of predictions with pseudo-labels in the logit space. Given the noisy nature of such labels, this naive supervision leads to error accumulation and limits performance. To address this, we propose an Uncertainty-Guided Reliable Learning (UGRL) framework that exerts dual control to reshape the learning process, achieving robust supervision that escapes the CAM shadow. The cornerstone of UGRL is a prototype-driven uncertainty modeling module that estimates the reliability of class-wise supervision. The modeled uncertainty enables two synergistic control mechanisms. First, it adaptively modulates classification and segmentation losses, encouraging the model to learn from more trustworthy signals. Second, it guides the structuring of the decoder’s feature space. Rather than relying solely on superficial alignment, UGRL enforces deeper representation alignment by applying contrastive learning on reliable pixels. This enables rich semantic transfer to fine-grained segmentation details. Extensive experiments on PASCAL VOC and MS COCO demonstrate that our method surpasses other state-of-the-art WSSS methods. Luyao Chang, Leiting Chen, Chuan Zhou 0004 |
AAAI | 2 |
| 2026 | Shedding the Facades, Connecting the Domains: Detecting Shifting Multimodal Hate Video with Test-Time AdaptationabstractHate Video Detection (HVD) is crucial for online ecosystems. Existing methods assume identical distributions between training (source) and inference (target) data. However, hateful content often evolves into irregular and ambiguous forms to evade censorship, resulting in substantial semantic drift and rendering previously trained models ineffective. Test-Time Adaptation (TTA) offers a solution by adapting models during inference to narrow the cross-domain gap, while conventional TTA methods target mild distribution shifts and struggle with the severe semantic drift in HVD. To tackle these challenges, we propose SCANNER, the first TTA framework tailored for HVD. Motivated by the insight that, despite the evolving nature of hateful manifestations, their underlying cores remain largely invariant (i.e., targeting is still based on characteristics like gender, race, etc), we leverage these stable cores as a bridge to connect the source and target domains. Specifically, SCANNER initially reveals the stable cores from the ambiguous layout in evolving hateful content via a principled centroid-guided alignment mechanism. To alleviate the impact of outlier-like samples that are weakly correlated with centroids during the alignment process, SCANNER enhances the prior by incorporating a sample-level adaptive centroid alignment strategy, promoting more stable adaptation. Furthermore, to mitigate semantic collapse from overly uniform outputs within clusters, SCANNER introduces an intra-cluster diversity regularization that encourages the cluster-wise semantic richness. Experiments show that SCANNER outperforms all baselines, with an average gain of 4.69% in Macro-F1 over the best. Jian Lang, Xikai Tang, Wenzheng Shu, Ting Zhong, Qiang Gao 0003, Yong Wang 0046, Leiting Chen, Fan Zhou 0002 |
AAAI | 8 |
| 2026 | Modality-Balanced Collaborative Distillation for Multi-Modal Domain GeneralizationabstractWeight Averaging (WA) has emerged as a powerful technique for enhancing generalization by promoting convergence to a flat loss landscape, which correlates with stronger out-of-distribution performance. However, applying WA directly to multi-modal domain generalization (MMDG) is challenging: differences in optimization speed across modalities lead WA to overfit to faster-converging ones in early stages, suppressing the contribution of slower yet complementary modalities, thereby hindering effective modality fusion and skewing the loss surface toward sharper, less generalizable minima. To address this issue, we propose MBCD, a unified collaborative distillation framework that retains WA's flatness-inducing advantages while overcoming its shortcomings in multi-modal contexts. MBCD begins with adaptive modality dropout in the student model to curb early-stage bias toward dominant modalities. A gradient consistency constraint then aligns learning signals between uni-modal branches and the fused representation, encouraging coordinated and smoother optimization. Finally, a WA-based teacher conducts cross-modal distillation by transferring fused knowledge to each uni-modal branch, which strengthens cross-modal interactions and steer convergence toward flatter solutions. Extensive experiments on MMDG benchmarks show that MBCD consistently outperforms existing methods, achieving superior accuracy and robustness across diverse unseen domains. Zhangtao Cheng, Ting Zhong, Leiting Chen, Fan Zhou 0002 |
AAAI | 4 |
| 2026 | From Shallow Humor to Metaphor: Towards Label-Free Harmful Meme Detection via LMM Agent Self-ImprovementabstractThe proliferation of harmful memes on online media poses significant risks to public health and stability. Existing detection methods heavily rely on large-scale labeled data for training, which necessitates substantial manual annotation efforts and limits their adaptability to the continually evolving nature of harmful content. To address these challenges, we present ALARM, the first lAbeL-free hARmful Meme detection framework powered by Large Multimodal Model (LMM) agent self-improvement. The core innovation of ALARM lies in exploiting the expressive information from "shallow" memes to iteratively enhance its ability to tackle more complex and subtle ones. ALARM consists of a novel Confidence-based Explicit Meme Identification mechanism that isolates the explicit memes from the original dataset and assigns them pseudo-labels. Besides, a new Pairwise Learning Guided Agent Self-Improvement paradigm is introduced, where the explicit memes are reorganized into contrastive pairs (positive vs. negative) to refine a learner LMM agent. This agent autonomously derives high-level detection cues from these pairs, which in turn empower the agent itself to handle complex and challenging memes effectively. Experiments on three diverse datasets demonstrate the superior performance and strong adaptability of ALARM to newly evolved memes. Notably, our method even outperforms label-driven methods. These results highlight the potential of label-free frameworks as a scalable and promising solution for adapting to novel forms and topics of harmful memes in dynamic online environments. Jian Lang, Rongpei Hong, Ting Zhong, Leiting Chen, Qiang Gao 0003, Fan Zhou 0002 |
KDD (1) | 4 |
| 2026 | Towards more efficient and better multi-view and multi-modal retinopathy assisted diagnosis
Yonghao Huang, Chuan Zhou 0004, Leiting Chen |
Artif. Intell. Medicine | 3 |
| 2026 | A Mamba-based multi-modal and multi-view ophthalmologic image analysis framework for correspondence relationships and complementary information modeling
Yonghao Huang, Leiting Chen, Chuan Zhou 0004 |
J. Vis. Commun. Image Represent. | 2 |
| 2025 | Uncertainty-aware Correspondence Distillation for Deep Image ClusteringabstractImage clustering is a challenging task in computer vision, with performance heavily dependent on the quality of feature representations due to the inherent complexity of images. However, current image clustering methods overlook the underlying semantic information during representation learning, leading to low-quality feature representations. Moreover, the absence of ground-truth labels amplifies the detrimental effects of unreliable data on semantic guidance, steering the model towards incorrect learning directions. In this work, we propose a novel deep image clustering method named Uncertainty-aware Correspondence Distillation (UCD) to address these issues. Specifically, we introduce the concept of representation correspondence to establish cross-level connections between instances and semantics, which is further employed as a distillation target to improve the network’s feature learning by complementing semantic information. To mitigate unnecessary similarity penalties arising from unreliable data, we develop robust dynamic weights for semantic guidance by modeling the uncertainty of image semantics. Extensive experiments on five benchmark datasets demonstrate the superiority of the proposed method. The code is available at https://github.com/YL616/UCD. Luyao Chang, Leiting Chen, Chuan Zhou 0004 |
ICASSP | 2 |
| 2025 | Uncertainty-Aware Contrastive Learning for deep clustering
Luyao Chang, Leiting Chen, Chuan Zhou 0004 |
Neurocomputing | 2 |
| 2024 | Adaptive attribute distribution similarity for few-shot learning
Anping Cai, Leiting Chen, Ziyu He, Shuqing Tao, Chuan Zhou 0004 |
Image Vis. Comput. | 2 |
| 2024 | Weighted Graph-Structured Semantics Constraint Network for Cross-Modal RetrievalabstractCross-modal retrieval aims to retrieve relevant content of different modalities by giving a query of another modality. The biggest difficulty is how to bridge the heterogeneous gap between different modalities. The commonly-used methods tend to focus on exploiting individual image-text pair and mining the relations of cross-modality data thereof, but ignore the role of multi-sample correlation. Moreover, more global, structural inter-pair knowledge contained by the training dataset will be under-used. To fully exploit graph-structured semantics and mine the semantic information in the dataset for learning discriminative representations, we propose Weighted Graph-structured Semantics Constraint Network (WGSCN), a unified, graph-based, semantic-constrained learning framework, in which GCN is used to mine comprehensive relation information from cross modality data. Our main inspiration is to design a novel two-branch GCN-based Cross-modal Semantic Encoding (GCSE) module to produce semantic embeddings with the both modality-specific and modality-shared correlation. Moreover, a GAN-based dual learning approach is used to further improve the discriminability and model the joint distribution across different modalities. Our proposed GDL uses semantic embeddings as supervisory signal to make the common representation semantically discriminative while adversarial learning and dual learning are used to make the common representation modality-invariant. Through comparative experiments on five commonly used cross-modal datasets, we have shown the superior retrieval accuracy of our WGSCN. Lei Zhang 0005, Leiting Chen, Chuan Zhou 0004, Xin Li 0079, Fan Yang 0054, Zhang Yi 0001 |
IEEE Trans. Multim. | 2 |
| 2023 | Model long-range dependencies for multi-modality and multi-view retinopathy diagnosis through transformers
Yonghao Huang, Leiting Chen, Chuan Zhou 0004, Lifeng Qiao, Shanlin Lan |
Knowl. Based Syst. | 2 |
| 2022 | Pre-MocoDiagnosis: Few-Shot Ophthalmic Diseases Recognition using Contrastive LearningabstractThe use of deep learning networks for medical image classification is becoming increasingly popular. However, annotated datasets for medical diagnosis are still difficult to obtain due to the Limitations of expertise and expensive consumption. In the absence of datasets, the effectiveness, and robustness of the networks are weak for traditional deep networks for medical diagnosis and problems with these kinds of data issues are quite common. Due to the long-tailed distribution of data and the lack of labeling data for some diseases, traditional deep networks and training techniques can suffer from significant overfitting and poor generalization. We examine the issue of locating the locations of impacted lesions in the absence of data and long-tailed distributions by building on recent developments in contrastive learning. The processing of each image—that is, its division into multiple smaller pictures and comparison between them—improves the ability to identify illnesses. Due to the modest variations between the sick and healthy areas of the disease and the necessity for the network to concentrate on the diseased portions of the disease, contrastive learning allows for the maintenance of intra-class undistorted properties. Our technique can swiftly identify an image’s essential details while masking as much superfluous information a spossible. A s a result of the pre-trained network’s increased sensitivity to image attributes, the network model is better able to generalize to samples and is more resilient when dealing with small sample sizes. Classification models when there are few training samples and a long tail distribution; however, when the training dataset is bigger, our method does not necessarily worse than conventional approaches, because few-shot learning is based on conventional techniques, few-shot learning still trains the classical network with the basis tasks. On the ISIC 2018 and internal ophthalmology datasets, our for the first time tests comparing few-shot learning and traditional classification network approachesd emonstrate that the classical few-shot learning performs better than the traditional deep network approach when training samples are few, and our suggested approach outperforms the classical few-shot learning in task setting on both datasets. Anping Cai, Leiting Chen, Jiahao Fang, Mengqi Sun, Chuan Zhou 0004 |
BIBM | 2 |
| 2022 | Few-shot Learning Framework Based on Adaptive Subspace for Skin Disease ClassificationabstractSkin disease classification from images is crucial to dermatological diagnosis. It is an important task to develop Computer Aided Detection(CAD) systems that can help dermatologists improve classification performances. One challenge limits the adoption of such systems so far: traditional CAD can’t handle new emerging diseases. The reason is that standard deep learning models can only identify the categories that appear in the training set. However, this limitation can be addressed with Few-shot learning, a machine learning paradigm where a classifier has to generalize to a new category not seen in training, only given a few examples of this category. However, because of the complexity of lesions, most Few-shot learning methods do not work well in medical tasks. We find that most methods assume a single similarity measure and only obtain a single feature space. Motivated by this, we propose a three-stage learning paradigm. In the second stage, we introduce the subspace method to construct a symmetric function. In the third stage, we propose a metric module that consists of two similarity measures. In this way, the model enables to learn more discriminative features from few shots of skin disease images and has better generalization ability. The results demonstrate that our method produces a substantial improvement on the ISIC-2019 dataset. Chuan Zhou 0004, Mengqi Sun, Leiting Chen, Anping Cai, Jiahao Fang |
BIBM | 3 |
| 2022 | Domain Adaptation for Medical Image Classification without Source DataabstractAlthough deep learning has achieved promising results on medical image classification, the domain shift between training and testing datasets leads to a low prediction accuracy. Domain adaptation is a effective solution. However, due to privacy issues and the lack of annotated data, it’s hard for conventional domain adaptation methods to access source images and labeled target images. To tackle this issue, we propose a novel framework that only requires unlabeled target domain data. This framework has two modules, one is based on class conditional generative adversarial net for source domain generation and another is for classification m odel t raining. S pecifically, the generator can generate target-style data as the pseudo-source data using random noise and a given label to improve the classifier. The increased a ccuracy of the classifier also can guide the generator. Besides, we introduce weight regularization and clustering-based regularization to keep the training process stable and fully explore the discriminative information. We take diabetic retinopathy grade classification as our task and conduct experiments on three datasets which are EysPACS, MESSIDOR and IDRiD. The experimental results show that our method performs well on only unlabeled target data, which proves that it is a general method and can be widely used in the field of medical image classification. Chuan Zhou 0004, Leiting Chen |
BIBM | 4 |
| 2022 | Pixel-wise triplet learning for enhancing boundary discrimination in medical image segmentation
Leiting Chen, Yu Deng 0005, Zhong Zhang 0004, Chuan Zhou 0004 |
Knowl. Based Syst. | 2 |
| 2022 | Deep Active Autoencoders for Outlier Detection
Jin Ning 0001, Leiting Chen, Chuan Zhou 0004 |
Neural Process. Lett. | 2 |
| 2021 | Automatic Medical Lesion Annotation via Feature Fusion Correlation NetworkabstractImplicit correlations between lesions in medical images are the key to improve the performance of automatic medical lesion annotation tasks, as additional information they provide can improve the accuracy of the annotation results. Existing methods use single-scale features to model the correlations. However, single-scale features are not conducive to the annotation of small and variable scale lesions. In this paper, we propose a feature fusion correlation network with a feature fusion module and a correlation learning module. The feature fusion module refines the feature maps extracted from different blocks of the backbone and obtains multi-scale features, while the correlation learning module uses these multi-scale features to capture the correlations between different scales of lesions. We conduct experiments on an in-house dataset, and the experiment results show that our method achieves the best performance. Chuan Zhou 0004, Junjing Chen, Ximan Tang, Siying Dai, Leiting Chen |
BIBM | 6 |
| 2021 | Medical Frequency Domain Learning: Consider Inter-class and Intra-class Frequency for Medical Image Segmentation and ClassificationabstractMedical image segmentation and classification tasks have become increasing accurate by employing deep neural networks. However, existing convolution neural networks models (CNNs) are challenging to achieve quite satisfactory results as medical objects and backgrounds are usually indistinguishable in spatial-domain images. In comparison, it is easier to analyze complex objects in frequency-domain images as different object information is retained in different frequency components. However, training CNNs in the frequency domain requires complex modification for network architecture. Thus, this paper proposes a method of learning in the frequency domain to train CNNs called Frequency domain attention (FDAM) Workflow, which only requires little parameters rise and modification in CNNs. FDAM utilizes the relationship of intra-class frequency to retain valuable frequency information and suppress trivial ones. Furthermore, to reduce computation, a Gate module is designed for deleting redundant frequency channels by exploiting the relationship of inter-class frequency. The proposed methods can be applied in various CNNs, such as U-Net, ResNet and DenseNet, while accepting frequency-domain data as input. Experiment results show a significant performance improvement compared to original CNNs for retinal vessel segmentation, glaucoma classification and pneumonia classification. Specifically, Gate module can improve accuracy while using less input data size. Yonghao Huang, Chuan Zhou 0004, Leiting Chen, Junjing Chen, Shanlin Lan |
BIBM | 3 |
| 2021 | Automatic Report Generation based on Multi-modal and Multi-view Model for Fundus ImagesabstractAutomatic generation of medical reports has gained increasing research interests, as the significant potential to assist physicians in decision-making. Medical reports generated by most existing methods are often inaccurate and inconsistent with clinical practice because these methods ignored multiple views and only focused on a single modal. Therefore, we introduced multi-view and multi-modal images that contain more valuable information into this task to generate more accurate and reliable diagnostic reports, the first attempt in this domain. In this paper, we propose a multi-task method to generate diagnostic reports for multi-view and multi-modal fundus images, which can simultaneously predict retinal-related diseases and diagnostic descriptions. Moreover, a multi-view and multi-modal features fusion module (MVMFF) is designed to fuse visual features extracted from different views and modalities, where channel attention mechanism (CAM) and a 3D convolution module are utilized to improve accuracy. Extensive experiments conducted on an internal dataset demonstrate that our method can generate accurate and reliable diagnostic reports that are clinically relevant. Shanlin Lan, Chuan Zhou 0004, Leiting Chen, Huqiu Fan, Yonghao Huang |
BIBM | 3 |
| 2021 | Enhancing Medical Image Classification via Augmentation-based Pre-trainingabstractCurrent deep learning-based medical image classification models are typically pre-trained on large-scale natural image databases (e.g., ImageNet) with random image augmentation processing, and then fine-tuned on medical image datasets with relatively small datasets to achieve satisfactory performance. However, this method ignores the differences in data augmentation operations applied to medical images and natural images, which sometimes leads to poor classification results. In this study, a self-supervised learning framework based on augmentation is designed to boost the final classification performance with three modules. First, security evaluating module is used to filter out unsafe data augmentation operations. Then augmentation learning module follows the augmentation strategy and extracts a priori knowledge using the rest of safe operations. Finally, classification performance is verified by classification module. We conducted thorough experiments on three public medical image datasets and evaluate four parametric levels of networks. Experimental results demonstrate that our method is superior to the previous state-of-the-art, while not requiring any external search time. Ximan Tang, Chuan Zhou 0004, Leiting Chen |
BIBM | 3 |
| 2021 | Cross-Modal Guidance for Hyperfluorescence Segmentation in Fundus Fluorescein AngiographyabstractFor lesions required segmentation are very similar to other background tissues, existing methods often incorrectly segment these tissues into foregrounds. As the fact that diagnostic descriptions can improve the performance of lesion segmentation, but existing methods do not sufficiently capture the correlations between images and the corresponding diagnostic descriptions. In this paper, we propose a novel cross-modal mutual-aware network which utilizes diagnostic descriptions to guide lesion segmentation. Concretely, the proposed network integrates several mutually aware feature fusion module to learn the co-relationship between visual and linguistic features at multiple levels, yielding the final mask which are more closely to the descriptions. We have carried out experiments on an internal dataset and the experimental results show a significant performance improvement compared to other traditional and cross-modal approaches. Chuan Zhou 0004, Leiting Chen, Lei Zhang 0005, Junjing Chen |
ICME | 4 |
| 2021 | Non-Local Attention Learning for Medical Image ClassificationabstractIn recent years, deep convolutional neural networks (CNNs) have been used with great success in medical image classification. However, within CNNs, the convolutional operation only considers localized regions and the stacking of pooling layers can lead to the loss of information about tiny lesions. In this paper, we propose a non-local attention learning method that models long-range dependencies between pixels to pre-serve global information and help CNNs better identify the tiny lesions. It consists of two main parts, the non-local attention module and the non-local visual context fusion module, one for improving global understanding of the visual scene and one for aggregating non-local visual features. These two modules are used in parallel with the CNN backbone as an auxiliary branch. We conducted extensive experiments on two public medical image datasets and showed that our model achieves the best performance in terms of accuracy, precision, and sensitivity on both datasets compared to other recent models. Leiting Chen, Haisheng Chen, Ximan Tang, Yu Deng 0005, Yongbiao Chen, Chuan Zhou 0004 |
ICME | 2 |
| 2021 | Let's Find Fluorescein: Cross-Modal Dual Attention Learning For Fluorescein Leakage Segmentation In Fundus Fluorescein AngiographyabstractAutomatic segmentation of fluorescein leakage in fundus fluorescein angiography images is important in the clinical diagnosis of advanced diabetic retinopathy. Despite the recent success of deep-learning-based models in improving medical image segmentation, segmentation of fluorescein leakage has been ignored owing to (1) a lack of publicly available data with sufficient annotations for training a segmentation network and (2) incapability of supervised models to accurately localize fluorescein leakage at different imaging angles. To address these issues, we studied the automatic segmentation of fluorescein leakage in fundus fluorescein angiography images and devised a method involving (1) a cross-modal learning framework for fluorescein leakage segmentation using both image and text data, (2) a dual attention learning module for identifying important linguistic and visual features, and (3) fluorescein-related-keyword classification for identifying meaningful textual expressions pertaining to the location and type of fluorescein leakage. We demonstrate the effectiveness of the proposed method for an in-house fundus fluorescein angiography image data set. Leiting Chen, Lifeng Qiao, Yu Deng 0005, Haisheng Chen, Chuan Zhou 0004 |
ICME | 2 |
| 2021 | Towards Efficient Medical Image Segmentation Via Boundary-Guided Knowledge DistillationabstractIn recent years, knowledge distillation for semantic segmentation has been extensively studied in order to obtain satisfactory performance while reducing computational costs. Compared with natural images, segmentation targets in medical images have fuzzy boundaries that are difficult to determine, but current knowledge distillation methods fail to explicitly transfer boundary information, resulting in poor boundary discrimination in compact models. Therefore, in this paper, we propose two knowledge distillation modules, namely boundary-guided deep supervision and output space boundary embeddng alignment, to explicitly transfer boundary information. The validity of our knowledge distillation approaches is demonstrated by extensive experiemnts on four public medical image data sets, namely, Montgomery County, CHAOS, GlaS and DRISHTI-GS. Leiting Chen, Shuo Xi, Yu Deng 0005, Ximan Tang, Chuan Zhou 0004 |
ICME | 2 |
| 2021 | Exploring Graph-Structured Semantics for Cross-Modal RetrievalabstractWe study and address the cross-modal retrieval problem which lies at the heart of visual-textual processing. Its major challenge lies in how to effectively learn a shared multi-modal feature space where the discrepancies of semantically related pairs, such as images and texts, are minimized regardless of their modalities. Most current methods focus on reasoning about cross-modality semantic relations within individual image-text pair to learn the common representation. However, they overlook more global, structural inter-pair knowledge within the dataset, i.e., the graph-structured semantics within each training batch. In this paper, we introduce a graph-based, semantic-constrained learning framework to comprehensively explore the intra- and inter-modality information for cross-modal retrieval. Our idea is to maximally explore the structures of labeled data in graph latent space, and use them as semantic constraints to enforce feature embeddings from the semantically-matched (image-text) pairs to be more similar and vice versa. It raises a novel graph-constrained common embedding learning paradigm for cross-modal retrieval, which is largely under-explored up to now. Moreover, a GAN-based dual learning approach is used to further improve the discriminability and model the joint distribution across different modalities. Our fully-equipped approach, called Graph-constrained Cross-modal Retrieval (GCR), is able to mine intrinsic structures of training data for model learning and enable reliable cross-modal retrieval. We empirically demonstrate that our GCR can achieve higher accuracy than existing state-of-the-art approaches on Wikipedia, NUS-WIDE-10K, PKU XMedia and Pascal Sentence datasets. Our code will be made publicly available. Code is available at https://github.com/neoscheung/GCR. Lei Zhang 0005, Leiting Chen, Chuan Zhou 0004, Fan Yang 0054, Xin Li 0079 |
ACM Multimedia | 2 |
| 2021 | Towards better semantic consistency of 2D medical image segmentation
Leiting Chen, Yu Deng 0005, Jin Ning 0001, Chuan Zhou 0004 |
J. Vis. Commun. Image Represent. | 2 |
| 2021 | Rethinking pre-training on medical imaging
Leiting Chen, Yu Deng 0005, Chuan Zhou 0004 |
J. Vis. Commun. Image Represent. | 2 |
| 2021 | Multi-object Spatial-Temporal Anomaly Detection Using an LSTM-Based Framework
Jin Ning 0001, Leiting Chen, Chuan Zhou 0004, Defu Liu 0001 |
Neural Process. Lett. | 2 |
| 2020 | Attention Based Detection for Central Serious Chorioretinopathy in Fundus ImageabstractThe detection of Central Serious Chorioretinopathy (CSCR) which is one of the main causes of vision loss so far still relies on manual evaluation of fundus images which requires experienced clinicians and is time-consuming. In this paper, we aim to utilize deep learning to overcome the difficulty of the absence of automated diagnostic methods for CSCR. However, there still exists two issues in general neural networks when applied to CSCR detection: 1) The lesion area that contributes the most for the final result cannot get adequate attention in one-stream networks. 2)Lesion regions become blurred when high-resolution fundus images are resized into lower resolution. Attention mechanism has been proved to be effective in addressing the problem of insufficient attention to significant regions, so we propose a Crop Attention Network (CA-Net) to screen CSCR automatically. CA-Net is based on attention framework and tackles the issues mentioned above by cropping the whole image into patches and adding weights on each patch. Experiment results on in-house database show that the proposed method outperforms all baseline methods. Chuan Zhou 0004, Leiting Chen, Junjing Chen |
BIBM | 3 |
| 2020 | SEAUNet: Domain Adaptation for Biomedical Image SegmentationabstractSemantic segmentation plays an important role in biomedical image analysis applications. Convolutional neural networks (CNN) have become a promising approach to segment biomedical images. Nevertheless, the accuracy of these methods is highly dependent on the training data. These models may not generalize well to unseen image domains due to the phenomenon of domain shift. We propose a self-ensembling attention networks to address domain shift for biomedical image segmentation. There are two main components in the proposed network: a student network which plays a role of the base networks and a teacher network which plays a role of the ensemble networks. As the iteration goes on, the student network becomes more accurate, and the ensemble predictions in the teacher network also get closer to the correct labels in the target domain. In this way, domain-invariant features can be learnt correspondingly. We use the DRIVE, STARE, HRF and CHASEDB eye vasculature segmentation datasets and show that our approach can significantly improve results where we only use labels of one domain in training and test on the other domain. Leiting Chen, Yuchu Chen, Minghao Fan, Chuan Zhou 0004 |
BIBM | 2 |
| 2020 | On the Effective Transfer Learning Strategy for Medical Image Analysis in Deep LearningabstractIn this study, we focus on exploring different strategies of transfer learning for medical applications. Firstly, we report competitive results indicating that convolutional neural networks (CNNs) that were pre-trained with different annotations could have diverse effects on the performance of medical image analysis, especially for segmentation tasks. Then, we present our further explorations of transferring different components of the CNNs, which revealed the importance of the decoder on medical segmentation. Finally, we demonstrate the advantages and disadvantages of transfer learning methods based on model integration. These observations present novel aspects of transfer learning for visual tasks in the medical field, and we expect that these discoveries will encourage the exploration of more effective transfer learning strategies for CNN-based medical image analysis. Leiting Chen, Chuan Zhou 0004, Yu Deng 0005, Huiru Zeng, Shuo Xi |
BIBM | 2 |
| 2020 | An Efficient Weakly-Supervised Learning Method for Optic Disc SegmentationabstractAccurate optic disc segmentation plays an essential role in the early diagnosis of glaucoma, which has been a major cause of irreversible blindness for the past decade. Recently, U-shape Convolutional Neural Network (CNN) models have achieved favourable performance in optic disc segmentation. However, it is worth noting that these models require a large number of pixel-level annotations while these annotations are difficult to obtain in clinical practice. As a solution, weakly-supervised training methods are commonly implemented, but it will provoke U-shape CNN generating inaccurate, diluted, and grid-like segmentation results. In this paper, we propose a novel Hybrid Network (HyNet) to solve the issue above. HyNet consists of a U-shape backbone hybridized with a cross-scale connection structure, which makes better use of multi-scale visual semantics. Nevertheless, the generalization ability of HyNet is affected by the domain shift among different datasets. Therefore, we innovatively combine weakly- and fully-supervised training methods, namely Hybrid Process (HyProcess), to solve the domain shift problem. Experimental results on ONHSD, DRIONS-DB, and DRISHTI-GS datasets show that our model outperforms the state-of-the-art, reaching Dice of 82.39(%), 93.72(%), and 95.34(%) respectively. Additionally, our ablation study validates the effectiveness of HyNet along with HyProcess, and further analysis reflects their value in clinical practice. Leiting Chen, Lifeng Qiao, Chuan Zhou 0004, Shuo Xi, Yu Deng 0005 |
BIBM | 2 |
| 2020 | On the Deep Learning-based Age Prediction of Color Fundus Images and Correlation with Ophthalmic DiseasesabstractColor fundus imaging is an important modality used for ophthalmic disease screening and provides a non-invasive invivo method for assessing body condition. We aimed to assess deep learning models for age prediction from fundus images of normal patients and patients with ophthalmic diseases. In addition, we sought to investigate interpretable clues regarding the salient regions between normal and pathological changes as determined by deep learning models during age prediction. In this study, we used a convolutional neural network model for age prediction and evaluated it on an in-house database of fundus images of the Chinese population. The results of the experiment revealed some conclusions as follows: (1) deep learning-based classification models have better age prediction performance than deep learning-based regression models of fundus images; (2) deep learning-based models tend to use holistic information of the fundus for age prediction; (3) ophthalmic diseases that cause damage to the structure of the fundus and change its appearance will result in a decline in age prediction performance. Leiting Chen, Lifeng Qiao, Yu Deng 0005, Chuan Zhou 0004 |
BIBM | 2 |
| 2020 | Symptom and Pathology Report Generation for Ophthalmic Diseases in Fundus ImagesabstractWhile fundus image interpretation and report writing are routine procedures in the diagnosis of ophthalmic diseases, they are error-prone for inexperienced ophthalmologists and laborious and tedious for experienced ophthalmologists. Therefore, there is an urgent need for an automated ophthalmic report generation tool. Existing methods for captioning medical images are confined to the field of radiological images, and there is no such method for fundus images. To address this issue, we studied automated report generation for fundus images. Unlike the radiological field, this task presents several challenges. First, there are a wide variety of ophthalmic diseases, and identifying diseases in fundus images is sometimes difficult since the diseases can have different manifestations. Second, an ophthalmic disease may have multiple subtypes, which can be distinguished from one another only through accurate identification o f specific ophthalmic symptoms. To overcome these challenges, we propose a method with the following features: (1) a multitask learning framework for predicting both ophthalmic symptoms and pathologies and for generating paragraphs, (2) two-branch architecture for handling ophthalmic symptoms and pathologies separately, and (3) a causal linkage module for transferring the symptom information for accurately predicting ophthalmic pathologies and thereby facilitating the identification of subtypes. We demonstrate the effectiveness of the proposed method for an in-house fundus image data set. Leiting Chen, Lifeng Qiao, Yu Deng 0005, Siying Dai, Junjing Chen, Chuan Zhou 0004 |
BIBM | 2 |
| 2020 | On Automatic Detection of Central Serous Chorioretinopathy and Central Exudative Chorioretinopathy in Fundus ImagesabstractAutomatic detection of chorioretinopathy plays an important role in clinical practice, but the detection of a major chorioretinopathy of central serous chorioretinopathy based on fundus photography images has rarely been studied, let alone distinguishing it from another chorioretinopathy of central exudative chorioretinopathy. Due to the high degree of similarity between the two chorioretinopathies on fundus images, it is difficult for the latest automatic methods to accurately distinguish between them. In this study, we design a deep neural network with two branches for different classification tasks, where the first one is to distinguish the normal and abnormal while the other is to classify the two chorioretinopathies. We manage to improve the classification accuracy by combining focal loss and discriminative loss. Extensive experiments are conducted for comparison between our method and other universal classification models using a private retinal fundus dataset. The results demonstrate that our method achieves the best performance with 97.69%, 99.58% and 98.87% on the accuracy, precision and sensitivity, respectively. Leiting Chen, Lifeng Qiao, Yu Deng 0005, Siying Dai, Junjing Chen, Chuan Zhou 0004 |
BIBM | 2 |
| 2020 | Enhancing Tiny Tissues Segmentation via Self-DistillationabstractAlthough the wide deployment of convolutional networks has greatly promoted the progress in the field of medical image segmentation, the performance of these method on tiny tissues, such as cell and fundus vessel, still needs to be improved. Most approaches focus on modifying the network architecture to overcome the problem of missing details in segmented images. In this paper, we try to solve this problem from a new perspective, that is, introducing self-distillation mechanism to fully utilize the features extracted from the network. Our method can be viewed as a combination of a novel loss function and a specific training strategy. It can be easily integrated into most existing encoderdecoder structured networks with few additional computational cost. We conduct experiments on four datasets, which are DRIVE, CHASEDB, GlaS and TNBC, and serval commonly used models to prove the effectiveness of our method. Experiments show that the performance of these models has been improved, which proves that our method is a general method and can be widely used in the field of m edical image segmentation. Chuan Zhou 0004, Yuchu Chen, Minghao Fan, Leiting Chen |
BIBM | 6 |
| 2020 | Automatic Detection Of Pathological Myopia And High Myopia On Fundus ImagesabstractAutomatic detection of myopia plays a significant role in clinical practice. Few studies have been done on the detection of pathological myopia, and no attention has been paid to the distinguishment between it and high myopia. Additionally, they are hard to differentiate because of the high similarity between them. In this paper, we design a network with two branches for different classification tasks, where the first one is to distinguish the normal and abnormal while the other is to classify pathological myopia and high myopia. We manage to improve the classification accuracy by combining Binary Cross-Entropy loss and Triplet loss. Extensive experiments are conducted for comparison between our method and other universal classification models using a private retinal fundus dataset. The results demonstrate that our method achieves the best performance with 81.82%, 83.61% and 83.52% on the accuracy, precision and sensitivity, respectively. Siying Dai, Leiting Chen, Chuan Zhou 0004 |
ICME | 2 |
| 2020 | Semi-supervised cross-modal representation learning with GAN-based Asymmetric Transfer Network
Lei Zhang 0005, Leiting Chen, Weihua Ou, Chuan Zhou 0004 |
J. Vis. Commun. Image Represent. | 2 |
| 2019 | Weakly Supervised Learning of Recurrent Residual ConvNets for Pancreas Segmentation in CT ScansabstractDeep neural networks trained by medical images with dense annotations have revealed favourable performance on accurate organ segmentation. The current supervised methods demand voxel-level annotations which are not easily accessible due to the consuming of time and requirements of specialized knowledge and skills. In this paper, we propose a weakly supervised method based on a recurrent residual convolutional neural network trained only with image-level labels to generate voxel-level segmentation. The recurrent residual convolutional units take advantage of contextual information of successive slices and a spatial pooling layer is introduced after the last convolutional layer to aggregate local features and learn accurate localization. The final segmentation mask is computed by applying a conditional random field for spatial prediction. Our method shows competitive performance to fully supervised methods on the public NIH-CT-82 dataset for pancreas segmentation. Huiru Zeng, Xiaohua Hu 0001, Leiting Chen, Chuan Zhou 0004 |
BIBM | 3 |
| 2018 | Multi-Scale Bidirectional FCN for Object Skeleton ExtractionabstractObject skeleton detection is a challenging problem with wide application. Recently, deep Convolutional Neural Networks (CNNs) have substantially improved the performance of the state-of-the-art in this task. However, most of the existing CNN-Based methods are based on a skip-layer structure where low-level and high-level features are combined and learned so as to gather multi-level contextual information. As shallow features are too messy and lack semantic knowledge, they may cause errors and inaccuracy. Therefore, we propose a novel network architecture, Multi-Scale Bidirectional Fully Convolutional Network (MSB-FCN), to better capture and consolidate multi-scale high-level context information for object skeleton detection. Our network uses only deep features to build multi-scale feature representations, and employs a bidirectional structure to collect contextual knowledge. Hence the proposed MSB-FCN has the ability to learn the semantic-level information from different sub-regions. Furthermore, we introduce dense connections into the bidirectional structure of our MSB-FCN to ensure that the learning process at each scale can directly encode information from all other scales. Extensive experiments on various commonly used benchmarks demonstrate that the proposed MSB-FCN has achieved significant improvements over the state-of-the-art algorithms. Fan Yang 0054, Xin Li 0079, Hong Cheng 0002, Yuxiao Guo 0001, Leiting Chen |
AAAI | 5 |
| 2018 | A canonical form-based approach to affine registration of DTI
Leiting Chen, Hongbin Cai, Qihe Liu, Nanxi Fei |
Multim. Tools Appl. | 2 |
| 2018 | Parameter k search strategy in outlier detection
Jin Ning 0001, Leiting Chen, Chuan Zhou 0004 |
Pattern Recognit. Lett. | 2 |
| 2017 | Object-Aware Dense Semantic Correspondence
Fan Yang 0054, Xin Li 0079, Hong Cheng 0002, Leiting Chen |
CVPR | 5 |
| 2017 | Multi-Scale Cascade Network for Salient Object DetectionabstractIn this paper we present a novel network architecture, called Multi-Scale Cascade Network (MSC-Net), to identify the most visually conspicuous objects in an image. Our network consists of several stages (sub-networks) for handling saliency detection across different scales. All these sub-networks form a cascade structure (in a coarse-to-fine manner) where the same underlying convolutional feature representations are fully shared. Compared with existing CNN-based saliency models, the MSC-Net can naturally enable the learning process in the finer cascade stages to encode more global contextual information while progressively incorporating the saliency prior knowledge obtained from coarser stages and thus lead to better detection accuracy. We also design a novel refinement module to further filter out errors by considering the intermediate feedback information. Our MSC-Net is highly integrated, end-to-end trainable, and very powerful. The proposed method achieves state-of-the-art performance on five widely-used salient object detection benchmarks, outperforming existing methods and also maintaining high efficiency. Code and pre-trained models are available at https://github.com/lixin666/MSC-NET. Xin Li 0079, Fan Yang 0054, Hong Cheng 0002, Junyu Chen 0002, Yuxiao Guo 0001, Leiting Chen |
ACM Multimedia | 6 |
| 2016 | Saliency Transfer: An Example-Based Method for Salient Object Detection
Xin Li 0079, Fan Yang 0054, Leiting Chen, Hongbin Cai |
IJCAI | 3 |
| 2016 | Evolution prediction of multi-scale information diffusion dynamics
Tao Wu 0003, Leiting Chen, Xingping Xian, Yuxiao Guo 0001 |
Knowl. Based Syst. | 2 |
| 2014 | Illumination-based nighttime video contrast enhancement using genetic algorithm
Yunbo Rao, Zhihui Wang 0001, Leiting Chen |
Multim. Tools Appl. | 4 |
| 2011 | Real-time control of individual agents for crowd simulation
Yunbo Rao, Leiting Chen, Qihe Liu, Weiyao Lin |
Multim. Tools Appl. | 2 |
| 2011 | Diffusion Kurtosis Imaging Based on Adaptive Spherical IntegralabstractDiffusion kurtosis imaging (DKI) is a recent approach in medical engineering that has potential value for both neurological diseases and basic neuroscience research. In this letter, we develop a robust method based on adaptive spherical integral that can compute kurtosis based quantities more precisely and efficiently. Our method integrates spherical trigonometry with a recursive computational scheme to make numerical estimations in kurtosis imaging convergent. Our algorithm improves the efficiency of computing integral invariants based on reconstructed diffusion kurtosis tensors and makes DKI better prepared for further clinical applications. Yugang Liu, Leiting Chen, Yizhou Yu |
IEEE Signal Process. Lett. | 2 |
| 2006 | Knowledge Reduction in Inconsistent Decision Tables
Qihe Liu, Leiting Chen, Fan Min 0001 |
ADMA | 2 |
| 2006 | Wafer Yield Estimation Using Support Vector Machines
Leiting Chen, Dan Muuniz, Chia-Jiu Wang |
ISNN (2) | 1 |
| 2006 | Audio Signal Classification Using Support Vector Machines
Leiting Chen, Ming-Jen Wang, Chia-Jiu Wang, Heng-Ming Tai |
ISNN (2) | 1 |