Chen Li 0011

dblp:l/ChenLi11 · DBLP profile ↗
← Back
81ranked-venue papers
2as first author
61since 2021 · last 2026
0000-0002-0079-3106ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 54 · 2 first-author · 37 since 2021Artificial intelligence and machine learning · 23 · 20 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 8 since 2021Databases, data management, data science and information retrieval · 8 · 7 since 2021Theory of computation · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 HAAF: Hierarchical Adaptation and Alignment of Foundation Models for Few-Shot Pathology Anomaly Detection
Chunze Yang, Junbo Lu, Jiusong Ge, Qidong Liu 0008, Zeyu Gao 0001, Chen Li 0011
WWW8
2026 MegaSeg: Towards scalable semantic segmentation for megapixel images
Solomon Kefas Kaura, Jialun Wu, Zeyu Gao 0001, Chen Li 0011
Medical Image Anal.4
2026 PH2ST: Prompt-guided hypergraph learning for spatial transcriptomics prediction in whole slide images
abstract
Spatial Transcriptomics (ST) reveals the spatial distribution of gene expression in tissues, offering critical insights into biological processes and disease mechanisms. However, the high cost, limited coverage, and technical complexity of current ST technologies restrict their widespread use in clinical and research settings, making obtaining high-resolution transcriptomic profiles across large tissue areas challenging. Predicting ST from H&E-stained histology images has emerged as a promising alternative to address these limitations but remains challenging due to the heterogeneous relationship between histomorphology and gene expression, which is affected by substantial variability across patients and tissue sections. In response, we propose PH2ST, a prompt-guided hypergraph learning framework, which leverages limited ST signals to guide multi-scale histological representation learning for accurate and robust spatial gene expression prediction. Extensive evaluations on two public ST datasets and multiple prompt sampling strategies simulating real-world scenarios demonstrate that PH2ST not only outperforms existing state-of-the-art methods, but also shows strong potential for practical applications such as imputing missing spots, ST super-resolution, and local-to-global prediction, highlighting its value for scalable and cost-effective spatial gene expression mapping in biomedical contexts.
Jiashuai Liu 0001, Yingkang Zhan, Jiangbo Shi, Marika Reinius, Inês Machado, Mireia Crispin-Ortuzar, Jialun Wu, Chen Li 0011, Zeyu Gao 0001
Medical Image Anal.10
2026 ProGIS: Prototype-Guided Interactive Segmentation for Pathological Images
abstract
Interactive segmentation offers greater clinical potential in computational pathology compared to traditional automatic segmentation. By incorporating interactive input, it addresses the limitations of fully automatic segmentation models, which often fail to meet pathologists' requirements and rely heavily on large-scale, pixel-level annotated datasets. However, current interactive segmentation methods struggle to balance interaction cost and segmentation performance, and they fail to adapt effectively to slide-level segmentation, a task that is even more crucial in routine pathology analysis. In this study, we propose a Prototype-Guided Interactive Segmentation (ProGIS) framework for pathological image segmentation, designed to deliver precise segmentation results efficiently with minimal interaction signals. ProGIS identifies all same-type tissue connected components in a single interaction and supports multi-class segmentation without predefined categories during inference. Moreover, ProGIS can be easily adapted for slide-level interactive segmentation. Specifically, ProGIS consists of three modules: Prototype Initialization, Prototype Navigation, and Local Refinement. First, the Prototype Initialization module identifies categorical prototypes, which are then utilized in the Prototype Navigation module to identify all tissue connected components belonging to the same type. The local refinement module further refines the segmentation results using detailed correction signals to ensure the accuracy of challenging-to-distinguish regions. We evaluate our framework on two regions of interest level and two slide-level pathological segmentation datasets, achieving new state-of-the-art performance with fewer interactions than existing methods. Our code is available at https://github.com/JSGe-AI/ProGIS.
Jiusong Ge, Yingkang Zhan, Jiashuai Liu 0001, Tieliang Gong, Jialun Wu, Mireia Crispin-Ortuzar, Chen Li 0011, Zeyu Gao 0001
IEEE Trans. Medical Imaging8
2025 Learning Heterogeneous Embedding with Prototype-Aware Graph Attention for Whole Slide Image Classification
abstract
Whole Slide Images (WSIs) are the digital version of pathology slides, central to computational pathology. The WSI pyramid makes it can offer a wide range of diagnostic information, from global tissue structures to detailed cellular features. However, current multi-instance and graph representation learning methods struggle to create a unified framework that effectively captures both local spatial awareness and global WSI representation, limiting their performance in critical tasks such as tumor staging. To this end, we propose a Prototypeaware Heterogeneous Graph ATtention (PHGAT) network that enables each region within a WSI to perceive the representations of its diverse heterogeneous neighbors. This, in turn, guides the learning of WSI-level heterogeneous embedding through multilevel prototypes. Specifically, we introduce three node relations (i.e., local, non-local, and hierarchical) into WSI heterogeneous graph construction and design a novel Heterogeneous Calibration Graph ATtention (HC-GAT) module to propagate the various heterogeneous neighbor node representations within graphs. Then, a Level-aware Prototype Attention module is proposed to obtain prototypes from different levels by aggregating the node representations via a set of trainable query embeddings. Lastly, based on these learned prototypes, a prototype-aware hierarchical pooling module is designed to generate the final heterogeneous embedding of each WSI. Extensive experiments on six diverse datasets across three cancer types and two specific diagnostic tasks show that the proposed framework significantly outperforms the state-of-the-art tumor staging methods and performs comparably in cancer subtyping.
Jiashuai Liu 0001, Yingkang Zhan, Jiangbo Shi, Chen Li 0011, Zeyu Gao 0001
BIBM7
2025 Exactly Tight Information-theoretic Generalization Bounds via Binary Jensen-Shannon Divergence
abstract
Information-theoretic bounds, while achieving significant success in analyzing the generalization of randomized learning algorithms, have been criticized for their slow convergence rates and overestimation. This paper presents novel bounds that bridge the expected empirical and population risks through a binarized variant of the Jensen-Shannon divergence. Leveraging our foundational lemma that characterizes the interaction between an arbitrary and a binary variable, we derive hypothesis-based bounds that enhance existing conditional mutual information bounds by reducing the number of conditioned samples from $2$ to $1$. We additionally establish prediction-based bounds that surpass prior bounds based on evaluated loss mutual information measures. Thereafter, through a new binarization technique for the evaluated loss variables, we obtain exactly tight generalization bounds broadly applicable to general randomized learning algorithms for any bounded loss functions. Our results effectively address key limitations of previous results in analyzing certain stochastic convex optimization problems, without requiring additional stability or compressibility assumptions about the learning algorithm.
Yuxin Dong 0003, Haoran Guo, Tieliang Gong, Wen Wen 0013, Chen Li 0011
ICML5
2025 StaDis: Stability distance to detecting out-of-distribution data in computational pathology
Jiusong Ge, Jiashuai Liu 0001, Chunbao Wang 0002, Tieliang Gong, Zeyu Gao 0001, Chen Li 0011
Medical Image Anal.7
2025 How Does Distribution Matching Help Domain Generalization: An Information-Theoretic Analysis
abstract
Domain generalization aims to learn invariance across multiple source domains, thereby enhancing generalization against out-of-distribution data. While gradient or representation matching algorithms have achieved remarkable success in domain generalization, these methods generally lack generalization guarantees or depend on strong assumptions, leaving a gap in understanding the underlying mechanism of distribution matching. In this work, we formulate domain generalization from a novel probabilistic perspective, ensuring robustness while avoiding overly conservative solutions. Through comprehensive information-theoretic analysis, we provide key insights into the roles of gradient and representation matching in promoting generalization. Our results reveal the complementary relationship between these two components, indicating that existing works focusing solely on either gradient or representation alignment are insufficient to solve the domain generalization problem. In light of these theoretical findings, we introduce IDM to simultaneously align the inter-domain gradients and representations. Integrated with the proposed PDM method for complex distribution matching, IDM achieves superior performance over various baseline methods.
Yuxin Dong 0003, Tieliang Gong, Hong Chen 0004, Shuangyong Song, Weizhan Zhang, Chen Li 0011
IEEE Trans. Inf. Theory6
2025 CoD-MIL: Chain-of-Diagnosis Prompting Multiple Instance Learning for Whole Slide Image Classification
abstract
Multiple instance learning (MIL) has emerged as a prominent paradigm for processing the whole slide image with pyramid structure and giga-pixel size in digital pathology. However, existing attention-based MIL methods are primarily trained on the image modality and a pre-defined label set, leading to limited generalization and interpretability. Recently, vision language models (VLM) have achieved promising performance and transferability, offering potential solutions to the limitations of MIL-based methods. Pathological diagnosis is an intricate process that requires pathologists to examine the WSI step-by-step. In the field of natural language process, the chain-of-thought (CoT) prompting method is widely utilized to imitate the human reasoning process. Inspired by the CoT prompt and pathologists' clinic knowledge, we propose a chain-of-diagnosis prompting multiple instance learning (CoD-MIL) framework for whole slide image classification. Specifically, the chain-of-diagnosis text prompt decomposes the complex diagnostic process in WSI into progressive sub-processes from low to high magnification. Additionally, we propose a text-guided contrastive masking module to accurately localize the tumor region by masking the most discriminative instances and introducing the guidance of normal tissue texts in a contrastive way. Extensive experiments conducted on three real-world subtyping datasets demonstrate the effectiveness and superiority of CoD-MIL.
Jiangbo Shi, Chen Li 0011, Tieliang Gong, Chunbao Wang 0002, Huazhu Fu
IEEE Trans. Medical Imaging2
2025 Efficient Approximations for Matrix-Based Rényi's Entropy on Sequential Data
abstract
The matrix-based Rényi's entropy (MBRE) has recently been introduced as a substitute for the original Rényi's entropy that could be directly obtained from data samples, avoiding the expensive intermediate step of density estimation. Despite its remarkable success in a broad of information-related tasks, the computational cost of MBRE, however, becomes a bottleneck for large-scale applications. The challenge, when facing sequential data, is further amplified due to the requirement of large-scale eigenvalue decomposition on multiple dense kernel matrices constructed by sliding windows in the region of interest, resulting in overall time complexity, where and denote the number and the size of windows, respectively. To overcome this issue, we adopt the static MBRE estimator together with a variance reduction criterion to develop randomized approximations for the target entropy, leading to high accuracy with substantially lower query complexity by utilizing the historical estimation results. Specifically, assuming that the changes of adjacent sliding windows are bounded by , which is a trivial case in domains, e.g., time-series analysis, we lower the complexity by a factor of . Polynomial approximation techniques are further adopted to support arbitrary orders. In general, our algorithms achieve total computational complexity, where denote the number of vector queries and the polynomial degrees, respectively. Theoretical upper and lower bounds are established in terms of the convergence rate for both and , and large-scale experiments on both simulation and real-world data are conducted to validate the effectiveness of our algorithms. The results show that our methods achieve promising speedup with only a trivial loss in performance.
Yuxin Dong 0003, Tieliang Gong, Hong Chen 0004, Chen Li 0011
IEEE Trans. Neural Networks Learn. Syst.4
2024 Shallow-Deep Synergy: Boosting Cross-Domain Generalization in Histopathological Image Segmentation
abstract
Accurate histopathological image segmentation is crucial for precise disease diagnosis and prognosis. Yet, challenges like staining variations, imaging conditions, and tissue diversity impede model generalization across domains, such as different institutes or organs. Traditional domain generalization (DG) techniques, such as data augmentation and feature alignment, excel in classification tasks but face challenges in segmentation tasks due to their dense prediction requirements. These tasks are particularly computationally demanding, and are complicated due to the fine-grained feature variability that arises from the domain differences in histopathological images. To tackle this, we propose the Shallow-Deep Synergy (SDS) approach for the U-Net-based segmentation framework, which capitalizes on the distinctive characteristics of both shallow and deep layers of the U-Net. Specifically, we introduce the fine-grained domain variations in image intensities and textures for shallow layers, while focusing on aligning the pixel-level classification decision boundaries in deep layers by adjusting the optimization trajectory through class-wise gradient and feature alignment. Moreover, the SDS is equipped with a big-batch strategy further boosting alignment efficiency, achieving high accuracy without substantial GPU memory. Extensive experiments conducted on two histopathological segmentation datasets, each representing different domain types, demonstrate that the proposed SDS achieves superior generalization performance compared to existing domain generalization methods, even being competitive with intra-domain models in some cases.
Weiheng Su, Yuxing Dong, Yang Li 0139, Xianli Zhang, Tieliang Gong, Inês Machado, Mireia Crispin-Ortuzar, Chen Li 0011, Zeyu Gao 0001
BIBM9
2024 ViLa-MIL: Dual-scale Vision-Language Multiple Instance Learning for Whole Slide Image Classification
abstract
Multiple instance learning (MIL)-based framework has become the mainstream for processing the whole slide image (WSI) with giga-pixel size and hierarchical image context in digital pathology. However, these methods heavily depend on a substantial number of bag-level labels and solely learn from the original slides, which are easily affected by variations in data distribution. Recently, vision language model (VLM)-based methods introduced the language prior by pre-training on large-scale pathological image-text pairs. However, the previous text prompt lacks the consideration of pathological prior knowledge, there-fore does not substantially boost the model's performance. Moreover, the collection of such pairs and the pre-training process are very time-consuming and source-intensive. To solve the above problems, we propose a dual-scale vision-language multiple instance learning (ViLa-MIL) framework for whole slide image classification. Specifically, we propose a dual-scale visual descriptive text prompt based on the frozen large language model (LLM) to boost the performance of VLM effectively. To transfer the VLM to process WSI efficiently, for the image branch, we propose a prototype-guided patch decoder to aggregate the patch features progressively by grouping similar patches into the same prototype; for the text branch, we introduce a context-guided text decoder to enhance the text features by incorporating the multi-granular image contexts. Extensive studies on three multi-cancer and multi-center subtyping datasets demonstrate the superiority of ViLa-MIL.
Jiangbo Shi, Chen Li 0011, Tieliang Gong, Yefeng Zheng 0001, Huazhu Fu
CVPR2
2024 Rethinking Information-theoretic Generalization: Loss Entropy Induced PAC Bounds
abstract
Information-theoretic generalization analysis has achieved astonishing success in characterizing the generalization capabilities of noisy and iterative learning algorithms. However, current advancements are mostly restricted to average-case scenarios and necessitate the stringent bounded loss assumption, leaving a gap with regard to computationally tractable PAC generalization analysis, especially for long-tailed loss distributions. In this paper, we bridge this gap by introducing a novel class of PAC bounds through leveraging loss entropies. These bounds simplify the computation of key information metrics in previous PAC information-theoretic bounds to one-dimensional variables, thereby enhancing computational tractability. Moreover, our data-independent bounds provide novel insights into the generalization behavior of the minimum error entropy criterion, while our data-dependent bounds improve over previous results by alleviating the bounded loss assumption under both leave-one-out and supersample settings. Extensive numerical studies indicate strong correlations between the generalization error and the induced loss entropy, showing that the presented bounds adeptly capture the patterns of the true generalization gap under various learning scenarios.
Yuxin Dong 0003, Tieliang Gong, Hong Chen 0004, Shujian Yu, Chen Li 0011
ICLR5
2024 Towards Generalization beyond Pointwise Learning: A Unified Information-theoretic Perspective
abstract
The recent surge in contrastive learning has intensified the interest in understanding the generalization of non-pointwise learning paradigms. While information-theoretic analysis achieves remarkable success in characterizing the generalization behavior of learning algorithms, its applicability is largely confined to pointwise learning, with extensions to the simplest pairwise settings remaining unexplored due to the challenges of non-i.i.d losses and dimensionality explosion. In this paper, we develop the first series of information-theoretic bounds extending beyond pointwise scenarios, encompassing pointwise, pairwise, triplet, quadruplet, and higher-order scenarios, all within a unified framework. Specifically, our hypothesis-based bounds elucidate the generalization behavior of iterative and noisy learning algorithms via gradient covariance analysis, and our prediction-based bounds accurately estimate the generalization gap with computationally tractable low-dimensional information metrics. Comprehensive numerical studies then demonstrate the effectiveness of our bounds in capturing the generalization dynamics across diverse learning scenarios.
Yuxin Dong 0003, Tieliang Gong, Hong Chen 0004, Zhongjiang He, Mengxiang Li, Shuangyong Song, Chen Li 0011
ICML7
2024 A Semantic Segmentation Method for SAR Image with Assistance of Self-Supervised Scene Classification
abstract
Unlike natural images, synthetic aperture radar (SAR) images exhibit a more scattered and uneven spatial distribution of objects, making semantic segmentation of SAR images a valuable topic of research. This paper presents a SAR image semantic segmentation method that incorporates the attention mechanism assisted by self-supervised scene classification. The self-supervised scene classification provides coarse scene classification at a higher semantic level, while the attention mechanism utilizes high-level semantic features to guide fine-grained classification of lower-level spatial structures. Overall, this approach improves the pixel-level classification performance of SAR images. We validate this method on the WHU-OPT-SAR dataset and compare its performance with previous works, providing a detailed analysis of its effectiveness.
Chen Li 0011, Zenghui Zhang, Wenxian Yu
IGARSS2
2024 PAMIL: Prototype Attention-Based Multiple Instance Learning for Whole Slide Image Classification
Jiashuai Liu 0001, Anyu Mao, Xianli Zhang, Tieliang Gong, Chen Li 0011, Zeyu Gao 0001
MICCAI (4)6
2024 E2-MIL: An explainable and evidential multiple instance learning framework for whole slide image classification
Jiangbo Shi, Chen Li 0011, Tieliang Gong, Huazhu Fu
Medical Image Anal.2
2024 Integrating K+ Entities Into Coreference Resolution on Biomedical Texts
abstract
Biomedical Coreference Resolution focuses on identifying the coreferences in biomedical texts, which normally consists of two parts: (i) mention detection to identify textual representation of biological entities and (ii) finding their coreference links. Recently, a popular approach to enhance the task is to embed knowledge base into deep neural networks. However, the way in which these methods integrate knowledge leads to the shortcoming that such knowledge may play a larger role in mention detection than coreference resolution. Specifically, they tend to integrate knowledge prior to mention detection, as part of the embeddings. Besides, they primarily focus on mention-dependent knowledge (KBase), i.e., knowledge entities directly related to mentions, while ignores the correlated knowledge (K+) between mentions in the mention-pair. For mentions with significant differences in word form, this may limit their ability to extract potential correlations between those mentions. Thus, this paper develops a novel model to integrate both KBase and K+ entities and achieves the state-of-the-art performance on BioNLP and CRAFT-CR datasets. Empirical studies on mention detection with different length reveals the effectiveness of the KBase entities. The evaluation on cross-sentence and match/mismatch coreference further demonstrate the superiority of the K+ entities in extracting background potential correlation between mentions.
Yufei Li 0002, Xiaoyong Ma, Penghzhen Cheng, Kai He 0001, Tieliang Gong, Chen Li 0011
IEEE ACM Trans. Comput. Biol. Bioinform.7
2024 Markov Subsampling Based on Huber Criterion
abstract
Subsampling is an important technique to tackle the computational challenges brought by big data. Many subsampling procedures fall within the framework of importance sampling, which assigns high sampling probabilities to the samples appearing to have big impacts. When the noise level is high, those sampling procedures tend to pick many outliers and thus often do not perform satisfactorily in practice. To tackle this issue, we design a new Markov subsampling strategy based on Huber criterion (HMS) to construct an informative subset from the noisy full data; the constructed subset then serves as refined working data for efficient processing. HMS is built upon a Metropolis-Hasting procedure, where the inclusion probability of each sampling unit is determined using the Huber criterion to prevent over scoring the outliers. Under mild conditions, we show that the estimator based on the subsamples selected by HMS is statistically consistent with a sub-Gaussian deviation bound. The promising performance of HMS is demonstrated by extensive studies on large-scale simulations and real data examples.
Tieliang Gong, Yuxin Dong 0003, Hong Chen 0004, Bo Dong 0001, Chen Li 0011
IEEE Trans. Neural Networks Learn. Syst.5
2024 Template-Free Prompting for Few-Shot Named Entity Recognition via Semantic-Enhanced Contrastive Learning
abstract
Prompt tuning has achieved great success in various sentence-level classification tasks by using elaborated label word mappings and prompt templates. However, for solving token-level classification tasks, e.g., named entity recognition (NER), previous research, which utilizes N-gram traversal for prompting all spans with all possible entity types, is time-consuming. To this end, we propose a novel prompt-based contrastive learning method for few-shot NER without template construction and label word mappings. First, we leverage external knowledge to initialize semantic anchors for each entity type. These anchors are simply appended with input sentence embeddings as template-free prompts (TFPs). Then, the prompts and sentence embeddings are in-context optimized with our proposed semantic-enhanced contrastive loss. Our proposed loss function enables contrastive learning in few-shot scenarios without requiring a significant number of negative samples. Moreover, it effectively addresses the issue of conventional contrastive learning, where negative instances with similar semantics are erroneously pushed apart in natural language processing (NLP)-related tasks. We examine our method in label extension (LE), domain-adaption (DA), and low-resource generalization evaluation tasks with six public datasets and different settings, achieving state-of-the-art (SOTA) results in most cases.
Kai He 0001, Rui Mao 0010, Tieliang Gong, Chen Li 0011, Erik Cambria
IEEE Trans. Neural Networks Learn. Syst.5
2023 Robust and Fast Measure of Information via Low-Rank Representation
abstract
The matrix-based Rényi's entropy allows us to directly quantify information measures from given data, without explicit estimation of the underlying probability distribution. This intriguing property makes it widely applied in statistical inference and machine learning tasks. However, this information theoretical quantity is not robust against noise in the data, and is computationally prohibitive in large-scale applications. To address these issues, we propose a novel measure of information, termed low-rank matrix-based Rényi's entropy, based on low-rank representations of infinitely divisible kernel matrices. The proposed entropy functional inherits the specialty of of the original definition to directly quantify information from data, but enjoys additional advantages including robustness and effective calculation. Specifically, our low-rank variant is more sensitive to informative perturbations induced by changes in underlying distributions, while being insensitive to uninformative ones caused by noises. Moreover, low-rank Rényi's entropy can be efficiently approximated by random projection and Lanczos iteration techniques, reducing the overall complexity from O(n³) to O(n²s) or even O(ns²), where n is the number of data samples and s ≪ n. We conduct large-scale experiments to evaluate the effectiveness of this new information measure, demonstrating superior results compared to matrix-based Rényi's entropy in terms of both performance and computational efficiency.
Yuxin Dong 0003, Tieliang Gong, Shujian Yu, Hong Chen 0004, Chen Li 0011
AAAI5
2023 Enhancing Cross-Lingual Few-Shot Named Entity Recognition by Prompt-Guiding
Tieliang Gong, Chen Li 0011
ICANN (1)4
2023 Understanding the Generalization Ability of Deep Learning Algorithms: A Kernelized Rényi's Entropy Perspective
abstract
Recently, information-theoretic analysis has become a popular framework for understanding the generalization behavior of deep neural networks. It allows a direct analysis for stochastic gradient / Langevin descent (SGD/SGLD) learning algorithms without strong assumptions such as Lipschitz or convexity conditions. However, the current generalization error bounds within this framework are still far from optimal, while substantial improvements on these bounds are quite challenging due to the intractability of high-dimensional information quantities. To address this issue, we first propose a novel information theoretical measure: kernelized Rényi's entropy, by utilizing operator representation in Hilbert space. It inherits the properties of Shannon's entropy and can be effectively calculated via simple random sampling, while remaining independent of the input dimension. We then establish the generalization error bounds for SGD/SGLD under kernelized Rényi's entropy, where the mutual information quantities can be directly calculated, enabling evaluation of the tightness of each intermediate step. We show that our information-theoretical bounds depend on the statistics of the stochastic gradients evaluated along with the iterates, and are rigorously tighter than the current state-of-the-art (SOTA) results. The theoretical findings are also supported by large-scale empirical studies.
Yuxin Dong 0003, Tieliang Gong, Hong Chen 0004, Chen Li 0011
IJCAI4
2023 Dual Attention and Patient Similarity Network for drug recommendation
abstract
MOTIVATION: Artificially making clinical decisions for patients with multi-morbidity has long been considered a thorny problem due to the complexity of the disease. Drug recommendations can assist doctors in automatically providing effective and safe drug combinations conducive to treatment and reducing adverse reactions. However, the existing drug recommendation works ignored two critical information. (i) Different types of medical information and their interrelationships in the patient's visit history can be used to construct a comprehensive patient representation. (ii) Patients with similar disease characteristics and their corresponding medication information can be used as a reference for predicting drug combinations. RESULTS: To address these limitations, we propose DAPSNet, which encodes multi-type medical codes into patient representations through code- and visit-level attention mechanisms, while integrating drug information corresponding to similar patient states to improve the performance of drug recommendation. Specifically, our DAPSNet is enlightened by the decision-making process of human doctors. Given a patient, DAPSNet first learns the importance of patient history records between diagnosis, procedure and drug in different visits, then retrieves the drug information corresponding to similar patient disease states for assisting drug combination prediction. Moreover, in the training stage, we introduce a novel information constraint loss function based on the information bottleneck principle to constrain the learned representation and enhance the robustness of DAPSNet. We evaluate the proposed DAPSNet on the public MIMIC-III dataset, our model achieves relative improvements of 1.33%, 1.20% and 2.03% in Jaccard, F1 and PR-AUC scores, respectively, compared to state-of-the-art methods. AVAILABILITY AND IMPLEMENTATION: The source code is available at the github repository: https://github.com/andylun96/DAPSNet.
Jialun Wu, Yuxin Dong 0003, Zeyu Gao 0001, Tieliang Gong, Chen Li 0011
Bioinform.5
2023 Virtual prompt pre-training for prototype-based few-shot relation extraction
Kai He 0001, Rui Mao 0010, Tieliang Gong, Chen Li 0011, Erik Cambria
Expert Syst. Appl.5
2023 A semi-supervised multi-task learning framework for cancer classification with weak annotation in whole-slide images
Zeyu Gao 0001, Bangyang Hong, Yang Li 0139, Xianli Zhang, Jialun Wu, Chunbao Wang 0002, Xiangrong Zhang, Tieliang Gong, Yefeng Zheng 0001, Deyu Meng, Chen Li 0011
Medical Image Anal.11
2023 Meta-Based Self-Training and Re-Weighting for Aspect-Based Sentiment Analysis
abstract
Aspect-based sentiment analysis (ABSA) means to identify fine-grained aspects, opinions, and sentiment polarities. Recent ABSA research focuses on utilizing multi-task learning (MTL) to achieve less computational costs and better performance. However, there are certain limits in MTL-based ABSA. For example, unbalanced labels and sub-task learning difficulties may result in the biases that some labels and sub-tasks are overfitting, while the others are underfitting. To address these issues, inspired by neuro-symbolic learning systems, we propose a meta-based self-training method with a meta-weighter (MSM). We believe that a generalizable model can be achieved by appropriate symbolic representation selection (in-domain knowledge) and effective learning control (regulation) in a neural system. Thus, MSM trains a teacher model to generate in-domain knowledge (e.g., unlabeled data selection and pseudo-label generation), where the generated pseudo-labels are used by a student model for supervised learning. Then, the meta-weighter of MSM is jointly trained with the student model to provide each instance with sub-task-specific weights to coordinate their convergence rates, balancing class labels, and alleviating noise impacts introduced from self-training. The following experiments indicate that MSM can utilize 50% labeled data to achieve comparable results to state-of-arts models in ABSA and outperform them with all labeled data.
Kai He 0001, Rui Mao 0010, Tieliang Gong, Chen Li 0011, Erik Cambria
IEEE Trans. Affect. Comput.4
2023 Optimal Randomized Approximations for Matrix-Based Rényi's Entropy
abstract
The Matrix-based Rényi’s entropy enables us to directly measure information quantities from given data without the costly probability density estimation of underlying distributions, thus has been widely adopted in numerous statistical learning and inference tasks. However, exactly calculating this new information quantity requires access to the eigenspectrum of a semi-positive definite (SPD) matrix$A$which grows linearly with the number of samples$n$, resulting in a$O(n^{3})$time complexity that is prohibitive for large-scale applications. To address this issue, this paper takes advantage of stochastic trace approximations for matrix-based Rényi’s entropy with arbitrary$\alpha \in \mathbb {R}^{+}$orders, lowering the complexity by converting the entropy approximation to a matrix-vector multiplication problem. Specifically, we develop random approximations for integer-order$\alpha $cases and polynomial series approximations (Taylor and Chebyshev) for fractional$\alpha $cases, leading to a$O(n^{2}sm)$overall time complexity, where$s, m \ll n$denote the number of vector queries and the polynomial order respectively. We theoretically establish statistical guarantees for all approximation algorithms and give explicit order of$s$and$m$with respect to the approximation error$\epsilon $, showing optimal convergence rate for both parameters up to a logarithmic factor. Large-scale simulations and real-world applications validate the effectiveness of the developed approximations, demonstrating remarkable speedup with negligible loss in accuracy.
Yuxin Dong 0003, Tieliang Gong, Shujian Yu, Chen Li 0011
IEEE Trans. Inf. Theory4
2023 Semi-Supervised Pixel Contrastive Learning Framework for Tissue Segmentation in Histopathological Image
abstract
Accurate tissue segmentation in histopathological images is essential for promoting the development of precision pathology. However, the size of the digital pathological image is great, which needs to be tiled into small patches containing limited semantic information. To imitate the pathologist's diagnosis process and model the semantic relation of the whole slide image, We propose a semi-supervised pixel contrastive learning framework (SSPCL) which mainly includes an uncertainty-guided mutual dual consistency learning module (UMDC) and a cross image pixel-contrastive learning module (CIPC). The UMDC module enables efficient learning from unlabeled data through mutual dual-consistency and consensus-based uncertainty. The CIPC module aims at capturing the cross-patch semantic relationship by optimizing a contrastive loss between pixel embeddings. We also propose several novel domain-related sampling methods by utilizing the continuous spatial structure of adjacent image patches, which can avoid the problem of false sampling and improve the training efficiency. In this way, SSPCL significantly reduces the labeling cost on histopathological images and realizes the accurate quantitation of tissues. Extensive experiments on three tissue segmentation datasets demonstrate the effectiveness of SSPCL, which outperforms state-of-the-art up to 5.0% in mDice.
Jiangbo Shi, Tieliang Gong, Chunbao Wang 0002, Chen Li 0011
IEEE J. Biomed. Health Informatics4
2023 Childhood Leukemia Classification via Information Bottleneck Enhanced Hierarchical Multi-Instance Learning
abstract
Leukemia classification relies on a detailed cytomorphological examination of Bone Marrow (BM) smear. However, applying existing deep-learning methods to it is facing two significant limitations. Firstly, these methods require large-scale datasets with expert annotations at the cell level for good results and typically suffer from poor generalization. Secondly, they simply treat the BM cytomorphological examination as a multi-class cell classification task, thus failing to exploit the correlation among leukemia subtypes over different hierarchies. Therefore, BM cytomorphological estimation as a time-consuming and repetitive process still needs to be done manually by experienced cytologists. Recently, Multi-Instance Learning (MIL) has achieved much progress in data-efficient medical image processing, which only requires patient-level labels (which can be extracted from the clinical reports). In this paper, we propose a hierarchical MIL framework and equip it with Information Bottleneck (IB) to tackle the above limitations. First, to handle the patient-level label, our hierarchical MIL framework uses attention-based learning to identify cells with high diagnostic values for leukemia classification in different hierarchies. Then, following the information bottleneck principle, we propose a hierarchical IB to constrain and refine the representations of different hierarchies for better accuracy and generalization. By applying our framework to a large-scale childhood acute leukemia dataset with corresponding BM smear images and clinical reports, we show that it can identify diagnostic-related cells without the need for cell-level annotations and outperforms other comparison methods. Furthermore, the evaluation conducted on an independent test cohort demonstrates the high generalizability of our framework.
Zeyu Gao 0001, Anyu Mao, Kefei Wu, Yang Li 0139, Liebin Zhao, Xianli Zhang, Jialun Wu, Lisha Yu, Tieliang Gong, Yefeng Zheng 0001, Deyu Meng, Chen Li 0011
IEEE Trans. Medical Imaging14
2023 MG-Trans: Multi-Scale Graph Transformer With Information Bottleneck for Whole Slide Image Classification
abstract
Multiple instance learning (MIL)-based methods have become the mainstream for processing the megapixel-sized whole slide image (WSI) with pyramid structure in the field of digital pathology. The current MIL-based methods usually crop a large number of patches from WSI at the highest magnification, resulting in a lot of redundancy in the input and feature space. Moreover, the spatial relations between patches can not be sufficiently modeled, which may weaken the model's discriminative ability on fine-grained features. To solve the above limitations, we propose a Multi-scale Graph Transformer (MG-Trans) with information bottleneck for whole slide image classification. MG-Trans is composed of three modules: patch anchoring module (PAM), dynamic structure information learning module (SILM), and multi-scale information bottleneck module (MIBM). Specifically, PAM utilizes the class attention map generated from the multi-head self-attention of vision Transformer to identify and sample the informative patches. SILM explicitly introduces the local tissue structure information into the Transformer block to sufficiently model the spatial relations between patches. MIBM effectively fuses the multi-scale patch features by utilizing the principle of information bottleneck to generate a robust and compact bag-level representation. Besides, we also propose a semantic consistency loss to stabilize the training of the whole model. Extensive studies on three subtyping datasets and seven gene mutation detection datasets demonstrate the superiority of MG-Trans.
Jiangbo Shi, Lufei Tang, Zeyu Gao 0001, Yang Li 0139, Chunbao Wang 0002, Tieliang Gong, Chen Li 0011, Huazhu Fu
IEEE Trans. Medical Imaging7
2023 A Structure-Aware Hierarchical Graph-Based Multiple Instance Learning Framework for pT Staging in Histopathological Image
abstract
Pathological primary tumor (pT) stage focuses on the infiltration degree of the primary tumor to surrounding tissues, which relates to the prognosis and treatment choices. The pT staging relies on the field-of-views from multiple magnifications in the gigapixel images, which makes pixel-level annotation difficult. Therefore, this task is usually formulated as a weakly supervised whole slide image (WSI) classification task with the slide-level label. Existing weakly-supervised classification methods mainly follow the multiple instance learning paradigm, which takes the patches from single magnification as the instances and extracts their morphological features independently. However, they cannot progressively represent the contextual information from multiple magnifications, which is critical for pT staging. Therefore, we propose a structure-aware hierarchical graph-based multi-instance learning framework (SGMF) inspired by the diagnostic process of pathologists. Specifically, a novel graph-based instance organization method is proposed, namely structure-aware hierarchical graph (SAHG), to represent the WSI. Based on that, we design a novel hierarchical attention-based graph representation (HAGR) network to capture the critical patterns for pT staging by learning cross-scale spatial features. Finally, the top nodes of SAHG are aggregated by a global attention layer for bag-level representation. Extensive studies on three large-scale multi-center pT staging datasets with two different cancer types demonstrate the effectiveness of SGMF, which outperforms state-of-the-art up to 5.6% in the F1 score.
Jiangbo Shi, Lufei Tang, Yang Li 0139, Xianli Zhang, Zeyu Gao 0001, Yefeng Zheng 0001, Chunbao Wang 0002, Tieliang Gong, Chen Li 0011
IEEE Trans. Medical Imaging9
2022 Regularized Modal Regression on Markov-Dependent Observations: A Theoretical Assessment
abstract
Modal regression, a widely used regression protocol, has been extensively investigated in statistical and machine learning communities due to its robustness to outlier and heavy-tailed noises. Understanding modal regression's theoretical behavior can be fundamental in learning theory. Despite significant progress in characterizing its statistical property, the majority results are based on the assumption that samples are independent and identical distributed (i.i.d.), which is too restrictive for real-world applications. This paper concerns about the statistical property of regularized modal regression (RMR) within an important dependence structure - Markov dependent. Specifically, we establish the upper bound for RMR estimator under moderate conditions and give an explicit learning rate. Our results show that the Markov dependence impacts on the generalization error in the way that sample size would be discounted by a multiplicative factor depending on the spectral gap of the underlying Markov chain. This result shed a new light on characterizing the theoretical underpinning for robust regression.
Tieliang Gong, Yuxin Dong 0003, Hong Chen 0004, Wei Feng 0010, Bo Dong 0001, Chen Li 0011
AAAI6
2022 Uncertainty-based Model Acceleration for Cancer Classification in Whole-Slide Images
abstract
Computational Pathology (CPATH) offers the possibility for highly accurate and low-cost automated pathological diagnosis. However, the high time cost of model inference is one of the main issues limiting the application of CPATH methods. Due to the large size of Whole-Slide Image (WSI), commonly used CPATH methods divided a WSI into a large number of image patches at relatively high magnification, then predicted each image patch individually, which is time-consuming. In this paper, we propose a novel Uncertainty-based Model Acceleration (UMA) method for reducing the time cost of model inference, thereby relieving the deployment burden of CPATH applications. Enlightened by the slide-viewing process of pathologists, only a few high-uncertain regions are regarded as “suspicious” regions that need to be predicted at high magnification, and most of the regions in WSI are predicted at low magnification, thereby reducing the times of image patch extraction and prediction. Meanwhile, uncertainty estimation ensures prediction accuracy at low magnification. We take two fundamental CPATH classification tasks (i.e., cancer region detection and subtyping) as examples. Extensive experiments on two large-scale renal cell carcinoma classification datasets demonstrate that our UMA can significantly reduce the time cost of model inference while maintaining competitive classification performance.
Zeyu Gao 0001, Anyu Mao, Jialun Wu, Yang Li 0139, Chunbao Wang 0002, Caixia Ding, Tieliang Gong, Chen Li 0011
BIBM8
2022 Knowledge Enhanced Coreference Resolution via Gated Attention
abstract
Coreference resolution aims at linking all mentions that refer to the same entity, which are widely adopted in many biomedical and bioinformatics tasks, such as biomedical knowledge graph construction and metabolic pathway integration. Many recent studies focus on improving neural model structures. However, we argue that a practical method that integrates commonsense knowledge can further improve coreference resolution performance, because commonsense delivers extra prior knowledge for reasoning and can enhance related representations, rather than naive mention-context occurrence modeling. In this work, we propose an effective method to integrate external commonsense knowledge into a neural coreference resolution model. Specially, a gated attention mechanism is employed in our method to leverage commonsense according to different contexts. By using ConceptNet as the knowledge base in three span-ranking backbone models, the models can yield significant performance gains on used datasets. We also achieve improvements in tasks of long-term mention detection and cross-sentence coreferences after incorporating knowledge.
Kai He 0001, Yufei Li 0002, Tieliang Gong, Chen Li 0011, Jialun Wu
BIBM6
2022 Uncertainty-guided Mutual Consistency Training for Semi-supervised Biomedical Relation Extraction
abstract
Biomedical relation extraction seeks to automatically extract biomedical relations from biomedical text, which plays an important role in biomedical studies. However, constructing high-quality biomedical annotation data is not only time-consuming but also requires a high level of knowledge in the biomedical field. To alleviate this problem, Semi-supervised Biomedical Relation Extraction aims to extract relation facts from the limited labeled data and the more readily available unlabeled samples. Existing works can be roughly categorized as self-training methods and self-ensembling methods. The former aims to generate pseudo labels, which may lead to the gradual drift problem. The latter aims to encourage the output of one model to be consistent with the other model, where the acquisition of the model is tedious. To alleviate these issues, we propose a novel Uncertainty-Guided Mutual Consistency Training framework(UG-MCT) for semi-supervised Biomedical relation extraction. Specifically, our framework consists of two models with the same structure, which differ only when updating their weights, and then an intersecting pseudo-label mechanism is designed to convert the prediction discrepancies of the two models into mutual consistency training loss, thus promoting the consistency of model predictions. In addition, we utilize uncertainty as guided information to assist the model in focusing on the confident pseudo labels and mitigate the noise of inaccurate pseudo labeling during training. Thus, our model is very simple and efficient while mitigating the noise introduced by pseudo-labels. UG-MCT is evaluated on multiple datasets in different settings and the experimental results demonstrate that our method is highly effective in semi-supervised biomedical relation extraction compared to the state-of-the-art.
Chang Jia, Kai He 0001, Jialun Wu, Tieliang Gong, Chen Li 0011
BIBM7
2022 Leveraging Multiple Types of Domain Knowledge for Safe and Effective Drug Recommendation
abstract
Predicting drug combinations according to patients' electronic health records is an essential task in intelligent healthcare systems, which can assist clinicians in ordering safe and effective prescriptions. However, existing work either missed/underutilized the important information lying in the drug molecule structure in drug encoding or has insufficient control over Drug-Drug Interactions (DDIs) rates within the predictions. To address these limitations, we propose CSEDrug, which enhances the drug encoding and DDIs controlling by leveraging multi-faceted drug knowledge, including molecule structures of drugs, Synergistic DDIs (SDDIs), and Antagonistic DDIs (ADDIs). We integrate these types of knowledge into CSEDrug by a graph-based drug encoder and multiple loss functions, including a novel triplet learning loss and a comprehensive DDI controllable loss. We evaluate the performance of CSEDrug in terms of accuracy, effectiveness, and safety on the public MIMIC-III dataset. The experimental results demonstrate that CSEDrug outperforms several state-of-the-art methods and achieves a 2.93% and a 2.77% increase in the Jaccard similarity scores and F1 scores, meanwhile, a 0.68% reduction of the ADDI rate (safer drug combinations), and 0.69% improvement of the SDDI rate (more effective drug combinations).
Jialun Wu, Buyue Qian, Yang Li 0139, Zeyu Gao 0001, Meizhi Ju, Yifan Yang 0008, Yefeng Zheng 0001, Tieliang Gong, Chen Li 0011, Xianli Zhang
CIKM9
2022 COPNER: Contrastive Learning with Prompt Guiding for Few-shot Named Entity Recognition
abstract
Distance metric learning has become a popular solution for few-shot Named Entity Recognition (NER). The typical setup aims to learn a similarity metric for measuring the semantic similarity between test samples and referents, where each referent represents an entity class. The effect of this setup may, however, be compromised for two reasons. First, there is typically a limited optimization exerted on the representations of entity tokens after initing by pre-trained language models. Second, the referents may be far from representing corresponding entity classes due to the label scarcity in the few-shot setting. To address these challenges, we propose a novel approach named COntrastive learning with Prompt guiding for few-shot NER (COPNER). We introduce a novel prompt composed of class-specific words to COPNER to serve as 1) supervision signals for conducting contrastive learning to optimize token representations; 2) metric referents for distance-metric inference on test samples. Experimental results demonstrate that COPNER outperforms state-of-the-art models with a significant margin in most cases. Moreover, COPNER shows great potential in the zero-shot setting.
Kai He 0001, Xianli Zhang, Tieliang Gong, Rui Mao 0010, Chen Li 0011
COLING7
2022 Learning Representations from Local to Global for Fine-grained Patient Similarity Measuring in Intensive Care Unit
abstract
Patient similarity measurement is an essential step in discovering clinically meaningful subgroups and building case retrieval systems. Most existing studies implement this procedure using similarity measurement algorithms on the multivariate clinical time-series (input space) or the low-dimensional patient representation (representation space) learned by a representation learning model. However, they either suffer from the adverse effects of irrelevant variables in the data or fail to assess the fine-grained similarity underneath the disease progress. In this paper, we propose a method to measure more fine-grained patient similarity in the state space, where each patient is represented by a series of state representations that reveal the dynamic health status. We discuss three desiderata, including stability, personality, and interpretability, for the state representations, and on this basis, develop a supervised predictive model that learns good state representations for identifying similar patients and predicting patient outcomes. Experimental results on the publicly available dataset MIMIC-III show that our method offers a promising direction for precisely identifying similar patients at the state trajectory level, as well as accurately predicting outcomes.
Xianli Zhang, Buyue Qian, Yang Li 0139, Zeyu Gao 0001, Chong Guan, Renzhen Wang, Yefeng Zheng 0001, Hansen Zheng, Chen Li 0011
ICDM9
2022 JCBIE: a joint continual learning neural network for biomedical information extraction
abstract
Extracting knowledge from heterogeneous data sources is fundamental for the construction of structured biomedical knowledge graphs (BKGs), where entities and relations are represented as nodes and edges in the graphs, respectively. Previous biomedical knowledge extraction methods simply considered limited entity types and relations by using a task-specific training set, which is insufficient for large-scale BKGs development and downstream task applications in different scenarios. To alleviate this issue, we propose a joint continual learning biomedical information extraction (JCBIE) network to extract entities and relations from different biomedical information datasets. By empirically studying different joint learning and continual learning strategies, the proposed JCBIE can learn and expand different types of entities and relations from different datasets. JCBIE uses two separated encoders in joint-feature extraction, hence can effectively avoid the feature confusion problem comparing with using one hard-parameter sharing encoder. Specifically, it allows us to adopt entity augmented inputs to establish the interaction between named entity recognition and relation extraction. Finally, a novel evaluation mechanism is proposed for measuring cross-corpus generalization errors, which was ignored by traditional evaluation methods. Our empirical studies show that JCBIE achieves promising performance when continual learning strategy is adopted with multiple corpora.
Kai He 0001, Rui Mao 0010, Tieliang Gong, Erik Cambria, Chen Li 0011
BMC Bioinform.5
2022 Semantic Attention and Scale Complementary Network for Instance Segmentation in Remote Sensing Images
abstract
In this article, we focus on the challenging multicategory instance segmentation problem in remote sensing images (RSIs), which aims at predicting the categories of all instances and localizing them with pixel-level masks. Although many landmark frameworks have demonstrated promising performance in instance segmentation, the complexity in the background and scale variability instances still remain challenging, for instance, segmentation of RSIs. To address the above problems, we propose an end-to-end multicategory instance segmentation model, namely, the semantic attention (SEA) and scale complementary network, which mainly consists of a SEA module and a scale complementary mask branch (SCMB). The SEA module contains a simple fully convolutional semantic segmentation branch with extra supervision to strengthen the activation of interest instances on the feature map and reduce the background noise's interference. To handle the undersegmentation of geospatial instances with large varying scales, we design the SCMB that extends the original single mask branch to trident mask branches and introduces complementary mask supervision at different scales to sufficiently leverage the multiscale information. We conduct comprehensive experiments to evaluate the effectiveness of our proposed method on the iSAID dataset and the NWPU Instance Segmentation dataset and achieve promising performance.
Tianyang Zhang 0002, Xiangrong Zhang, Peng Zhu 0004, Xu Tang 0004, Chen Li 0011, Licheng Jiao, Huiyu Zhou 0001
IEEE Trans. Cybern.5
2022 Recurrent Attention and Semantic Gate for Remote Sensing Image Captioning
abstract
The remote sensing image captioning has attracted wide spread attention in remote sensing field due to its application potentiality. However, most existing approaches model limited interactions between image content and sentence and fail to exploit special characteristics of the remote sensing images. We introduce a novel recurrent attention and semantic gate (RASG) framework to facilitate the remote sensing image captioning in this article, which integrates competitive visual features and a recurrent attention mechanism to generate a better context vector for the images every time as well as enhances the representations of the current word state. Specifically, we first project each image into competitive visual features by taking the advantage of both static visual features and multiscale features. Then, a novel recurrent attention mechanism is developed to extract the high-level attentive maps from encoded features and nonvisual features, which can help the decoder recognize and focus on the effective information for understanding the complex content of the remote sensing images. Finally, the hidden states from the long short-term memory (LSTM) and other semantic references are incorporated into a semantic gate, which contributes to more comprehensive and precise semantic understanding. Comprehensive experiments on three widely used datasets, Sydney-Captions, UCM-Captions, and Remote Sensing Image Captioning Dataset, have demonstrated the superiority of the proposed RASG over a series of attentive models based on image captioning methods.
Yunpeng Li 0010, Xiangrong Zhang, Chen Li 0011, Xin Wang 0068, Xu Tang 0004, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.4
2022 Foreground Refinement Network for Rotated Object Detection in Remote Sensing Images
abstract
Object detection has been a fundamental task in the field of remote sensing and has made considerable progress in recent years. However, the high background complexity in remote sensing images (RSIs) remains challenging. In this article, we propose a refined rotation detector, namely, the Foreground Refinement Network (FoRDet), to alleviate the above problem by leveraging the information of foreground regions from the perspectives of feature and optimization. Specifically, we propose a foreground relation module (FRL) that aggregates the foreground-contextual representations from the coarse stage and improves the discrimination of foreground regions on feature maps in the refined stage. Besides, considering the risk of the potential foreground anchors being overwhelmed in the training phase, we design a foreground anchor reweighting (FRW) loss that integrates the classification confidence and localization accuracy of each foreground anchor from the coarse stage to dynamically regulate their contributions in the refined stage, which highlights the potential foreground anchors. The comprehensive experimental results on three public datasets for rotated object detection DOTA, HRSC2016, and UCAS-AOD demonstrate the effectiveness of our proposed method.
Tianyang Zhang 0002, Xiangrong Zhang, Peng Zhu 0004, Puhua Chen, Xu Tang 0004, Chen Li 0011, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.6
2022 Unsupervised Representation Learning for Tissue Segmentation in Histopathological Images: From Global to Local Contrast
abstract
Tissue segmentation is an essential task in computational pathology. However, relevant datasets for such a pixel-level classification task are hard to obtain due to the difficulty of annotation, bringing obstacles for training a deep learning-based segmentation model. Recently, contrastive learning has provided a feasible solution for mitigating the heavy reliance of deep learning models on annotation. Nevertheless, applying contrastive loss to the most abstract image representations, existing contrastive learning frameworks focus on global features, therefore, are less capable of encoding finer-grained features (e.g., pixel-level discrimination) for the tissue segmentation task. Enlightened by domain knowledge, we design three contrastive learning tasks with multi-granularity views (from global to local) for encoding necessary features into representations without accessing annotations. Specifically, we construct: (1) an image-level task to capture the difference between tissue components, i.e., encoding the component discrimination; (2) a superpixel-level task to learn discriminative representations of local regions with different tissue components, i.e., encoding the prototype discrimination; (3) a pixel-level task to encourage similar representations of different tissue components within a local region, i.e., encoding the spatial smoothness. Through our global-to-local pre-training strategy, the learned representations can reasonably capture the domain-specific and fine-grained patterns, making them easily transferable to various tissue segmentation tasks in histopathological images. We conduct extensive experiments on two tissue segmentation datasets, while considering two real-world scenarios with limited or sparse annotations. The experimental results demonstrate that our framework is superior to existing contrastive learning methods and can be easily combined with weakly supervised and semi-supervised segmentation methods.
Zeyu Gao 0001, Chang Jia, Yang Li 0139, Xianli Zhang, Bangyang Hong, Jialun Wu, Tieliang Gong, Chunbao Wang 0002, Deyu Meng, Yefeng Zheng 0001, Chen Li 0011
IEEE Trans. Medical Imaging11
2021 AEFNet: Adaptive Scale Feature Based on Elastic-and-Funnel Neural Network for Healthcare Representation
abstract
Healthcare Representation learning has been a key element to achieving state-of-the-art performance on healthcare prediction. Recent advances based Electronic Healthcare Records(EHRs) are mostly devoted to extracting temporal progression patterns with temporal model and their variants. Although these works have shown excellent performances in healthcare prediction, the unified temporal pattern may not be suitable for individuals in all healthcare conditions. Moreover, some studies ususally introduce complex Deep Neural Networks models and medical prior knowledge to get compact representation, causing great computational burden. In this paper, we propose a general health care representation model, named AEFNet. We only leverage three simple convolution operations and a set of up and down sampling to ensure performance and model complexity equally, which achieves adaptively extract distinct individual key feature in a light manner. AEFNet can shrink and refine highly suitable scale information adaptively and comletely. Breaking traditional fixed convolution scale or multi-scale, AEFNet achieves scale adaptively to extract the most significant information and context relationship. Finally, We validate our method on the public dataset MIMIC-III, and the evaluation results indicate that our method can significantly outperform other remarkable baseline models.
Jialun Wu, Yuhua Wei, Chen Li 0011, Tieliang Gong
BIBM5
2021 W-Net: A Two-Stage Convolutional Network for Nucleus Detection in Histopathology Image
abstract
Pathological diagnosis is the gold standard for cancer diagnosis, but it is labor-intensive, in which tasks such as cell detection, classification, a nd c ounting a re particularly prominent. A common solution for automating these tasks is using nucleus segmentation technology. However, it is hard to train a robust nucleus segmentation model, due to several challenging problems,i.e., the nucleus adhesion, stacking, and excessive fusion with the background. Recently, some researchers proposed a series of automatic nucleus segmentation methods based on point annotation, which can significant i mprove t he m odel performance. Nevertheless, the point annotation needs to be marked by experienced pathologists. In order to take advantage of segmentation methods based on point annotation, further alleviate the manual workload, and make cancer diagnosis more efficient and accurate, it is necessary to develop an automatic nucleus detection algorithm, which can automatically and efficiently l ocate the position of the nucleus in the pathological image and extract valuable information for pathologists. In this paper, we propose a W-shaped network for automatic nucleus detection. Different from the traditional U-Net based method, mapping the original pathology image to the target mask directly, our proposed method split the detection task into two sub-tasks. The first sub-task maps the original pathology image to the binary mask, then the binary mask is mapped to the density mask in the second subtask. After the task is split, the task's difficulty i s significantly reduced, and the network's overall performance is improved. Our proposed network can automatic identify the center of each nucleus. Combined with the NuClick, a semi-supervised recognition model based on point annotation, we implement a fully automatic nucleus annotation framework.
Anyu Mao, Jialun Wu, Xinrui Bao, Zeyu Gao 0001, Tieliang Gong, Chen Li 0011
BIBM6
2021 Meta Mask Correction for Nuclei Segmentation in Histopathological Image
abstract
Nuclei segmentation is a fundamental task in digital pathology analysis and can be automated by deep learning-based methods. However, the development of such an automated method requires a large amount of data with precisely annotated masks which is hard to obtain. Training with weakly labeled data is a popular solution for reducing the workload of annotation. In this paper, we propose a novel meta-learning-based nuclei segmentation method which follows the label correction paradigm to leverage data with noisy masks. Specifically, we design a fully conventional meta-model that can correct noisy masks using a small amount of clean meta-data. Then the corrected masks can be used to supervise the training of the segmentation model. Meanwhile, a bi-level optimization method is adopted to alternately update the parameters of the main segmentation model and the meta-model in an end-to-end way. Extensive experimental results on two nuclear segmentation datasets show that our method achieves the state-of-the-art result. It even achieves comparable performance with the model training on supervised data in some noisy settings.
Jiangbo Shi, Chang Jia, Zeyu Gao 0001, Tieliang Gong, Chunbao Wang 0002, Chen Li 0011
BIBM6
2021 PIMIP: An Open Source Platform for Pathology Information Management and Integration
abstract
Digital pathology plays a crucial role in the development of artificial intelligence in the medical field. The digital pathology platform can make the pathological resources digital and networked, and realize the permanent storage of visual data and the synchronous browsing processing without the limitation of time and space. It has been widely used in various fields of pathology. However, there is still a lack of an open and universal digital pathology platform to assist doctors in the management and analysis of digital pathological sections, as well as the management and structured description of relevant patient information. Most platforms cannot integrate image viewing, annotation and analysis, and text information management. To solve the above problems, we propose a comprehensive and extensible platform, PIMIP (Pathology Information Management & Integration Platform). PIMIP has developed the image annotation functions based on the visualization of digital pathological sections. Our annotation functions support multi-user collaborative annotation and multi-device annotation, and realize the automation of some annotation tasks. In the annotation task, we invited a professional pathologist for guidance. We introduce a machine learning module for image analysis. The data we collected included public data from local hospitals and clinical examples. Our platform is more clinical and suitable for clinical use. In addition to image data, we also structured the management and display of text information. So our platform is comprehensive. The platform framework is built in a modular way to support users to add machine learning modules independently, which makes our platform extensible.
Jialun Wu, Anyu Mao, Xinrui Bao, Haichuan Zhang 0001, Zeyu Gao 0001, Chunbao Wang 0002, Tieliang Gong, Chen Li 0011
BIBM8
2021 A Precision Diagnostic Framework of Renal Cell Carcinoma on Whole-Slide Images using Deep Learning
abstract
Diagnostic pathology, which is the basis and gold standard of cancer diagnosis, provides essential information on the prognosis of the disease and vital evidence for clinical treatment. However, pathological diagnosis is subjective, and differences in observation and diagnosis between pathologists are common. This phenomenon is more evident in hospitals with insufficient medical resources. Deep learning (DL) can be used to identify and classify structures in digital pathology. In order to solve the above difficulties, in this work, we propose a DL framework for generating pathological diagnosis by analyzing histopathological images of renal cell carcinoma. A deep neural network is trained on a large high-quality annotated dataset for accurate tumor area detection, subtyping, and grading. The results show that our framework has achieved pathologist-level accuracy in diagnosis, can generate pathology reports with tumor indicators, and provide pathologists with interpretable auxiliary diagnoses
Jialun Wu, Tieliang Gong, Xinrui Bao, Zeyu Gao 0001, Haichuan Zhang 0001, Chunbao Wang 0002, Chen Li 0011
BIBM8
2021 BioIE: Biomedical Information Extraction with Multi-head Attention Enhanced Graph Convolutional Network
abstract
Constructing large-scaled medical knowledge graphs (MKGs) can significantly boost healthcare applications for medical surveillance, bring much attention from recent research. An essential step in constructing large-scale MKG is extracting information from medical reports. Recently, information extraction techniques have been proposed and show promising performance in biomedical information extraction. However, these methods only consider limited types of entity and relation due to the noisy biomedical text data with complex entity correlations. Thus, they fail to provide enough information for constructing MKGs and restrict the downstream applications. To address this issue, we propose Biomedical Information Extraction (BioIE), a hybrid neural network to extract relations from biomedical text and unstructured medical reports. Our model utilizes a multi-head attention enhanced graph convolutional network (GCN) to capture the complex relations and context information while resisting the noise from the data. We evaluate our model on two major biomedical relationship extraction tasks, chemical-disease relation (CDR) and chemical-protein interaction (CPI), and a cross-hospital pan-cancer pathology report corpus. The results show that our method achieves superior performance than baselines. Furthermore, we evaluate the applicability of our method under a transfer learning setting and show that BioIE achieves promising performance in processing medical text from different formats and writing styles.
Jialun Wu, Tieliang Gong, Chunbao Wang 0002, Chen Li 0011
BIBM6
2021 A Personalized Diagnostic Generation Framework Based on Multi-source Heterogeneous Data
abstract
Personalized diagnoses have not been possible due to a sear amount of data pathologists have to bear during the day-to-day routine, leading to the current generalized standards being continuously updated as new findings are reported. It is noticeable that these practical standards are developed based on multi-source heterogeneous data, including whole-slide images and pathology and clinical reports. In this study, we propose a framework that combines pathological images and medical reports to generate a personalized diagnosis result for an individual patient. We use nuclei-level image feature similarity and content-based deep learning method to search for a personalized group of populations with similar pathological characteristics, extract structured prognostic information from descriptive pathology reports of the similar patient population, and assign importance of different prognostic factors to generate a personalized pathological diagnosis result. We use multi-source heterogeneous data from TCGA (The Cancer Genome Atlas) database. The result demonstrates that our framework matches the performance of pathologists in the diagnosis of renal cell carcinoma. This framework is designed to be generic, and this could be applied to other types of cancer. The weights could provide insights into the known prognostic factors and further guide more precise clinical treatment protocols.
Jialun Wu, Tieliang Gong, Haichuan Zhang 0001, Chunbao Wang 0002, Chen Li 0011
BIBM6
2021 BaT: Beat-aligned Transformer for Electrocardiogram Classification
abstract
Electrocardiogram (ECG) is one of the critical diagnostic tools in healthcare. Various deep learning models, except Transformers, have been explored and applied to map ECG patterns to heart abnormalities. Transformer models have been adopted from natural language processing to computer vision with advanced features. Most recently, vision transformers show exceptional performances, even on moderate-scale datasets. However, naively applying vision transformers on electrocardiogram datasets leads to poor results. In this paper, we propose a novel network called Beat-aligned Transformer (BaT), a hierarchical Transformer that sufficiently exploits the cyclicity of ECG. We organize and treat an input ECG as multiple aligned beats instead of a single time series. In the BaT, shifted-window-based Transformer blocks (SW Block) are adopted to learn the representation for each beat, and aggregation blocks are designed to exchange information among the beat representations. Nested SW Blocks and aggregation blocks form a beat-aware hierarchical structure of BaT. In this way, the new data format and the BaT hierarchical structure boost Transformer performance on ECG classification. From the experiments on public ECG datasets, we observe BaT outperforms other Transformer-based models and achieves competitive performance compared with other state-of-the-art methods.
Xiaoyu Li 0007, Chen Li 0011, Yuhua Wei, Yuyao Sun, Jishang Wei, Xiang Li 0013, Buyue Qian
ICDM2
2021 Towards Interpretability and Personalization: A Predictive Framework for Clinical Time-series Analysis
abstract
Clinical time-series is receiving long-term attention in data mining and machine learning communities and has boosted a variety of data-driven applications. Identifying similar patients or subgroups from clinical time-series is an essential step to design tailored treatments in clinical practice. However, most of the existing methods are either purely unsupervised that tend to neglect the patient outcome information or cannot generate personalized patient representation through supervised learning, thus may fail to identify ‘truly similar patients’ (i.e., patients who similar in both outcomes and individual outcome-related clinical variables). To tackle these limitations, we propose a novel predictive clinical time-series analysis framework. Specifically, our framework uses task-specific information to rule out the task-irrelevant factors in each patient data individually and generates the contribution scores that reveal the factors’ importance for the patient outcome. Then a patient representation construction method is proposed to generate task-related and personalized representations by combining remained factors and their contribution scores. At last, similarity measurement or cluster analysis can be conducted. We evaluate our framework on three real-world clinical time-series datasets, empirically demonstrate that our framework achieves improvements in prediction performance, similarity measurement, and clustering, thus potentially benefiting patient-similarity-based precision medicine applications.
Yang Li 0139, Xianli Zhang, Buyue Qian, Zeyu Gao 0001, Chong Guan, Yefeng Zheng 0001, Hansen Zheng, Fenglang Wu, Chen Li 0011
ICDM9
2021 Learning to Reweight Samples with Offline Loss Sequence
abstract
Deep neural networks (DNNs) provide the best of class solutions to many supervised tasks due to their powerful function fitting capabilities. However, it is challenging to handle data bias, such as label noise and class imbalance, when applying DNNs to solve real-world problems. Sample reweighting is a popular strategy to tackle data bias, which assigns higher weights to informative samples or samples with clean labels. However, conventional reweighting methods require prior knowledge of the distribution information of data bias, which is intractable in practice. In recent years, meta-learning-based methods have been proposed to learn to assign weights to training samples adaptively by using their online training loss or gradient directions. However, the latent bias distribution cannot be adequately characterized in an online fashion. The online loss distribution changes over the training procedure, making it even harder to perform the sample weight learning. In contrast to past methods, we propose a two-stage training strategy to tackle the above problems. In the first stage, the loss sequences of samples are collected. In the second stage, a subnet with convolutional layers is utilized to learn the mapping from offline sample loss sequence to sample weight adaptively. Guided by a small unbiased meta dataset, this subnet is optimized iteratively with the main classifier network in a meta-learning manner. Empirical results show that our method, called Meta Reweighting with Offline Loss Sequence (MROLS), outperforms state-of-the-art reweighting techniques on most benchmarks. Moreover, the weights of training samples learned via MROLS can be well utilized by other classifiers, which can directly enhance the standard training schema. Our source code is available at https://github.com/Neronjust2017/MROLS.
Yuhua Wei, Xiaoyu Li 0007, Jishang Wei, Buyue Qian, Chen Li 0011
ICDM5
2021 Instance-Based Vision Transformer for Subtyping of Papillary Renal Cell Carcinoma in Histopathological Image
Zeyu Gao 0001, Bangyang Hong, Xianli Zhang, Yang Li 0139, Chang Jia, Jialun Wu, Chunbao Wang 0002, Deyu Meng, Chen Li 0011
MICCAI (8)9
2021 Nuclei Grading of Clear Cell Renal Cell Carcinoma in Histopathological Image by Composite High-Resolution Network
Zeyu Gao 0001, Jiangbo Shi, Xianli Zhang, Yang Li 0139, Haichuan Zhang 0001, Jialun Wu, Chunbao Wang 0002, Deyu Meng, Chen Li 0011
MICCAI (8)9
2021 Adaptive Affinity Loss and Erroneous Pseudo-Label Refinement for Weakly Supervised Semantic Segmentation
abstract
Semantic segmentation has been continuously investigated in the last ten years, and majority of the established technologies are based on supervised models. In recent years, image-level weakly supervised semantic segmentation (WSSS), including single- and multi-stage process, has attracted large attention due to data labeling efficiency. In this paper, we propose to embed affinity learning of multi-stage approaches in a single-stage model. To be specific, we introduce an adaptive affinity loss to thoroughly learn the local pairwise affinity. As such, a deep neural network is used to deliver comprehensive semantic information in the training phase, whilst improving the performance of the final prediction module. On the other hand, considering the existence of errors in the pseudo labels, we propose a novel label reassign loss to mitigate over-fitting. Extensive experiments are conducted on the PASCAL VOC 2012 dataset to evaluate the effectiveness of our proposed approach that outperforms other standard single-stage methods and achieves comparable performance against several multi-stage methods.
Xiangrong Zhang, Zelin Peng, Peng Zhu 0004, Tianyang Zhang 0002, Chen Li 0011, Huiyu Zhou 0001, Licheng Jiao
ACM Multimedia5
2021 Learning Robust Patient Representations from Multi-modal Electronic Health Records: A Supervised Deep Learning Approach
Xianli Zhang, Buyue Qian, Yang Li 0139, Xi Chen 0003, Chong Guan, Chen Li 0011
SDM7
2021 Knowledge enhanced LSTM for coreference resolution on biomedical texts
abstract
MOTIVATION: Bio-entity Coreference Resolution focuses on identifying the coreferential links in biomedical texts, which is crucial to complete bio-events' attributes and interconnect events into bio-networks. Previously, as one of the most powerful tools, deep neural network-based general domain systems are applied to the biomedical domain with domain-specific information integration. However, such methods may raise much noise due to its insufficiency of combining context and complex domain-specific information. RESULTS: In this article, we explore how to leverage the external knowledge base in a fine-grained way to better resolve coreference by introducing a knowledge-enhanced Long Short Term Memory network (LSTM), which is more flexible to encode the knowledge information inside the LSTM. Moreover, we further propose a knowledge attention module to extract informative knowledge effectively based on contexts. The experimental results on the BioNLP and CRAFT datasets achieve state-of-the-art performance, with a gain of 7.5 F1 on BioNLP and 10.6 F1 on CRAFT. Additional experiments also demonstrate superior performance on the cross-sentence coreferences. AVAILABILITY AND IMPLEMENTATION: The source code will be made available at https://github.com/zxy951005/KB-CR upon publication. Data is avaliable at http://2011.bionlp-st.org/ and https://github.com/UCDenver-ccp/CRAFT/releases/tag/v3.1.3. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Yufei Li 0002, Xiaoyong Ma, Pengzhen Cheng, Kai He 0001, Chen Li 0011
Bioinform.6
2021 A representation model for biological entities by fusing structured axioms with unstructured texts
abstract
MOTIVATION: Structured semantic resources, for example, biological knowledge bases and ontologies, formally define biological concepts, entities and their semantic relationships, manifested as structured axioms and unstructured texts (e.g. textual definitions). The resources contain accurate expressions of biological reality and have been used by machine-learning models to assist intelligent applications like knowledge discovery. The current methods use both the axioms and definitions as plain texts in representation learning (RL). However, since the axioms are machine-readable while the natural language is human-understandable, difference in meaning of token and structure impedes the representations to encode desirable biological knowledge. RESULTS: We propose ERBK, a RL model of bio-entities. Instead of using the axioms and definitions as a textual corpus, our method uses knowledge graph embedding method and deep convolutional neural models to encode the axioms and definitions respectively. The representations could not only encode more underlying biological knowledge but also be further applied to zero-shot circumstance where existing approaches fall short. Experimental evaluations show that ERBK outperforms the existing methods for predicting protein-protein interactions and gene-disease associations. Moreover, it shows that ERBK still maintains promising performance under the zero-shot circumstance. We believe the representations and the method have certain generality and could extend to other types of bio-relation. AVAILABILITY AND IMPLEMENTATION: The source code is available at the gitlab repository https://gitlab.com/BioAI/erbk. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Peiliang Lou, Yuxin Dong 0003, Antonio Jimeno-Yepes, Chen Li 0011
Bioinform.4
2021 GRS-Det: An Anchor-Free Rotation Ship Detector Based on Gaussian-Mask in Remote Sensing Images
abstract
Ship detection is a significant and challenging task in remote sensing. Due to the arbitrary-oriented property and large aspect ratio of ships, most of the existing detectors adopt rotation boxes to represent ships. However, manual-designed rotation anchors are needed in these detectors, which causes multiplied computational cost and inaccurate box regression. To address the abovementioned problems, an anchor-free rotation ship detector, named GRS-Det, is proposed, which mainly consists of a feature extraction network with selective concatenation module (SCM), a rotation Gaussian-Mask model, and a fully convolutional network-based detection module. First, a U-shape network with SCM is used to extract multiscale feature maps. With the help of SCM, the channel unbalance problem between different-level features in feature fusion is solved. Then, a rotation Gaussian-Mask is designed to model the ship based on its geometry characteristics, which aims at solving the mislabeling problem of rotation bounding boxes. Meanwhile, the Gaussian-Mask leverages context information to strengthen the perception of ships. Finally, multiscale feature maps are fed to the detection module for classification and regression of each pixel. Our proposed method, evaluated on ship detection benchmarks, including HRSC2016 and DOTA Ship data sets, achieves state-of-the-art results.
Xiangrong Zhang, Guanchun Wang, Peng Zhu 0004, Tianyang Zhang 0002, Chen Li 0011, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.5
2020 Dual sentence representation model integrating prior knowledge for bio-text-mining
abstract
Data mining, especially the extraction of the relationship between genes and proteins, plays an important role in the biomedical field. Several related models have been proposed for data mining in the biomedical domain. Furthermore, manually curated biomedical knowledge bases, which could assist the task, have been used to enhance the data-mining model. However, due to the limitation of methods, much prior knowledge information is not be fully exploited. In this work, we propose a novel method that reasonably applied the curated prior knowledge for biomedical text mining by dual sentence representation models; one model is for the experimental data and the other one is for the prior knowledge information sentence. We evaluated our method on two community-supported datasets; BioNLP and BioCreative corpora. The experimental results demonstrate that the dual sentence representation model can successfully utilize external prior knowledge information to extract relationship from biomedical text. Our method can achieve state-of-art results and it could be an application of biomedical relation extraction in the future.
Zhijing Li 0005, YangYang Lan, Saikat Chatterjee, Pargorn Puttapirat, Xiangrong Zhang, Chen Li 0011
BIBM6
2020 Structured Information Extraction of Pathology Reports with Attention-based Graph Convolutional Network
abstract
Electronic medical data contains biochemical, imaging, pathological information during diagnosis and treatment. The pathology report is a kind of highly liberalized unstructured textual data, which is the basis and gold standard of cancer diagnosis and is very important for the prognosis and treatment of patients. The application of information extraction technology to pathological reports can obtain structured data that can be understood and analyzed by computers, helping pathologists make appropriate decisions. In this work, we proposed an attention-based graph convolutional network (GCN) for converting unstructured pathological reports into a structured form suitable for computer analysis to improve the current pathologist's workflow, collected medical data from different platforms, and provided more accurate assistance for diagnosis and treatment. We used pathology reports data from TCGA (The Cancer Genome Atlas) database with fine-grained annotations on 3632 pathology reports including four types of cancers. Our method performs better in our pathology report dataset with higher F1 score than traditional methods and deep learning methods. The results indicate that our method is robust, thus may work with other types of cancer pathology report.
Jialun Wu, Kaiwen Tang, Haichuan Zhang 0001, Chunbao Wang 0002, Chen Li 0011
BIBM5
2020 Scene Attention Mechanism for Remote Sensing Image Caption Generation
abstract
Remote sensing images play an important role in various applications. To make it easier for humans to understand remote sensing images, the task of remote sensing image captioning attracts more and more researchers' attention. Inspired from the way human receives visual information, attention mechanism has been widely used in remote sensing image understanding. To catch more scene information and improve the stability of the generated sentences, a new attention mechanism called scene attention is proposed. Except for the current attention via the current hidden state of the long shortterm memory network (LSTM), our proposed method simultaneously explores the global visual information from the mean feature of all convolutional features. The effectiveness of the proposed method is evaluated on UCM-captions, Sydney-captions and RSICD datasets. The results of our experiment show that comparing with some other captioning methods, our method is more stable and obtains a better performance.
Shiqi Wu, Xiangrong Zhang, Xin Wang 0068, Chen Li 0011, Licheng Jiao
IJCNN4
2020 Discriminative Feature Pyramid Network For Object Detection In Remote Sensing Images
abstract
Multi-class geospatial object detection in remote sensing images suffer great challenges, such as large scales variability and complex background. Although feature pyramid network (FPN) can alleviate the problem of scale variation to some extent, it causes the loss of spatial and semantic information which is not conducive to object location. To address the above problem, this paper proposes a discriminative feature pyramid network (DFPN) by introducing a global guidance module (GGM) and a feature aggregation module (FAM). Specifically, the global guidance module delivers the high-level semantic information to lower layers, so as to obtain feature maps with stronger semantic information to eliminate the interference caused by complex background. The feature aggregation module enhances the interflow of information between different layers and better captures the discrimination information at each layer. We validate the effectiveness of our method on the NWPU VHR-10 and RSOD datasets, the results outperform baseline by 2.06 and 3.88 points respectively.
Xiaoqian Zhu, Xiangrong Zhang, Tianyang Zhang 0002, Peng Zhu 0004, Xu Tang 0004, Chen Li 0011
IJCNN6
2020 Renal Cell Carcinoma Detection and Subtyping with Minimal Point-Based Annotation in Whole-Slide Images
Zeyu Gao 0001, Pargorn Puttapirat, Jiangbo Shi, Chen Li 0011
MICCAI (5)4
2020 BioNorm: deep learning-based event normalization for the curation of reaction databases
abstract
MOTIVATION: A biochemical reaction, bio-event, depicts the relationships between participating entities. Current text mining research has been focusing on identifying bio-events from scientific literature. However, rare efforts have been dedicated to normalize bio-events extracted from scientific literature with the entries in the curated reaction databases, which could disambiguate the events and further support interconnecting events into biologically meaningful and complete networks. RESULTS: In this paper, we propose BioNorm, a novel method of normalizing bio-events extracted from scientific literature to entries in the bio-molecular reaction database, e.g. IntAct. BioNorm considers event normalization as a paraphrase identification problem. It represents an entry as a natural language statement by combining multiple types of information contained in it. Then, it predicts the semantic similarity between the natural language statement and the statements mentioning events in scientific literature using a long short-term memory recurrent neural network (LSTM). An event will be normalized to the entry if the two statements are paraphrase. To the best of our knowledge, this is the first attempt of event normalization in the biomedical text mining. The experiments have been conducted using the molecular interaction data from IntAct. The results demonstrate that the method could achieve F-score of 0.87 in normalizing event-containing statements. AVAILABILITY AND IMPLEMENTATION: The source code is available at the gitlab repository https://gitlab.com/BioAI/leen and BioASQvec Plus is available on figshare https://figshare.com/s/45896c31d10c3f6d857a.
Peiliang Lou, Antonio Jimeno-Yepes, Zai Zhang 0002, Xiangrong Zhang, Chen Li 0011
Bioinform.6
2020 Bio-semantic relation extraction with attention-based external knowledge reinforcement
abstract
BACKGROUND: Semantic resources such as knowledge bases contains high-quality-structured knowledge and therefore require significant effort from domain experts. Using the resources to reinforce the information retrieval from the unstructured text may further exploit the potentials of such unstructured text resources and their curated knowledge. RESULTS: The paper proposes a novel method that uses a deep neural network model adopting the prior knowledge to improve performance in the automated extraction of biological semantic relations from the scientific literature. The model is based on a recurrent neural network combining the attention mechanism with the semantic resources, i.e., UniProt and BioModels. Our method is evaluated on the BioNLP and BioCreative corpus, a set of manually annotated biological text. The experiments demonstrate that the method outperforms the current state-of-the-art models, and the structured semantic information could improve the result of bio-text-mining. CONCLUSION: The experiment results show that our approach can effectively make use of the external prior knowledge information and improve the performance in the protein-protein interaction extraction task. The method should be able to be generalized for other types of data, although it is validated on biomedical texts.
Zhijing Li 0005, Yuchen Lian, Xiaoyong Ma, Xiangrong Zhang, Chen Li 0011
BMC Bioinform.5
2020 Fully Convolutional Network-Based Ensemble Method for Road Extraction From Aerial Images
abstract
This letter proposed a road extraction method based on fully convolutional networks (FCNs) with an ensemble strategy in order to solve the imbalance of road and background areas in aerial images. By utilizing the FCN, we consider road extraction as a semantic segmentation problem. In the network, the weight of the loss function is modified because of the imbalance between the roads and backgrounds, and there will be a larger punishment if roads are wrongly classified as background. Since it is difficult to determine an appropriate weight of the loss function for a given image, an ensemble method based on spatial consistency (SC) is proposed. The result maps that are obtained from the FCNs with different loss functions are fused in our proposed ensemble strategy, which also avoids the determination of weights. Our method is tested using the Massachusetts road data set, and it was proven to be effective compared with the base fully convolutional model according to our experimental result.
Xiangrong Zhang, Wenkang Ma, Chen Li 0011, Jie Wu 0016, Xu Tang 0004, Licheng Jiao
IEEE Geosci. Remote. Sens. Lett.3
2019 OpenHI2 - Open source histopathological image platform
abstract
Transition from conventional to digital pathology requires a new category of biomedical informatic infrastructure which could facilitate delicate pathological routine. Pathological diagnoses are sensitive to many external factors and is known to be subjective. Only systems that can meet strict requirements in pathology would be able to run along pathological routines and eventually digitized the area, and the developed platform should comply with existing pathological routines and international standards. Currently, there are a number of available software tools which can perform histopathological tasks including virtual slide viewing, annotating, and basic image analysis, however, none of them can serve as a digital platform for pathology. Here we describe OpenHI2, an enhanced version Open Histopathological Image platform which is capable of supporting all basic pathological tasks and file formats; ready to be deployed in medical institutions on a standard server environment or cloud computing infrastructure. In this paper, we also describe the development decisions for the platform and propose solutions to overcome technical challenges including responsive region retrieval and viewing, virtual slide magnification, recording of diagnostic areas. These factors would promote OpenHI2 be used as a platform for histopathological images in real-world clinical settings. Furthermore, in research, OpenHI2 inherited the annotation functionality from the previous version, thus acquired annotations can be directly utilized by the newly added machine learning module which include popular machine learning models to perform tasks such as histology image classification and segmentation in the same environment. Addition can be made to the platform since each component is modularized and fully documented. OpenHI2 is free, open-source, and available at https://gitlab.com/BioAI/OpenHI.
Pargorn Puttapirat, Chen Li 0011, Haichuan Zhang 0001, Jingyi Deng, Yuxin Dong 0003, Jiangbo Shi, Zeyu Gao 0001, Chunbao Wang 0002, Xiangrong Zhang
BIBM2
2019 Effects of annotation granularity in deep learning models for histopathological images
abstract
Pathological is crucial to cancer diagnosis. Usually, Pathologists draw their conclusion based on observed cell and tissue structure on histology slides. Rapid development in machine learning, especially deep learning have established robust and accurate classifiers. They are being used to analyze histopathological slides and assist pathologists in diagnosis. Most machine learning systems rely heavily on annotated data sets to gain experiences and knowledge to correctly and accurately perform various tasks such as classification and segmentation. Generally, annotations made in pathology-related datasets have inherited annotation methods from natural scene images. This work investigates different granularity of annotations in histopathological data set including image-wise, bounding box, ellipse-wise, and pixel-wise to verify the influence of annotation in pathological slide on deep learning models. We design corresponding experiments to test classification and segmentation performance of deep learning models based on annotations with different annotation granularity. In classification, state-of-the-art deep learning-based classifiers perform better when trained by pixel-wise annotation dataset. On average, precision, recall and F1-score improves by 7.87%, 8.83% and 7.85% respectively. Thus, it is suggested that finer granularity annotations are better utilized by deep learning algorithms in classification tasks. Similarly, semantic segmentation algorithms can achieve 8.33% better segmentation accuracy when trained by pixel-wise annotations. Our study shows not only that finer-grained annotation can improve the performance of deep learning models, but also help they extract more accurate phenotypic information from histopathological slides. The accurate and spatially precise acquisitions of phenotypic information can improve the reliability of the model prediction. Intelligence systems trained on granular annotations may help pathologists inspecting certain regions and features in the slide that were mainly used to calculate the prediction. The compartmentalized prediction approach similar to this work may contribute to phenotype and genotype association studies.
Jiangbo Shi, Zeyu Gao 0001, Haichuan Zhang 0001, Pargorn Puttapirat, Chunbao Wang 0002, Xiangrong Zhang, Chen Li 0011
BIBM7
2019 Comparing digital histology slides with multiple staining based on decoloring and dyeing technique
abstract
Information in histology slides are usually visualized by different staining techniques, each of them unveils specific chemical and biological substances within tissue samples. Correlations between different stains can be useful to predict how certain tissue slides may look like if they were stained by other staining techniques. This work investigates two stains including hematoxylin and eosin (H&E) and immunohistochemistry (IHC) in digital pathological slides. Four cases of surgical biopsies were used in this work. The specimens were subjected to two consecutive stains with a decoloring process based on ethanol and potassium permanganate in between. After each stain, slides were digitized and archived as results. Comparing the effects of the two staining pipelines, IHC slides after decoloring of H&E showed that the cell structure was clear, the positive IHC staining was accurate, the background of the slide was clean, there was no DAB residue, and tissue fragments were intact. However, the other pipeline where IHC was stained before H&E showed that the nuclear border was blurred. Eosin is lightly colored resulting in low contrast visualization of nucleoplasm, DAB is not completely decolored, and parts of tissue were fragmented. We conclude that, from the proposed staining and decoloring technique, tissue slides could be stained with IHC more effectively on decolored H&E slides than those stained with H&E after IHC. Utilizing digital section scanning technology, we can obtain pairs of tissue images stained differently while preserving the exact same tissue structure.
Chunbao Wang 0002, Pargorn Puttapirat, Chen Li 0011
BIBM5
2018 OpenHI - An open source framework for annotating histopathological image
Pargorn Puttapirat, Haichuan Zhang 0001, Yuchen Lian, Chunbao Wang 0002, Xiangrong Zhang, Lixia Yao, Chen Li 0011
BIBM7
2018 Spatial-Spectral Graph-Based Nonlinear Embedding Dimensionality Reduction for Hyperspectral Image Classificaiton
abstract
Dimensionality reduction (DR) is one of the most important tasks to improve the performance of hyperspectral images classification. Recently, a sparse and low-rank graph embedding based method (SLGE) has been proposed to describe the intrinsic structure of data combined with the local and global constraint simultaneously, which is effective to reduce the dimension of hyperspectral data and obtain a better classification accuracy. However, SLGE is based on an assumption that low-dimensional feature can be obtained utilizing a linear projection. Its performance may degrade under nonlinearly distributed data. Moreover, spatial prior of HSI is not considered in the framework. In this paper, we proposed a novel dimensionality reduction method named spatial-spectral graph-based non-linear embedding (SSGNE). To generate a new graph-trained data, the segmentation strategy based on superpixel is adopted. The spatial-spectral graph is constructed by constraining the sparsity and low-rankness simultaneously on graph-trained data set. Finally, the kernel trick is adopted to extend the general graph embedding framework to nonlinearly space, which fully considers the complexity of real data. Experimental results show that the proposed method outperforms the state-of-the-art methods in terms of the classification accuracy.
Xiangrong Zhang, Yaru Han, Ning Huyan, Chen Li 0011, Jie Feng 0003, Xiaoxiao Ma 0003
IGARSS4
2018 Hybrid Unmixing Based on Adaptive Region Segmentation for Hyperspectral Imagery
abstract
Unmixing is an important issue of hyperspectral images. Most unmixing methods adopt linear mixing models for simplicity. However, multiple scattering usually occurs between vegetation and soil in a bilinear scene. Thus, nonlinear mixing problems which are difficult to be solved should be taken into consideration under this circumstance. In practice, both linear and nonlinear spectral mixtures exist in hyperspectral scenes. Considering the characteristics of different regions in images, we propose a hybrid unmixing algorithm for hyperspectral images based on region adaptive segmentation. Our method uses a standard K-means clustering algorithm to obtain different regions, including homogeneous regions and detailed regions. The model of the homogeneous regions is assumed to be linear, which will be pursued using the method of sparse-constrained nonnegative matrix factorization (NMF), and the mixing in the detailed regions is assumed to be based on a nonlinear model. We also propose a new nonlinear unmixing method, called graph-regularized semi-NMF, which considers the manifold structure of hyperspectral data as the unmixing method to deal with the detailed regions. Finally, by combining the two regions, we obtain the abundance of the whole hyperspectral image. The proposed method can not only achieve more precise abundance but also be good at keeping the edge information of the bilinear abundance. The experimental results on both synthetic and real data also show that the proposed method is effective for improving the unmixing accuracy of hyperspectral remote-sensing images.
Xiangrong Zhang, Chen Li 0011, Cai Cheng, Licheng Jiao, Huiyu Zhou 0001
IEEE Trans. Geosci. Remote. Sens.3
2017 Improving Chinese Sentiment Analysis via Segmentation-Based Representation Using Parallel CNN
Yazhou Hao, YangYang Lan, Yufei Li 0002, Meng Wang 0009, Sen Wang 0001, Chen Li 0011
ADMA7
2017 Natural language description of remote sensing images based on deep learning
abstract
The semantic description of remote sensing image is a useful and meaningful task, which can help us to get a better understanding of the scene depicted in the remote sensing images and make better use of the remote sensing images. Nature language provides good solution for describing the semantic information of remote sensing images. Nature language description of a remote sensing image is to generate a meaningful sentence given a remote sensing image. This paper presents a novel method based on deep learning. First, a convolutional neural network is utilized to detect the main objects of the remote sensing images. Then a recurrent neural network language model is utilized to generate the natural language descriptions of the objects which are detected in the first step. Experimental results on a set of remote sensing images demonstrate that the proposed method is able to generate desirable description of the scene.
Xiangrong Zhang, Xiang Li 0013, Jinliang An, Biao Hou, Chen Li 0011
IGARSS6
2017 Recursive Autoencoders-Based Unsupervised Feature Learning for Hyperspectral Image Classification
abstract
For hyperspectral image (HSI) classification, it is very important to learn effective features for the discrimination purpose. Meanwhile, the ability to combine spectral and spatial information together in a deep level is also important for feature learning. In this letter, we propose an unsupervised feature learning method for HSI classification, which is based on recursive autoencoders (RAE) network. RAE utilizes the spatial and spectral information and produces high-level features from the original data. It learns features from the neighborhood of the investigated pixel to represent the whole local homogeneous area of the image. In addition, to obtain more accurate representation of the investigated pixel, a weighting scheme is adopted based on the neighboring pixels, where the weights are determined by the spectral similarity between the neighboring pixels and the investigated pixel. The effectiveness of our method is evaluated by the experiments on two hyperspectral data sets, and the results show that our proposed method has a better performance.
Xiangrong Zhang, Chen Li 0011, Ning Huyan, Licheng Jiao, Huiyu Zhou 0001
IEEE Geosci. Remote. Sens. Lett.3
2016 Automated Segmentation of MOOC Lectures towards Customized Learning
abstract
The sheer size of the student body for MOOC and the diversity of their learning styles and backgrounds demand that we develop alternatives to the one-size-fits-all pedagogy used in residential education. An important aspect of this endeavor is the segmentation of the video material, since it forms the omnipresent and central part of every course, and structuralized videos allow non-linear navigation as well as help learners with various needs find desired information efficiently. Here, we propose an automatic visual transition detection method to partition lecture videos into self-contained segments, which is the foundation to structuralize video and support non-linear navigation. Our method can be done at scale and has been proved being able to achieve reasonable quality.
Xiangrong Zhang, Chen Li 0011, Shang-Wen Li 0001, Victor Zue
ICALT2
2014 Biological network extraction from scientific literature: state of the art and challenges
abstract
Networks of molecular interactions explain complex biological processes, and all known information on molecular events is contained in a number of public repositories including the scientific literature. Metabolic and signalling pathways are often viewed separately, even though both types are composed of interactions involving proteins and other chemical entities. It is necessary to be able to combine data from all available resources to judge the functionality, complexity and completeness of any given network overall, but especially the full integration of relevant information from the scientific literature is still an ongoing and complex task. Currently, the text-mining research community is steadily moving towards processing the full body of the scientific literature by making use of rich linguistic features such as full text parsing, to extract biological interactions. The next step will be to combine these with information from scientific databases to support hypothesis generation for the discovery of new knowledge and the extension of biological networks. The generation of comprehensive networks requires technologies such as entity grounding, coordination resolution and co-reference resolution, which are not fully solved and are required to further improve the quality of results. Here, we analyse the state of the art for the extraction of network information from the scientific literature and the evaluation of extraction methods against reference corpora, discuss challenges involved and identify directions for future research.
Chen Li 0011, Maria Liakata, Dietrich Rebholz-Schuhmann
Briefings Bioinform.1
2010 BioModels.net Web Services, a free and integrated toolkit for computational modelling software
abstract
Exchanging and sharing scientific results are essential for researchers in the field of computational modelling. BioModels.net defines agreed-upon standards for model curation. A fundamental one, MIRIAM (Minimum Information Requested in the Annotation of Models), standardises the annotation and curation process of quantitative models in biology. To support this standard, MIRIAM Resources maintains a set of standard data types for annotating models, and provides services for manipulating these annotations. Furthermore, BioModels.net creates controlled vocabularies, such as SBO (Systems Biology Ontology) which strictly indexes, defines and links terms used in Systems Biology. Finally, BioModels Database provides a free, centralised, publicly accessible database for storing, searching and retrieving curated and annotated computational models. Each resource provides a web interface to submit, search, retrieve and display its data. In addition, the BioModels.net team provides a set of Web Services which allows the community to programmatically access the resources. A user is then able to perform remote queries, such as retrieving a model and resolving all its MIRIAM Annotations, as well as getting the details about the associated SBO terms. These web services use established standards. Communications rely on SOAP (Simple Object Access Protocol) messages and the available queries are described in a WSDL (Web Services Description Language) file. Several libraries are provided in order to simplify the development of client software. BioModels.net Web Services make one step further for the researchers to simulate and understand the entirety of a biological system, by allowing them to retrieve biological models in their own tool, combine queries in workflows and efficiently analyse models.
Chen Li 0011, Mélanie Courtot, Nicolas Le Novère, Camille Laibe
Briefings Bioinform.1