VLDB 2026 Research / reviewers in the wild / expert
Lan Huang 0002
dblp:74/1260-2
· DBLP profile ↗
81ranked-venue papers
23as first author
63since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 38 · 10 first-author · 31 since 2021Artificial intelligence and machine learning · 25 · 6 first-author · 19 since 2021Databases, data management, data science and information retrieval · 16 · 4 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 6 since 2021Systems, architecture and hardware · 4 · 1 first-author · 2 since 2021Computer networks · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multiple feature similarities based heterogeneous graph representation
Lan Huang 0002, Yihang Geng, Rui Zhang 0084 |
Knowl. Based Syst. | 1 |
| 2026 | HpMiX: A Disease ceRNA biomarker prediction framework driven by graph topology-constrained Mixup and hypergraph residual enhancement
Xinfei Wang 0001, Lan Huang 0002, Yan Wang 0028, Renchu Guan, Zhu-Hong You, Fengfeng Zhou |
Neural Networks | 2 |
| 2026 | A pre-trained language model-based cross-modal fusion framework for predicting miRNA-drug resistance and sensitivity associationsabstractMicroRNAs (miRNAs) are pivotal regulators of drug resistance and sensitivity in cancer cells, functioning as tumor suppressors or oncogenes that modulate the cellular response to anticancer drugs. While experimental identification of miRNA-mediated drug resistance and sensitivity is both costly and laborious, computational methods present a promising alternative. Recent advances in pre-trained language models (PLMs) offer new opportunities to leverage large-scale unlabeled biomolecular data for enhanced relationship prediction. In this study, we introduce PLMF-MDA, a PLM-based cross-modal fusion model designed to predict miRNA-drug resistance (MDR) and miRNA-drug sensitivity (MDS) associations. PLMF-MDA integrates miRNA and drug multimodal embeddings derived from PLMs and intrinsic feature extractors, and employs a cross-modal attention fusion module to adaptively capture key interactions between modalities. To evaluate the performance of the approach, we manually constructed two benchmark datasets. Experimental results demonstrate that the PLMF-MDA achieves superior prediction performance. Furthermore, case studies on anticancer drug docetaxel and gefitinib demonstrate its potential in discovering novel MDR (MDS) associations. All data and source code are available on GitHub: https://github.com/sheng-n/PLMF-MDA. Nan Sheng, Yun-Zhi Liu, Wenju Hou, Lan Huang 0002, Yan Wang 0028 |
PLoS Comput. Biol. | 5 |
| 2026 | DeepSelective: Interpretable prognosis prediction via feature selection and compression in EHR data
Ruochi Zhang, Xiaoyang Wang 0009, Qiong Zhou, Ziqi Deng, Yueying Wang, Yusi Fan, Jiale Zhang 0002, Lan Huang 0002, Chang Liu 0082, Fengfeng Zhou |
Pattern Recognit. | 11 |
| 2026 | Structure and Semantics Aware Multi-View Contrastive Learning for Predicting Association Among lncRNAs, miRNAs and DiseasesabstractExploring associations among long non-coding RNAs (lncRNAs), microRNAs (miRNAs), and diseases is crucial for biomarker discovery and precision medicine. Existing computational methods are hindered by sparse known associations and the complexity of biological networks. To address this challenge, we propose SSMVCL (Structure- and Semantic-aware Multi-View Contrastive Learning), a unified framework for predicting lncRNA-disease associations (LDAs), miRNA-disease associations (MDAs), and lncRNA-miRNA interactions (LMIs). SSMVCL constructs a heterogeneous bioinformatics network from multi-source biological data and learns representations from two complementary views: a structure-aware view for local topology and a semantic-aware view using biologically meaningful meta-paths to capture high-order relationships. A cross-view contrastive alignment module with adaptive negative sampling enforces consistency between views and enhances discriminative capability. On two benchmark datasets, SSMVCL achieves state-of-the-art performance: for Dataset2, AUC/AUPR of 0.9736/0.9716 (LDA), 0.9364/0.9309 (MDA), and 0.9297/0.9234 (LMI) Case studies on gastric and prostate cancers further validated robustness and translational potential by identifying supported associations. Lan Huang 0002, Yujuan Zhang, Yan Wang 0028, Nan Sheng |
IEEE J. Biomed. Health Informatics | 1 |
| 2026 | A Dynamic Multi-Scale Hypergraph Learning Framework Driven by Features and Structures for ceRNA-Disease Association PredictionabstractCompetitive endogenous RNA (ceRNA) networks are pivotal for uncovering disease molecular mechanisms. Graph representation learning is a cornerstone for modeling biological regulatory networks and predicting disease-related biomarkers. However, current methods face challenges: traditional graph neural network (GNN) rely on low-order graph structures, which struggle to capture high-order molecular interactions, resulting in topological information loss; shallow GNN fail to model long-range dependencies, while deep architectures suffer from over-smoothing, limiting complex regulatory expression; static embeddings overlook dynamic molecular interactions, reducing biomarker accuracy. These limitations highlight the need for advanced graph learning frameworks. To address these challenges, we propose DMHLF, a Dynamic Multi-scale Hypergraph Learning Framework for predicting disease-associated ceRNA biomarkers. The framework first integrates multiple regulatory relationships among miRNAs, lncRNAs, circRNAs, mRNAs, and diseases to construct disease-specific ceRNA regulatory networks, capturing local and global regulatory patterns through multi-Hop hyperedges. Subsequently, we devise a Hypergraph-Weighted Dynamic Random Walk (HEDRW) method to dynamically extract node meta-embeddings that encode high-order regulatory information. Concurrently, we extend Eigen-GNN spectral analysis to hypergraph structures, incorporating a residual-enhanced hypergraph neural network to preserve the global topological properties of shallow hypergraphs. Finally, a cross-scale attention mechanism aligns and fuses multi-scale features to generate high-quality node embeddings for disease-ceRNA association prediction. Experiments on diverse datasets demonstrate that DMHLF significantly outperforms existing methods. Case study further validates the framework's efficacy in identifying disease-related ceRNA biomarkers, providing a reliable predictive tool for biomedical research. Xinfei Wang 0001, Lan Huang 0002, Yan Wang 0028, Renchu Guan, Zhu-Hong You, Fengfeng Zhou, Yu-Qing Li |
IEEE J. Biomed. Health Informatics | 2 |
| 2026 | Reliable Multimodal Cancer Survival Prediction With Confidence-Aware Risk ModelingabstractMultimodal survival methods that integrate histology whole-slide images and transcriptomic profiles hold significant promise for understanding patient prognostication and guiding personalized treatment strategies. However, existing approaches primarily focus on improving predictive performance through multimodal information fusion, often neglecting the reliability estimation of the prediction results and the inherent alignment noise across modalities. Thus, we propose ReCaSP, a novel and reliable cancer survival prediction framework that effectively integrates histology and transcriptomics data via multimodal alignment and fusion, providing the auxiliary confidence levels for survival predictions through a confidence-aware risk modeling mechanism. Specifically, our approach incorporates a fine-grained risk classifier that models risk labels jointly over both multiple time intervals and censorship status, utilizing evidential deep learning to yield fine-grained risk predictions accompanied by confidence scores. Additionally, to mitigate the inherent noise in multimodal data alignment, we introduce a cross-attention alignment module that effectively aligns histology data with transcriptomics data prior to multimodal fusion, thereby facilitating cross-modal interaction learning. Extensive experiments on five datasets demonstrate that ReCaSP significantly outperforms state-of-the-art methods, achieving a 4.58% improvement in the overall C-Index. Xuping Xie, Qixing Yang, Lan Huang 0002, Fengfeng Zhou, Yan Wang 0028 |
IEEE J. Biomed. Health Informatics | 5 |
| 2025 | A Simple Graph Contrastive Learning Framework for Short Text ClassificationabstractShort text classification has gained significant attention in the information age due to its prevalence and real-world applications. Recent advancements in graph learning combined with contrastive learning have shown promising results in addressing the challenges of semantic sparsity and limited labeled data in short text classification. However, existing models have certain limitations. They rely on explicit data augmentation techniques to generate contrastive views, resulting in semantic corruption and noise. Additionally, these models only focus on learning the intrinsic consistency between the generated views, neglecting valuable discriminative information from other potential views. To address these issues, we propose a Simple graph contrastive learning framework for Short Text Classification (SimSTC). Our approach involves performing graph learning on multiple text-related component graphs to obtain multi-view text embeddings. Subsequently, we directly apply contrastive learning on these embeddings. Notably, our method eliminates the need for data augmentation operations to generate contrastive views while still leveraging the benefits of multi-view contrastive learning. Despite its simplicity, our model achieves outstanding performance, surpassing large language models on various datasets. Yonghao Liu 0001, Fausto Giunchiglia, Lan Huang 0002, Ximing Li 0002, Xiaoyue Feng, Renchu Guan |
AAAI | 3 |
| 2025 | Boosting Short Text Classification with Multi-Source Information Exploration and Dual-Level Contrastive LearningabstractShort text classification, as a research subtopic in natural language processing, is more challenging due to its semantic sparsity and insufficient labeled samples in practical scenarios. We propose a novel model named MI-DELIGHT for short text classification in this work. Specifically, it first performs multi-source information (i.e., statistical information, linguistic information, and factual information) exploration to alleviate the sparsity issues. Then, the graph learning approach is adopted to learn the representation of short texts, which are presented in graph forms. Moreover, we introduce a dual-level (i.e., instance-level and cluster-level) contrastive learning auxiliary task to effectively capture different-grained contrastive information within massive unlabeled data. Meanwhile, previous models merely perform the main task and auxiliary tasks in parallel, without considering the relationship among tasks. Therefore, we introduce a hierarchical architecture to explicitly model the correlations between tasks. We conduct extensive experiments across various benchmark datasets, demonstrating that MI-DELIGHT significantly surpasses previous competitive models. It even outperforms popular large language models on several datasets. Yonghao Liu 0001, Wei Pang 0001, Fausto Giunchiglia, Lan Huang 0002, Xiaoyue Feng, Renchu Guan |
AAAI | 5 |
| 2025 | GSAM-MRI: Frequency-Based Domain Randomization for Generalized MR Image Segmentation with Segment Anything ModelabstractMagnetic resonance imaging (MRI) data segmentation plays a critical role in clinical diagnosis and treatment planning. However, the performance of deep learning-based segmentation models is often hindered by domain shifts caused by variations in imaging factors. To address this challenge, we propose GSAM-MRI, a Generalized Segment Anything Model for robust MRI segmentation in the scenario of single-source domain generalization (SDG). GSAM-MRI integrates multiple components to enhance domain generalization: (1) Frequency-based Domain Randomization module that simulates inter-site variability by perturbing the frequency domain; (2) Domain Adversarial Block that promotes domain-invariant feature learning through adversarial training; (3) General Embedding Generator that fuses multi-scale hierarchical features to produce dense prompt embeddings; Additionally, a hybrid loss function is employed for output consistency. Experiments on prostate segmentation and white matter hyperintensity segmentation tasks demonstrate that GSAM-MRI consistently outperforms state-of-the-art SDG methods and baseline models, achieving superior generalization across unseen domains. Lan Huang 0002, Yinglu Sun, Xinfei Wang 0001, Qixing Yang, Xuping Xie, Wenju Hou, Chunjie Guo, Yan Wang 0028 |
BIBM | 1 |
| 2025 | Enhancing Unsupervised Graph Few-shot Learning via Set Functions and Optimal TransportabstractGraph few-shot learning has garnered significant attention for its ability to rapidly adapt to downstream tasks with limited labeled data, sparking considerable interest among researchers. Recent advancements in graph few-shot learning models have exhibited superior performance across diverse applications. Despite their successes, several limitations still exist. First, existing models in the meta-training phase predominantly focus on instance-level features within tasks, neglecting crucial set-level features essential for distinguishing between different categories. Second, these models often utilize query sets directly on classifiers trained with support sets containing only a few labeled examples, overlooking potential distribution shifts between these sets and leading to suboptimal performance. Finally, previous models typically require necessitate abundant labeled data from base classes to extract transferable knowledge, which is typically infeasible in real-world scenarios. To address these issues, we propose a novel model named STAR, which leverages Set funcTions and optimAl tRansport for enhancing unsupervised graph few-shot learning. Specifically, STAR utilizes expressive set functions to obtain set-level features in an unsupervised manner and employs optimal transport principles to align the distributions of support and query sets, thereby mitigating distribution shift effects. Theoretical analysis demonstrates that STAR can capture more task-relevant information and enhance generalization capabilities. Empirically, extensive experiments across multiple datasets validate the effectiveness of STAR. Our code can be found here. Yonghao Liu 0001, Fausto Giunchiglia, Ximing Li 0002, Lan Huang 0002, Xiaoyue Feng, Renchu Guan |
KDD (1) | 4 |
| 2025 | Dual-level Mixup for Graph Few-shot Learning with Fewer TasksabstractGraph neural networks have been demonstrated as a powerful paradigm for effectively learning graph-structured data on the web and mining content from it. %the wide web. for downstream task analysis. Current leading graph models require a large number of labeled samples for training, which unavoidably leads to overfitting in few-shot scenarios. Recent research has sought to alleviate this issue by simultaneously leveraging graph learning and meta-learning paradigms. However, these graph meta-learning models assume the availability of numerous meta-training tasks to learn transferable meta-knowledge. Such assumption may not be feasible in the real world due to the difficulty of constructing tasks and the substantial costs involved. Therefore, we propose a SiMple yet effectIve approach for graph few-shot Learning with fEwer tasks, named SMILE. We introduce a dual-level mixup strategy, encompassing both within-task and across-task mixup, to simultaneously enrich the available nodes and tasks in meta-learning. Moreover, we explicitly leverage the prior information provided by the node degrees in the graph to encode expressive node representations. Theoretically, we demonstrate that SMILE can enhance the model generalization ability. Empirically, SMILE consistently outperforms other competitive models by a large margin across all evaluated datasets with in-domain and cross-domain settings. Our anonymous code can be found https://github.com/KEAML-JLU/SMILE. Yonghao Liu 0001, Fausto Giunchiglia, Lan Huang 0002, Ximing Li 0002, Xiaoyue Feng, Renchu Guan |
WWW | 4 |
| 2025 | Efficient feature selection for pre-trained vision transformers
Lan Huang 0002, Mengqiang Yu, Weiping Ding 0001, Xingyu Bai, Kangping Wang |
Comput. Vis. Image Underst. | 1 |
| 2025 | TaiChiNet: PCA-based Ying-Yang dilution of inter- and intra-BERT layers to represent anti-coronavirus peptides
Shiying Ding, Yusi Fan, Yannan Sun, Gongyou Zhang, Ruochi Zhang, Lan Huang 0002, Fengfeng Zhou |
Expert Syst. Appl. | 9 |
| 2025 | Supervised contrastive knowledge graph learning for ncRNA-disease association prediction
Yan Wang 0028, Xuping Xie, Nan Sheng, Lan Huang 0002, Chunman Zuo |
Expert Syst. Appl. | 5 |
| 2025 | Prompt-guided orthogonal multimodal fusion for cancer survival prediction
Lan Huang 0002, Shuyu Guo, Tian Bai 0002, Ruihong Zhao, Ke Tao |
Inf. Sci. | 1 |
| 2025 | PCG-CAM: Enhanced class activation map using principal components of gradients and its applications in brain MRI
Lan Huang 0002, Yangguang Shao, Wenju Hou, Yan Wang 0028, Nan Sheng, Yinglu Sun, Yao Wang 0010 |
Inf. Sci. | 1 |
| 2025 | Unsupervised Adversarial Domain Adaptation with Hierarchical Semantic Consistency for Cross-Modal Nuclei Detection
Shuyu Guo, Lan Huang 0002, Yu-Hao Mu, Tian Bai 0002 |
J. Comput. Sci. Technol. | 2 |
| 2025 | Multi-view fusion based on graph convolutional network with attention mechanism for predicting miRNA related to drugsabstractMicroRNAs (miRNAs) play crucial roles in cancer progression, invasion, and response to treatment, particularly in regulating anticancer drug resistance and sensitivity. Identifying potential human miRNA-drug associations (MDAs) that manifest as resistance or sensitivity relationships offers valuable insights for cancer treatment and drug development. With the growing availability of biological data, computational methods have emerged as powerful tools to complement experimental approaches. However, limited attention has been paid to computational prediction of MDAs. Furthermore, existing approaches typically rely on known MDA information, overlooking the valuable insights available from multi-source data related to miRNAs and drugs. In this study, we present a multi-view fusion-based graph convolutional network with attention mechanism (MGCNA) to predict miRNA-associated drug resistance/sensitivity. Specifically, MGCNA integrates macro- and micro- level information of miRNAs and drugs to construct multi-view node features from different perspectives. The proposed multi-view graph convolutional network (GCN) encoder obtains miRNA and disease features from different views and learns adaptive importance weights of the embedding using an attention mechanism. Extensive experiments on manually curated benchmark datasets demonstrate that MGCNA outperforms existing baseline methods. Case studies of two common drugs further establish MGCNA's effectiveness in discovering novel MDAs. Nan Sheng, Yun-Zhi Liu, Lei Wang 0121, Lan Huang 0002, Yan Wang 0028 |
PLoS Comput. Biol. | 5 |
| 2025 | A framework for global role-based author name disambiguation
Lan Huang 0002, Rui Zhang 0084 |
Pattern Recognit. | 1 |
| 2025 | MCSSA: A Stream-Based Multiconcurrency Systolic Sorting Array Combining Merge TreeabstractThe exploration of utilizing reconfigurable circuits with parallel computing capabilities has been conducted to enhance sorting performance and reduce power consumption. However, most sorting algorithms using dedicated processors are based on parallelization designs of serial algorithms without considering the design method of large-scale integrated circuits. This results in various issues, including the overuse of$I/O$interface resources, on-chip storage resources, and complex layout wiring. In this article, we extend the 2-tuple relation in the uniform recurrence equation (URE) structure used to define the systolic array to n-tuples, and the extended structure is flexible in defining$I/O$bandwidth and concurrency. Then we define the multiconcurrency systolic sorter array (MCSSA) algorithm based on the extended URE structure, which has a flexible$4N/n$time complexity based on the n-tuple relation. Moreover, this systolic array can simultaneously sort two independent sequences, increasing the reuse of resources. Afterwards, we encapsulate each n-tuple into a processing element (PE) cell. The entire MCSSA consists of these interconnected PE cells, each of which can be customized in terms of data bit width and type. Last but not least, we have improved the merge tree structure called MC-merge tree. The concurrency of this algorithm can also be flexibly defined, we use this algorithm combined with MCSSA to cope with large-scale sorting scenarios. In our experiments, we have demonstrated the speed-up ratio of MCSSA relative to other state of the art (SOTA) sorting algorithms. Inheriting the unity and simplicity from the Systolic Array architecture, MCSSA achieves a maximum$73.17\times $acceleration ratio on the U200. In addition, the MC-merge tree expands the MCSSA sorting scale with a maximum of 450.56 times while maintaining the advantage of the acceleration ratio. The results of our study demonstrate that MCSSA and MC-merge tree have better acceleration, throughput and scalability advantages over other SOTA algorithms. Lan Huang 0002, Teng Gao, Kangping Wang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2025 | TSLAmy: A Novel Amyloid Hexapeptide Aggregation Prediction Approach Based on Two-Stage LearningabstractIdentifying aggregation-prone proteins or peptides is essential for advancing our understanding of amyloid aggregation processes and their related pathogenic mechanisms. Recognizing potential amyloid hexapeptides can also support peptide-based drug design and reduce experimental costs. In this study, we proposed TSLAmy, a computational model designed to predict amyloid hexapeptides using a two-stage learning framework. In the first stage, we performed feature extraction on the hexapeptides, and in the second stage, we presented prediction model for amyloid hexapeptide aggregation. Firstly, to ensure balanced dataset partitioning, we applied a clustering-based method by training two autoencoders on all possible hexapeptides using their sequence and physicochemical features, respectively. The resulting clusters were used to stratify the data into training and testing datasets. Then, in the first stage, we extracted features from hexapeptides based on their sequence and physicochemical properties. The feature extraction module was used to obtain physicochemical features, while the ESM-2 module was responsible for extracting sequence features for each hexapeptide. Finally, in the second stage, the aggregation prediction module was employed to predict the aggregation potential of hexapeptides. The experimental results demonstrated that the accuracy of TSLAmy reached 0.8493 (0.8447-0.8539), outperforming other state-of-the-art methods. Furthermore, we predicted the aggregation potential of all 64,000,000 possible hexapeptides and analyzed the amino acids that form aggregation-prone hexapeptides. We anticipate that TSLAmy can offer new insights into the identification of aggregation-prone peptides, contributing to advancements in peptide drug development. Lan Huang 0002, Qingchen Jiang, Yucong Xiong, Guangzhao Zhang, Dan Shao |
IEEE J. Biomed. Health Informatics | 1 |
| 2025 | Self-Supervised Contrastive Learning on Attribute and Topology Graphs for Predicting Relationships Among lncRNAs, miRNAs and DiseasesabstractExploring associations between long non-coding RNAs (lncRNAs), microRNAs (miRNAs) and diseases is crucial for disease prevention, diagnosis and treatment. While determining these relationships experimentally is resource-intensive and time-consuming, computational methods have emerged as an attractive way. However, existing computational methods tend to focus on single tasks, neglecting the benefits of leveraging multiple biomolecular interactions and domain-specific knowledge for multi-task prediction. Furthermore, the scarcity of labeled data for lncRNA-disease associations (LDAs), miRNA-disease associations (MDAs) and lncRNA-miRNA interactions (LMIs) poses challenges for comprehensive node embedding learning. This paper proposes a multi-task prediction model (called SSCLMD) that employs self-supervised contrastive learning on attribute and topology graphs to identify potential LDAs, MDAs and LMIs. Firstly, domain knowledge of lncRNAs, miRNAs and diseases as well as their interactions are exploited to construct attribute graph and topology graph, respectively. Then, the nodes are encoded in the attribute and topology spaces to extract the specific and common feature. Meanwhile, the attention mechanism is performed to adaptively fuse the embedding from different views. SSCLMD incorporates contrastive self-supervised learning as a regularize to guide node embedding learning in both attribute and topology space without relying on labels. Severing as a regularize in multi-task learning paradigm, it to improves the model.s generalization capabilities. Extensive experiments on 2 manually curated datasets demonstrate that SSCLMD significantly outperforms baseline methods in LDA, MDA and LMI prediction tasks. Case studies on both old and new datasets further supported SSCLMD's ability to uncover novel disease-related lncRNAs and miRNAs. Lan Huang 0002, Nan Sheng, Lei Wang 0121, Wenju Hou, Yan Wang 0028 |
IEEE J. Biomed. Health Informatics | 1 |
| 2024 | Improved Graph Contrastive Learning for Short Text ClassificationabstractText classification occupies an important role in natural language processing and has many applications in real life. Short text classification, as one of its subtopics, has attracted increasing interest from researchers since it is more challenging due to its semantic sparsity and insufficient labeled data. Recent studies attempt to combine graph learning and contrastive learning to alleviate the above problems in short text classification. Despite their fruitful success, there are still several inherent limitations. First, the generation of augmented views may disrupt the semantic structure within the text and introduce negative effects due to noise permutation. Second, they ignore the clustering-friendly features in unlabeled data and fail to further utilize the prior information in few valuable labeled data. To this end, we propose a novel model that utilizes improved Graph contrastIve learning for short text classiFicaTion (GIFT). Specifically, we construct a heterogeneous graph containing several component graphs by mining from an internal corpus and introducing an external knowledge graph. Then, we use singular value decomposition to generate augmented views for graph contrastive learning. Moreover, we employ constrained kmeans on labeled texts to learn clustering-friendly features, which facilitate cluster-oriented contrastive learning and assist in obtaining better category boundaries. Extensive experimental results show that GIFT significantly outperforms previous state-of-the-art methods. Our code can be found in https://github.com/KEAML-JLU/GIFT. Yonghao Liu 0001, Lan Huang 0002, Fausto Giunchiglia, Xiaoyue Feng, Renchu Guan |
AAAI | 2 |
| 2024 | 2.5D ASF-UNet: Adjacent Slice Spatial Feature Fusion Model for WMH Segmentation from 3D MR Brain ImageabstractSegmenting brain white matter hyperintensities (WMH) from 3D Magnetic Resonance (MR) images is crucial for the diagnosis, treatment, and prognosis of Multiple Sclerosis (MS). Unlike common 2D images, this task is more challenging and time-consuming. Classical deep learning methods for 3D image segmentation face two main challenges: 1) Pure 3D networks have more parameters and are prone to overfitting. 2) When using 2D networks to segment slices of 3D images, the lack of 3D structural information results in suboptimal segmentation after reconstruction. To address these difficulties, we propose the 2.5D ASF-UNet, which employs the 2.5D workflow and uses adjacent slices as the input for 2D segmentation network ASF-UNet. In ASF-UNet, separate down-sampling paths are used for the adjacent slices, and the Local Spatial Attention Module (LSAM) is designed to more effectively integrate 3D spatial information into the 2D network. Additionally, the Conv_Spectral_Block (CSB) is designed to extract and integrate local and global features. It allows the model to capture global spatial structures while preserving detailed information. Experimental results on MICCAI MSSEG 2016 and Local MS datasets show that 2.5D ASF-UNet achieve better segmentation performance than other deep learning methods. Lan Huang 0002, Yinglu Sun, Chunjie Guo, Yan Wang 0028 |
BIBM | 1 |
| 2024 | Resolving Word Vagueness with Scenario-guided Adapter for Natural Language Inference
Yonghao Liu 0001, Di Liang, Ximing Li 0002, Fausto Giunchiglia, Lan Huang 0002, Xiaoyue Feng, Renchu Guan |
IJCAI | 6 |
| 2024 | A Simple but Effective Approach for Unsupervised Few-Shot Graph ClassificationabstractGraphs, as a fundamental data structure, have proven efficacy in modeling complex relationships between objects and are therefore found in wide web applications. Graph classification is an essential task in graph data analysis, which can effectively assist in extracting information and mining content from the web. Recently, few-shot graph classification, a more realistic and challenging task, has garnered great research interest. Existing few-shot graph classification models are all supervised, assuming abundant labeled data in base classes for meta-training. However, sufficient annotation is often challenging to obtain in practice due to high costs or demand for expertise. Moreover, they commonly adopt complicated meta-learning algorithms via episodic training to transfer prior knowledge from base classes. To break free from these constraints, in this paper, we propose a simple yet effective approach named SMART for unsupervised few-shot graph classification without using any labeled data. SMART employs transfer learning philosophy instead of the previously prevailing meta-learning paradigm, avoiding the need for sophisticated meta-learning algorithms. Additionally, we adopt a novel mixup strategy to augment the original graph data and leverage unsupervised pretraining on these data to obtain the expressive graph encoder. We also utilize the prompt tuning technique to alleviate the overfitting and low fine-tuning efficiency caused by the limited support samples of novel classes. Extensive experimental results demonstrate the superiority of our proposed approach, significantly surpassing even leading supervised few-shot graph classification models. Our code is available here. Yonghao Liu 0001, Lan Huang 0002, Bowen Cao, Ximing Li 0002, Fausto Giunchiglia, Xiaoyue Feng, Renchu Guan |
WWW | 2 |
| 2024 | A multi-task prediction method based on neighborhood structure embedding and signed graph representation learning to infer the relationship between circRNA, miRNA, and cancerabstractMOTIVATION: Research shows that competing endogenous RNA is widely involved in gene regulation in cells, and identifying the association between circular RNA (circRNA), microRNA (miRNA), and cancer can provide new hope for disease diagnosis, treatment, and prognosis. However, affected by reductionism, previous studies regarded the prediction of circRNA-miRNA interaction, circRNA-cancer association, and miRNA-cancer association as separate studies. Currently, few models are capable of simultaneously predicting these three associations. RESULTS: Inspired by holism, we propose a multi-task prediction method based on neighborhood structure embedding and signed graph representation learning, CMCSG, to infer the relationship between circRNA, miRNA, and cancer. Our method aims to extract feature descriptors of all molecules from the circRNA-miRNA-cancer regulatory network using known types of association information to predict unknown types of molecular associations. Specifically, we first constructed the circRNA-miRNA-cancer association network (CMCN), which is constructed based on the experimentally verified biomedical entity regulatory network; next, we combine topological structure embedding methods to extract feature representations in CMCN from local and global perspectives, and use denoising autoencoder for enhancement; then, combined with balance theory and state theory, molecular features are extracted from the point of social relations through the propagation and aggregation of signed graph attention network; finally, the GBDT classifier is used to predict the association of molecules. The results show that CMCSG can effectively predict the relationship between circRNA, miRNA, and cancer. Additionally, the case studies also demonstrate that CMCSG is capable of accurately identifying biomarkers across various types of cancer. The data and source code can be found at https://github.com/1axin/CMCSG. Lan Huang 0002, Xinfei Wang 0001, Yan Wang 0028, Renchu Guan, Nan Sheng, Xuping Xie, Lei Wang 0121 |
Briefings Bioinform. | 1 |
| 2024 | Multi-view learning framework for predicting unknown types of cancer markers via directed graph neural networks fitting regulatory networksabstractThe discovery of diagnostic and therapeutic biomarkers for complex diseases, especially cancer, has always been a central and long-term challenge in molecular association prediction research, offering promising avenues for advancing the understanding of complex diseases. To this end, researchers have developed various network-based prediction techniques targeting specific molecular associations. However, limitations imposed by reductionism and network representation learning have led existing studies to narrowly focus on high prediction efficiency within single association type, thereby glossing over the discovery of unknown types of associations. Additionally, effectively utilizing network structure to fit the interaction properties of regulatory networks and combining specific case biomarker validations remains an unresolved issue in cancer biomarker prediction methods. To overcome these limitations, we propose a multi-view learning framework, CeRVE, based on directed graph neural networks (DGNN) for predicting unknown type cancer biomarkers. CeRVE effectively extracts and integrates subgraph information through multi-view feature learning. Subsequently, CeRVE utilizes DGNN to simulate the entire regulatory network, propagating node attribute features and extracting various interaction relationships between molecules. Furthermore, CeRVE constructed a comparative analysis matrix of three cancers and adjacent normal tissues through The Cancer Genome Atlas and identified multiple types of potential cancer biomarkers through differential expression analysis of mRNA, microRNA, and long noncoding RNA. Computational testing of multiple types of biomarkers for 72 cancers demonstrates that CeRVE exhibits superior performance in cancer biomarker prediction, providing a powerful tool and insightful approach for AI-assisted disease biomarker discovery. Xinfei Wang 0001, Lan Huang 0002, Yan Wang 0028, Renchu Guan, Zhu-Hong You, Nan Sheng, Xuping Xie, Wenju Hou |
Briefings Bioinform. | 2 |
| 2024 | A multichannel graph neural network based on multisimilarity modality hypergraph contrastive learning for predicting unknown types of cancer biomarkersabstractIdentifying potential cancer biomarkers is a key task in biomedical research, providing a promising avenue for the diagnosis and treatment of human tumors and cancers. In recent years, several machine learning-based RNA-disease association prediction techniques have emerged. However, they primarily focus on modeling relationships of a single type, overlooking the importance of gaining insights into molecular behaviors from a complete regulatory network perspective and discovering biomarkers of unknown types. Furthermore, effectively handling local and global topological structural information of nodes in biological molecular regulatory graphs remains a challenge to improving biomarker prediction performance. To address these limitations, we propose a multichannel graph neural network based on multisimilarity modality hypergraph contrastive learning (MML-MGNN) for predicting unknown types of cancer biomarkers. MML-MGNN leverages multisimilarity modality hypergraph contrastive learning to delve into local associations in the regulatory network, learning diverse insights into the topological structures of multiple types of similarities, and then globally modeling the multisimilarity modalities through a multichannel graph autoencoder. By combining representations obtained from local-level associations and global-level regulatory graphs, MML-MGNN can acquire molecular feature descriptors benefiting from multitype association properties and the complete regulatory network. Experimental results on predicting three different types of cancer biomarkers demonstrate the outstanding performance of MML-MGNN. Furthermore, a case study on gastric cancer underscores the outstanding ability of MML-MGNN to gain deeper insights into molecular mechanisms in regulatory networks and prominent potential in cancer biomarker prediction. Xinfei Wang 0001, Lan Huang 0002, Yan Wang 0028, Renchu Guan, Zhu-Hong You, Nan Sheng, Xuping Xie, Qixing Yang |
Briefings Bioinform. | 2 |
| 2024 | MolFeSCue: enhancing molecular property prediction in data-limited and imbalanced contexts using few-shot and contrastive learningabstractMOTIVATION: Predicting molecular properties is a pivotal task in various scientific domains, including drug discovery, material science, and computational chemistry. This problem is often hindered by the lack of annotated data and imbalanced class distributions, which pose significant challenges in developing accurate and robust predictive models. RESULTS: This study tackles these issues by employing pretrained molecular models within a few-shot learning framework. A novel dynamic contrastive loss function is utilized to further improve model performance in the situation of class imbalance. The proposed MolFeSCue framework not only facilitates rapid generalization from minimal samples, but also employs a contrastive loss function to extract meaningful molecular representations from imbalanced datasets. Extensive evaluations and comparisons of MolFeSCue and state-of-the-art algorithms have been conducted on multiple benchmark datasets, and the experimental data demonstrate our algorithm's effectiveness in molecular representations and its broad applicability across various pretrained models. Our findings underscore MolFeSCues potential to accelerate advancements in drug discovery. AVAILABILITY AND IMPLEMENTATION: We have made all the source code utilized in this study publicly accessible via GitHub at http://www.healthinformaticslab.org/supp/ or https://github.com/zhangruochi/MolFeSCue. The code (MolFeSCue-v1-00) is also available as the supplementary file of this paper. Ruochi Zhang, Chang Liu 0082, Yan Wang 0028, Lan Huang 0002, Fengfeng Zhou |
Bioinform. | 7 |
| 2024 | BEROLECMI: a novel prediction method to infer circRNA-miRNA interaction from the role definition of molecular attributes and biological networksabstractCircular RNA (CircRNA)-microRNA (miRNA) interaction (CMI) is an important model for the regulation of biological processes by non-coding RNA (ncRNA), which provides a new perspective for the study of human complex diseases. However, the existing CMI prediction models mainly rely on the nearest neighbor structure in the biological network, ignoring the molecular network topology, so it is difficult to improve the prediction performance. In this paper, we proposed a new CMI prediction method, BEROLECMI, which uses molecular sequence attributes, molecular self-similarity, and biological network topology to define the specific role feature representation for molecules to infer the new CMI. BEROLECMI effectively makes up for the lack of network topology in the CMI prediction model and achieves the highest prediction performance in three commonly used data sets. In the case study, 14 of the 15 pairs of unknown CMIs were correctly predicted. Xinfei Wang 0001, Zhu-Hong You, Yan Wang 0028, Lan Huang 0002, Yan Qiao 0002, Lei Wang 0121, Zhengwei Li 0001 |
BMC Bioinform. | 5 |
| 2024 | MMGAT: a graph attention network framework for ATAC-seq motifs findingabstractBACKGROUND: Motif finding in Assay for Transposase-Accessible Chromatin using sequencing (ATAC-seq) data is essential to reveal the intricacies of transcription factor binding sites (TFBSs) and their pivotal roles in gene regulation. Deep learning technologies including convolutional neural networks (CNNs) and graph neural networks (GNNs), have achieved success in finding ATAC-seq motifs. However, CNN-based methods are limited by the fixed width of the convolutional kernel, which makes it difficult to find multiple transcription factor binding sites with different lengths. GNN-based methods has the limitation of using the edge weight information directly, makes it difficult to aggregate the neighboring nodes' information more efficiently when representing node embedding. RESULTS: To address this challenge, we developed a novel graph attention network framework named MMGAT, which employs an attention mechanism to adjust the attention coefficients among different nodes. And then MMGAT finds multiple ATAC-seq motifs based on the attention coefficients of sequence nodes and k-mer nodes as well as the coexisting probability of k-mers. Our approach achieved better performance on the human ATAC-seq datasets compared to existing tools, as evidenced the highest scores on the precision, recall, F1_score, ACC, AUC, and PRC metrics, as well as finding 389 higher quality motifs. To validate the performance of MMGAT in predicting TFBSs and finding motifs on more datasets, we enlarged the number of the human ATAC-seq datasets to 180 and newly integrated 80 mouse ATAC-seq datasets for multi-species experimental validation. Specifically on the mouse ATAC-seq dataset, MMGAT also achieved the highest scores on six metrics and found 356 higher-quality motifs. To facilitate researchers in utilizing MMGAT, we have also developed a user-friendly web server named MMGAT-S that hosts the MMGAT method and ATAC-seq motif finding results. CONCLUSIONS: The advanced methodology MMGAT provides a robust tool for finding ATAC-seq motifs, and the comprehensive server MMGAT-S makes a significant contribution to genomics research. The open-source code of MMGAT can be found at https://github.com/xiaotianr/MMGAT , and MMGAT-S is freely available at https://www.mmgraphws.com/MMGAT-S/ . Wenju Hou, Lan Huang 0002, Nan Sheng, Qixing Yang, Shuangquan Zhang, Yan Wang 0028 |
BMC Bioinform. | 4 |
| 2024 | FAM: Improving columnar vision transformer with feature attention mechanism
Lan Huang 0002, Xingyu Bai, Mengqiang Yu, Wei Pang 0001, Kangping Wang |
Comput. Vis. Image Underst. | 1 |
| 2024 | Prompt-based learning framework for zero-shot cross-lingual text classification
Lan Huang 0002, Kangping Wang, Wei Wei 0042, Rui Zhang 0084 |
Eng. Appl. Artif. Intell. | 2 |
| 2024 | FairCare: Adversarial training of a heterogeneous graph neural network with attention mechanism to learn fair representations of electronic health records
Yan Wang 0028, Ruochi Zhang, Qiong Zhou, Shengde Zhang, Yusi Fan, Lan Huang 0002, Fengfeng Zhou |
Inf. Process. Manag. | 7 |
| 2024 | A Survey of Deep Learning for Detecting miRNA- Disease Associations: Databases, Computational Methods, Challenges, and Future DirectionsabstractMicroRNAs (miRNAs) are an important class of non-coding RNAs that play an essential role in the occurrence and development of various diseases. Identifying the potential miRNA-disease associations (MDAs) can be beneficial in understanding disease pathogenesis. Traditional laboratory experiments are expensive and time-consuming. Computational models have enabled systematic large-scale prediction of potential MDAs, greatly improving the research efficiency. With recent advances in deep learning, it has become an attractive and powerful technique for uncovering novel MDAs. Consequently, numerous MDA prediction methods based on deep learning have emerged. In this review, we first summarize publicly available databases related to miRNAs and diseases for MDA prediction. Next, we outline commonly used miRNA and disease similarity calculation and integration methods. Then, we comprehensively review the 48 existing deep learning-based MDA computation methods, categorizing them into classical deep learning and graph neural network-based techniques. Subsequently, we investigate the evaluation methods and metrics that are frequently used to assess MDA prediction performance. Finally, we discuss the performance trends of different computational methods, point out some problems in current research, and propose 9 potential future research directions. Data resources and recent advances in MDA prediction methods are summarized in the GitHub repository https://github.com/sheng-n/DL-miRNA-disease-association-methods. Nan Sheng, Xuping Xie, Yan Wang 0028, Lan Huang 0002, Shuangquan Zhang |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2024 | Meta-GPS++: Enhancing Graph Meta-Learning with Contrastive Learning and Self-TrainingabstractNode classification is an essential problem in graph learning. However, many models typically obtain unsatisfactory performance when applied to few-shot scenarios. Some studies have attempted to combine meta-learning with graph neural networks to solve few-shot node classification on graphs. Despite their promising performance, some limitations remain. First, they employ the node encoding mechanism of homophilic graphs to learn node embeddings, even in heterophilic graphs. Second, existing models based on meta-learning ignore the interference of randomness in the learning process. Third, they are trained using only limited labeled nodes within the specific task, without explicitly utilizing numerous unlabeled nodes. Finally, they treat almost all sampled tasks equally without customizing them for their uniqueness. To address these issues, we propose a novel framework for few-shot node classification called Meta-GPS \(++\) . Specifically, we first adopt an efficient method to learn discriminative node representations on homophilic and heterophilic graphs. Then, we leverage a prototype-based approach to initialize parameters and contrastive learning for regularizing the distribution of node embeddings. Moreover, we apply self-training to extract valuable information from unlabeled nodes. Additionally, we adopt S \({}^{2}\) (scaling and shifting) transformation to learn transferable knowledge from diverse tasks. The results on real-world datasets show the superiority of Meta-GPS \(++\) . Our code is available here . Yonghao Liu 0001, Ximing Li 0002, Lan Huang 0002, Fausto Giunchiglia, Yanchun Liang 0001, Xiaoyue Feng, Renchu Guan |
ACM Trans. Knowl. Discov. Data | 4 |
| 2024 | SSA: A Uniformly Recursive Bidirection-Sequence Systolic Sorter ArrayabstractThe use of reconfigurable circuits with parallel computing capabilities has been explored to enhance sorting performance and reduce power consumption. Nonetheless, most sorting algorithms utilizing dedicated processors are designed solely based on the parallelization of the algorithm, lacking considerations of specialized hardware structures. This leads to problems, including but not limited to the consumption of excessive I/O interface resources, on-chip storage resources, and complex layout wiring. In this paper, we propose a Systolic Sorter Array, implemented by a Uniform Recurrence Equation (URE) with highly parameterised in terms of data size, bit width and type. Leveraging this uniformly recursive structure, the sorter can simultaneously sort two independent sequences. In addition, we implemented global and local control modes on the FPGA to achieve higher computational frequencies. In our experiments, we have demonstrated the speed-up ratio of SSA relative to other state of the art (SOTA) sorting algorithms using C++$std$::$sort()$as benchmark. Inheriting the benefits from the Systolic Array architecture, the SSA reaches up to 810 Mhz computing frequency on the U200. The results of our study show that SSA outperforms other sorting algorithms in terms of throughput, speed-up ratio, and computation frequency. Teng Gao, Lan Huang 0002, Kangping Wang |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2023 | Contrastive self-supervised graph convolutional network for detecting the relationship among lncRNAs, miRNAs, and diseasesabstractInferring potential relationships among long non-coding RNAs (lncRNAs), microRNAs (miRNAs), and diseases play a crucial role in investigation of disease aetiology and pathogenesis. Due to the high cost of laboratory experiments, there is a practical requirement to develop appropriate computational methods that promise to accelerate the experimental screening process for potential lncRNA-disease associations (LDAs), miRNA-disease associations (MDAs), and lncRNA-miRNA interactions (LMIs). However, most existing methods are applied to predict LDAs, MDAs, and LMIs in specific domains, neglecting the important benefits of integrating multiple sources data and limiting the ability of transferring models to other tasks. Furthermore, with the high sparsity of LDA, MDA, and LMI data, it is difficult for many computational models to exploit enough knowledge to learn the comprehensive patterns of node embedding. In this study, inspired by the recent success of graph contrastive learning, we develop a Contrastive Self-supervised Graph convolutional network to identify potential LDAs, MDAs, and LMIs (called CSGLMD). CSGLMD combines supervised learning and self-supervised learning to fully capture node features. Specifically, CSGLMD primarily leverages the rich association and similarity relationships among lncRNA, miRNA, and disease to construct a lncRNA-miRNA-disease heterogeneous graph (LMDHG) that contains three types of biological entities. It can effectively embed multi-source biological data and assist the model extension to other prediction tasks. In addition, we consider applying a label instantiation mechanism to make the LMDHG better adapt graph neural network structures and control the strength of similarity relationships between the same biological entities. Secondly, CSGLMD implements graph convolutional network (GCN) as encoder to extract node embedding features from the LMDHG, and utilizes a multi-relational modelling decoder to predict LDAs, MDAs, or LMIs. Finally, we designed a contrastive self-supervised learning task that guides the learning of node embeddings without relying on labels, and acts as a regularize in a multi-task learning paradigm to enhance the generalization ability of the model. Extensive results on two datasets (from the old and new versions of the database, respectively) show that CSGLMD significantly outperforms 12 state-of-the-art methods (5 LDA prediction and 7 MDA prediction) in predicting disease-associated lncRNAs and miRNAs. Case studies on old and new datasets can further demonstrate the capability of CSGLMD to discover disease-related new candidate lncRNAs and miRNAs. The source data and code for the proposed model are publicly available on https://github.com/sheng-n/CSGLMD. Nan Sheng, Lan Huang 0002, Yan Wang 0028, Huiyan Sun, Xuping Xie |
BIBM | 2 |
| 2023 | Deep Feature-Based Text Clustering and Its ExplanationabstractText clustering is a critical step in text data analysis and has been extensively studied by the text mining community. Most existing text clustering algorithms are based on the bag-of-words model, which faces the high-dimensional and sparsity problems and ignores text structural and sequence information. Deep learning-based models such as convolutional neural networks and recurrent neural networks regard texts as sequences but lack supervised signals and explainable results. In this paper, we propose a deep feature-based text clustering (DFTC) framework that incorporates pretrained text encoders into text clustering tasks. This model, which is based on sequence representations, breaks the dependency on supervision. The experimental results show that our model outperforms classic text clustering algorithms on almost all the considered datasets. In addition, the explanation of the clustering results is significant for understanding the principles of the deep learning approach. Our proposed clustering framework includes an explanation module that can help users understand the meaning and quality of the clustering results. Our code is available at https://github.com/KEAML-JLU/DeepTextClustering. Renchu Guan, Yanchun Liang 0001, Fausto Giunchiglia, Lan Huang 0002, Xiaoyue Feng |
ICDE | 5 |
| 2023 | Local and Global: Temporal Question Answering via Information FusionabstractMany models that leverage knowledge graphs (KGs) have recently demonstrated remarkable success in question answering (QA) tasks. In the real world, many facts contained in KGs are time-constrained thus temporal KGQA has received increasing attention. Despite the fruitful efforts of previous models in temporal KGQA, they still have several limitations. (I) They neither emphasize the graph structural information between entities in KGs nor explicitly utilize a multi-hop relation path through graph neural networks to enhance answer prediction. (II) They adopt pre-trained language models (LMs) to obtain question representations, focusing merely on the global information related to the question while not highlighting the local information of the entities in KGs. To address these limitations, we introduce a novel model that simultaneously explores both Local information and Global information for the task of temporal KGQA (LGQA). Specifically, we first introduce an auxiliary task in the temporal KG embedding procedure to make timestamp embeddings time-order aware. Then, we design information fusion layers that effectively incorporate local and global information to deepen question understanding. We conduct extensive experiments on two benchmarks, and LGQA significantly outperforms previous state-of-the-art models, especially in difficult questions. Moreover, LGQA can generate interpretable and trustworthy predictions. Yonghao Liu 0001, Di Liang, Fausto Giunchiglia, Ximing Li 0002, Lan Huang 0002, Xiaoyue Feng, Renchu Guan |
IJCAI | 8 |
| 2023 | Orchestrating information across tissues via a novel multitask GAT framework to improve quantitative gene regulation relation modeling for survival analysisabstractSurvival analysis is critical to cancer prognosis estimation. High-throughput technologies facilitate the increase in the dimension of genic features, but the number of clinical samples in cohorts is relatively small due to various reasons, including difficulties in participant recruitment and high data-generation costs. Transcriptome is one of the most abundantly available OMIC (referring to the high-throughput data, including genomic, transcriptomic, proteomic and epigenomic) data types. This study introduced a multitask graph attention network (GAT) framework DQSurv for the survival analysis task. We first used a large dataset of healthy tissue samples to pretrain the GAT-based HealthModel for the quantitative measurement of the gene regulatory relations. The multitask survival analysis framework DQSurv used the idea of transfer learning to initiate the GAT model with the pretrained HealthModel and further fine-tuned this model using two tasks i.e. the main task of survival analysis and the auxiliary task of gene expression prediction. This refined GAT was denoted as DiseaseModel. We fused the original transcriptomic features with the difference vector between the latent features encoded by the HealthModel and DiseaseModel for the final task of survival analysis. The proposed DQSurv model stably outperformed the existing models for the survival analysis of 10 benchmark cancer types and an independent dataset. The ablation study also supported the necessity of the main modules. We released the codes and the pretrained HealthModel to facilitate the feature encodings and survival analysis of transcriptome-based future studies, especially on small datasets. The model and the code are available at http://www.healthinformaticslab.org/supp/. Meiyu Duan, Yueying Wang, Gongyou Zhang, Haotian Zhang 0018, Lan Huang 0002, Ruochi Zhang, Fengfeng Zhou |
Briefings Bioinform. | 8 |
| 2023 | Multi-task prediction-based graph contrastive learning for inferring the relationship among lncRNAs, miRNAs and diseasesabstractMOTIVATION: Identifying the relationships among long non-coding RNAs (lncRNAs), microRNAs (miRNAs) and diseases is highly valuable for diagnosing, preventing, treating and prognosing diseases. The development of effective computational prediction methods can reduce experimental costs. While numerous methods have been proposed, they often to treat the prediction of lncRNA-disease associations (LDAs), miRNA-disease associations (MDAs) and lncRNA-miRNA interactions (LMIs) as separate task. Models capable of predicting all three relationships simultaneously remain relatively scarce. Our aim is to perform multi-task predictions, which not only construct a unified framework, but also facilitate mutual complementarity of information among lncRNAs, miRNAs and diseases. RESULTS: In this work, we propose a novel unsupervised embedding method called graph contrastive learning for multi-task prediction (GCLMTP). Our approach aims to predict LDAs, MDAs and LMIs by simultaneously extracting embedding representations of lncRNAs, miRNAs and diseases. To achieve this, we first construct a triple-layer lncRNA-miRNA-disease heterogeneous graph (LMDHG) that integrates the complex relationships between these entities based on their similarities and correlations. Next, we employ an unsupervised embedding model based on graph contrastive learning to extract potential topological feature of lncRNAs, miRNAs and diseases from the LMDHG. The graph contrastive learning leverages graph convolutional network architectures to maximize the mutual information between patch representations and corresponding high-level summaries of the LMDHG. Subsequently, for the three prediction tasks, multiple classifiers are explored to predict LDA, MDA and LMI scores. Comprehensive experiments are conducted on two datasets (from older and newer versions of the database, respectively). The results show that GCLMTP outperforms other state-of-the-art methods for the disease-related lncRNA and miRNA prediction tasks. Additionally, case studies on two datasets further demonstrate the ability of GCLMTP to accurately discover new associations. To ensure reproducibility of this work, we have made the datasets and source code publicly available at https://github.com/sheng-n/GCLMTP. Nan Sheng, Yan Wang 0028, Lan Huang 0002, Yangkun Cao, Xuping Xie |
Briefings Bioinform. | 3 |
| 2023 | Cross-domain endoscopic image translation and landmark detection based on consistency regularization cycle generative adversarial network
Lan Huang 0002, Yuzhao Wang, Yingfang Zhang, Shuyu Guo, Ke Tao, Tian Bai 0002 |
Expert Syst. Appl. | 1 |
| 2023 | GSRNet, an adversarial training-based deep framework with multi-scale CNN and BiGRU for predicting genomic signals and regions
Gancheng Zhu, Yusi Fan, Fei Li 0039, Annebella Tsz Ho Choi, Zhikang Tan, Yiruo Cheng, Changfan Luo, Gongyou Zhang, Zhaomin Yao, Lan Huang 0002, Fengfeng Zhou |
Expert Syst. Appl. | 14 |
| 2023 | U-DARTS: Uniform-space differentiable architecture search
Lan Huang 0002, Wencong Wang, Wei Pang 0001, Kangping Wang |
Inf. Sci. | 1 |
| 2023 | EvaGoNet: An integrated network of variational autoencoder and Wasserstein generative adversarial network with gradient penalty for binary classification tasks
Changfan Luo, Yongkang Shao, Jianzheng Hu, Meiyu Duan, Lan Huang 0002, Fengfeng Zhou |
Inf. Sci. | 9 |
| 2023 | A Hybrid VAE Based Network Embedding Method for Biomedical Relation Mining
Tian Bai 0002, Lan Huang 0002 |
Neural Process. Lett. | 4 |
| 2023 | A Survey of Computational Methods and Databases for lncRNA-MiRNA Interaction PredictionabstractLong non-coding RNAs (lncRNAs) and microRNAs (miRNAs) are two prevalent non-coding RNAs in current research. They play critical regulatory roles in the life processes of animals and plants. Studies have shown that lncRNAs can interact with miRNAs to participate in post-transcriptional regulatory processes, mainly involved in regulating cancer development, metastatic progression, and drug resistance. Additionally, these interactions have significant effects on plant growth, development, and responses to biotic and abiotic stresses. Deciphering the potential relationships between lncRNAs and miRNAs may provide new insights into our understanding of the biological functions of lncRNAs and miRNAs, and the pathogenesis of complex diseases. In contrast, gathering information on lncRNA-miRNA interactions (LMIs) through biological experiments is expensive and time-consuming. With the accumulation of multi-omics data, computational models are extremely attractive in systematically exploring potential LMIs. To the best of our knowledge, this is the first comprehensive review of computational methods for identifying LMIs. Specifically, we first summarized the available public databases for predicting animal and plant LMIs. Second, we comprehensively reviewed the computational methods for predicting LMIs and classified them into two categories, including network-based methods and sequence-based methods. Third, we analyzed the standard evaluation methods and metrics used in LMI prediction. Finally, we pointed out some problems in the current study and discuss future research directions. Relevant databases and the latest advances in LMI prediction are summarized in a GitHub repository https://github.com/sheng-n/lncRNA-miRNA-interaction-methods, and we'll keep it updated. Nan Sheng, Lan Huang 0002, Yangkun Cao, Xuping Xie, Yan Wang 0028 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2022 | Artificial intelligence in clinical research of cancersabstractSeveral factors, including advances in computational algorithms, the availability of high-performance computing hardware, and the assembly of large community-based databases, have led to the extensive application of Artificial Intelligence (AI) in the biomedical domain for nearly 20 years. AI algorithms have attained expert-level performance in cancer research. However, only a few AI-based applications have been approved for use in the real world. Whether AI will eventually be capable of replacing medical experts has been a hot topic. In this article, we first summarize the cancer research status using AI in the past two decades, including the consensus on the procedure of AI based on an ideal paradigm and current efforts of the expertise and domain knowledge. Next, the available data of AI process in the biomedical domain are surveyed. Then, we review the methods and applications of AI in cancer clinical research categorized by the data types including radiographic imaging, cancer genome, medical records, drug information and biomedical literatures. At last, we discuss challenges in moving AI from theoretical research to real-world cancer research applications and the perspectives toward the future realization of AI participating cancer treatment. Dan Shao, Yinfei Dai, Nianfeng Li, Xuqing Cao, Zhuqing Rong, Lan Huang 0002, Yan Wang 0028 |
Briefings Bioinform. | 8 |
| 2022 | Multi-channel graph attention autoencoders for disease-related lncRNAs predictionabstractMOTIVATION: Predicting disease-related long non-coding RNAs (lncRNAs) can be used as the biomarkers for disease diagnosis and treatment. The development of effective computational prediction approaches to predict lncRNA-disease associations (LDAs) can provide insights into the pathogenesis of complex human diseases and reduce experimental costs. However, few of the existing methods use microRNA (miRNA) information and consider the complex relationship between inter-graph and intra-graph in complex-graph for assisting prediction. RESULTS: In this paper, the relationships between the same types of nodes and different types of nodes in complex-graph are introduced. We propose a multi-channel graph attention autoencoder model to predict LDAs, called MGATE. First, an lncRNA-miRNA-disease complex-graph is established based on the similarity and correlation among lncRNA, miRNA and diseases to integrate the complex association among them. Secondly, in order to fully extract the comprehensive information of the nodes, we use graph autoencoder networks to learn multiple representations from complex-graph, inter-graph and intra-graph. Thirdly, a graph-level attention mechanism integration module is adopted to adaptively merge the three representations, and a combined training strategy is performed to optimize the whole model to ensure the complementary and consistency among the multi-graph embedding representations. Finally, multiple classifiers are explored, and Random Forest is used to predict the association score between lncRNA and disease. Experimental results on the public dataset show that the area under receiver operating characteristic curve and area under precision-recall curve of MGATE are 0.964 and 0.413, respectively. MGATE performance significantly outperformed seven state-of-the-art methods. Furthermore, the case studies of three cancers further demonstrate the ability of MGATE to identify potential disease-correlated candidate lncRNAs. The source code and supplementary data are available at https://github.com/sheng-n/MGATE. CONTACT: [email protected], [email protected]. Nan Sheng, Lan Huang 0002, Yan Wang 0028, Ping Xuan, Yangkun Cao |
Briefings Bioinform. | 2 |
| 2022 | HLAB: learning the BiLSTM features from the ProtBert-encoded proteins for the class I HLA-peptide binding predictionabstractHuman Leukocyte Antigen (HLA) is a type of molecule residing on the surfaces of most human cells and exerts an essential role in the immune system responding to the invasive items. The T cell antigen receptors may recognize the HLA-peptide complexes on the surfaces of cancer cells and destroy these cancer cells through toxic T lymphocytes. The computational determination of HLA-binding peptides will facilitate the rapid development of cancer immunotherapies. This study hypothesized that the natural language processing-encoded peptide features may be further enriched by another deep neural network. The hypothesis was tested with the Bi-directional Long Short-Term Memory-extracted features from the pretrained Protein Bidirectional Encoder Representations from Transformers-encoded features of the class I HLA (HLA-I)-binding peptides. The experimental data showed that our proposed HLAB feature engineering algorithm outperformed the existing ones in detecting the HLA-I-binding peptides. The extensive evaluation data show that the proposed HLAB algorithm outperforms all the seven existing studies on predicting the peptides binding to the HLA-A*01:01 allele in AUC and achieves the best average AUC values on the six out of the seven k-mers (k=8,9,...,14, respectively represent the prediction task of a polypeptide consisting of k amino acids) except for the 9-mer prediction tasks. The source code and the fine-tuned feature extraction models are available at http://www.healthinformaticslab.org/supp/resources.php. Gancheng Zhu, Fei Li 0039, Lan Huang 0002, Meiyu Duan, Fengfeng Zhou |
Briefings Bioinform. | 5 |
| 2022 | Research on Reverse Skyline Query Algorithm Based on Decision SetabstractReverse skyline query is an extension of the classical skyline query, widely used in the decision support in e-business. The vast burst of big data in e-business challenges the classical algorithms for such queries. This paper provides a novel definition of decision set and a decision set based reverse skyline query method called DRS on the double-layer R tree indexing in a map-reduce manner. Theoretical proofs are provided for the correctness and complexity of the DRS algorithm. Experiments made using several large data sets are presented and analyzed to illustrate the applicability and the outperformance of DRS over the state-of-the-art reverse skyline query methods. Lan Huang 0002, Yuanwei Zhao, Pedro Mestre, Laipeng Han, Kangping Wang, Wenjuan Gao, Rui Zhang 0084 |
J. Database Manag. | 1 |
| 2022 | Traditional Chinese medicine entity relation extraction based on CNN with segment attention
Tian Bai 0002, Haotian Guan, Lan Huang 0002 |
Neural Comput. Appl. | 5 |
| 2022 | Deep Feature-Based Text Clustering and its ExplanationabstractText clustering is a critical step in text data analysis and has been extensively studied by the text mining community. Most existing text clustering algorithms are based on the bag-of-words model, which faces the high-dimensional and sparsity problems and ignores text structural and sequence information. Deep learning-based models such as convolutional neural networks and recurrent neural networks regard texts as sequences but lack supervised signals and explainable results. In this paper, we propose adeepfeature-basedtextclustering (DFTC) framework that incorporates pretrained text encoders into text clustering tasks. This model, which is based on sequence representations, breaks the dependency on supervision. The experimental results show that our model outperforms classic text clustering algorithms and the state-of-the-art pretrained language model, i.e., BERT, on almost all the considered datasets. In addition, the explanation of the clustering results is significant for understanding the principles of the deep learning approach. Our proposed clustering framework includes an explanation module that can help users understand the meaning and quality of the clustering results. Renchu Guan, Yanchun Liang 0001, Fausto Giunchiglia, Lan Huang 0002, Xiaoyue Feng |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2021 | COVID-19 Knowledge Graph for Drug and Vaccine DevelopmentabstractThe worldwide spread of COVID-19 has made a severe impact on human health and life. It has shown rapid propagation, long in vitro survival, and a long incubation period. More seriously, COVID-19 is more susceptible to variation, as it is an RNA virus. Mutations of COVID-19 have been reported in multiple countries worldwide, which makes drug and vaccine development a significant challenge. To search for potential drugs and vaccines and reveal the atlas of COVID-19 evolution, we extract information from massive unstructured data and construct a COVID-19 knowledge graph using the COVID-19 data. Based on machine learning approaches, we infer and predict novel coronavirus pneumonia-related diseases, drug action targets, etc. to speculate on new and more effective treatment methods. In addition, to study transcriptome of SARS-CoV-2, new ideas can be provided to biomedical experts with flexible responses to viral variation. An in-depth analysis of the COVID-19 pathomechanism at the pharmaceutical, genetic, and protein levels provides effective means and tools for novel coronavirus pneumonia vaccines, drug development, and therapeutic program design. Lan Huang 0002, Hongrui Guan, Yanchun Liang 0001, Renchu Guan, Xiaoyue Feng |
BIBM | 1 |
| 2021 | Research and Application of Reinforcement Learning Recommendation Method for TaobaoabstractNowadays, many e-commerce companies are using reinforcement learning recommendation methods to maximize long-term benefits. Alibaba Group and Nanjing University build “Virtual Taobao”, a Taobao simulator. In this paper, we proposed TTD3 based on TD3 and trained it in Virtual Taobao. There are three important improvements in TTD3's training process. First, the current actor-network and target actor-network will predict two candidate actions for Virtual Taobao's current state, and the action with a larger value evaluated by the current critic-network is selected as the final execution action. Second, the Ornstein-Uhlenbeck (OU) process is used as the exploration noise to improve the agent's ability to explore Virtual Taobao. Third, prioritized experience replay is adopted to improve sampling efficiency. TTD3 achieves the highest average CTR of about 0.85 in Virtual Taobao which is superior to TD3 as well as DPPO, SAC, and DDPG used by Virtual Taobao's author. Lan Huang 0002, Yan Wang 0028, Xuping Xie |
ISCC | 1 |
| 2021 | A comprehensive comparison of residue-level methylation levels with the regression-based gene-level methylation estimations by ReGearabstractMOTIVATION: DNA methylation is a biological process impacting the gene functions without changing the underlying DNA sequence. The DNA methylation machinery usually attaches methyl groups to some specific cytosine residues, which modify the chromatin architectures. Such modifications in the promoter regions will inactivate some tumor-suppressor genes. DNA methylation within the coding region may significantly reduce the transcription elongation efficiency. The gene function may be tuned through some cytosines are methylated. METHODS: This study hypothesizes that the overall methylation level across a gene may have a better association with the sample labels like diseases than the methylations of individual cytosines. The gene methylation level is formulated as a regression model using the methylation levels of all the cytosines within this gene. A comprehensive evaluation of various feature selection algorithms and classification algorithms is carried out between the gene-level and residue-level methylation levels. RESULTS: A comprehensive evaluation was conducted to compare the gene and cytosine methylation levels for their associations with the sample labels and classification performances. The unsupervised clustering was also improved using the gene methylation levels. Some genes demonstrated statistically significant associations with the class label, even when no residue-level methylation features have statistically significant associations with the class label. So in summary, the trained gene methylation levels improved various methylome-based machine learning models. Both methodology development of regression algorithms and experimental validation of the gene-level methylation biomarkers are worth of further investigations in the future studies. The source code, example data files and manual are available at http://www.healthinformaticslab.org/supp/. Jinpu Cai, Yuyang Xu, Shiying Ding, Yuewei Sun, Jingyi Lyu, Meiyu Duan, Shuai Liu 0010, Lan Huang 0002, Fengfeng Zhou |
Briefings Bioinform. | 9 |
| 2021 | Human body-fluid proteome: quantitative profiling and computational predictionabstractEmpowered by the advancement of high-throughput bio technologies, recent research on body-fluid proteomes has led to the discoveries of numerous novel disease biomarkers and therapeutic drugs. In the meantime, a tremendous progress in disclosing the body-fluid proteomes was made, resulting in a collection of over 15 000 different proteins detected in major human body fluids. However, common challenges remain with current proteomics technologies about how to effectively handle the large variety of protein modifications in those fluids. To this end, computational effort utilizing statistical and machine-learning approaches has shown early successes in identifying biomarker proteins in specific human diseases. In this article, we first summarized the experimental progresses using a combination of conventional and high-throughput technologies, along with the major discoveries, and focused on current research status of 16 types of body-fluid proteins. Next, the emerging computational work on protein prediction based on support vector machine, ranking algorithm, and protein-protein interaction network were also surveyed, followed by algorithm and application discussion. At last, we discuss additional critical concerns about these topics and close the review by providing future perspectives especially toward the realization of clinical disease biomarker discovery. Lan Huang 0002, Dan Shao, Yan Wang 0028, Xueteng Cui, Juan Cui |
Briefings Bioinform. | 1 |
| 2021 | A dynamic recursive feature elimination framework (dRFE) to further refine a set of OMIC biomarkersabstractMOTIVATION: A feature selection algorithm may select the subset of features with the best associations with the class labels. The recursive feature elimination (RFE) is a heuristic feature screening framework and has been widely used to select the biological OMIC biomarkers. This study proposed a dynamic recursive feature elimination (dRFE) framework with more flexible feature elimination operations. The proposed dRFE was comprehensively compared with 11 existing feature selection algorithms and five classifiers on the eight difficult transcriptome datasets from a previous study, the ten newly collected transcriptome datasets and the five methylome datasets. RESULTS: The experimental data suggested that the regular RFE framework did not perform well, and dRFE outperformed the existing feature selection algorithms in most cases. The dRFE-detected features achieved Acc = 1.0000 for the two methylome datasets GSE53045 and GSE66695. The best prediction accuracies of the dRFE-detected features were 0.9259, 0.9424 and 0.8601 for the other three methylome datasets GSE74845, GSE103186 and GSE80970, respectively. Four transcriptome datasets received Acc = 1.0000 using the dRFE-detected features, and the prediction accuracies for the other six newly collected transcriptome datasets were between 0.6301 and 0.9917. AVAILABILITY AND IMPLEMENTATION: The experiments in this study are implemented and tested using the programming language Python version 3.7.6. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Lan Huang 0002, Fengfeng Zhou |
Bioinform. | 2 |
| 2021 | DeepSec: a deep learning framework for secreted protein discovery in human body fluidsabstractMOTIVATION: Human proteins that are secreted into different body fluids from various cells and tissues can be promising disease indicators. Modern proteomics research empowered by both qualitative and quantitative profiling techniques has made great progress in protein discovery in various human fluids. However, due to the large number of proteins and diverse modifications present in the fluids, as well as the existing technical limits of major proteomics platforms (e.g. mass spectrometry), large discrepancies are often generated from different experimental studies. As a result, a comprehensive proteomics landscape across major human fluids are not well determined. RESULTS: To bridge this gap, we have developed a deep learning framework, named DeepSec, to identify secreted proteins in 12 types of human body fluids. DeepSec adopts an end-to-end sequence-based approach, where a Convolutional Neural Network is built to learn the abstract sequence features followed by a Bidirectional Gated Recurrent Unit with fully connected layer for protein classification. DeepSec has demonstrated promising performances with average area under the ROC curves of 0.85-0.94 on testing datasets in each type of fluids, which outperforms existing state-of-the-art methods available mostly on blood proteins. As an illustration of how to apply DeepSec in biomarker discovery research, we conducted a case study on kidney cancer by using genomics data from the cancer genome atlas and have identified 104 possible marker proteins. AVAILABILITY: DeepSec is available at https://bmbl.bmi.osumc.edu/deepsec/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Dan Shao, Lan Huang 0002, Yan Wang 0028, Xueteng Cui, Yao Wang 0010, Qin Ma 0003, Juan Cui |
Bioinform. | 2 |
| 2021 | Recognizing art work image from natural type: a deep adaptive depiction fusion method
Lan Huang 0002, Yuzhao Wang, Tian Bai 0002 |
Vis. Comput. | 1 |
| 2020 | An Incremental Learning Network Model Based on Random Sample Distribution Fitting
Wencong Wang, Lan Huang 0002, Kainuo Li, Kangping Wang |
KSEM (2) | 2 |
| 2020 | Feature selection may improve deep neural networks for the bioinformatics problemsabstractMOTIVATION: Deep neural network (DNN) algorithms were utilized in predicting various biomedical phenotypes recently, and demonstrated very good prediction performances without selecting features. This study proposed a hypothesis that the DNN models may be further improved by feature selection algorithms. RESULTS: A comprehensive comparative study was carried out by evaluating 11 feature selection algorithms on three conventional DNN algorithms, i.e. convolution neural network (CNN), deep belief network (DBN) and recurrent neural network (RNN), and three recent DNNs, i.e. MobilenetV2, ShufflenetV2 and Squeezenet. Five binary classification methylomic datasets were chosen to calculate the prediction performances of CNN/DBN/RNN models using feature selected by the 11 feature selection algorithms. Seventeen binary classification transcriptome and two multi-class transcriptome datasets were also utilized to evaluate how the hypothesis may generalize to different data types. The experimental data supported our hypothesis that feature selection algorithms may improve DNN models, and the DBN models using features selected by SVM-RFE usually achieved the best prediction accuracies on the five methylomic datasets. AVAILABILITY AND IMPLEMENTATION: All the algorithms were implemented and tested under the programming environment Python version 3.6.6. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Shuainan Li, Xiaoyue Feng, Xin Feng 0004, Yexian Zhang, Meiyu Duan, Lan Huang 0002, Fengfeng Zhou |
Bioinform. | 11 |
| 2020 | sefOri: selecting the best-engineered sequence features to predict DNA replication originsabstractMOTIVATION: Cell divisions start from replicating the double-stranded DNA, and the DNA replication process needs to be precisely regulated both spatially and temporally. The DNA is replicated starting from the DNA replication origins. A few successful prediction models were generated based on the assumption that the DNA replication origin regions have sequence level features like physicochemical properties significantly different from the other DNA regions. RESULTS: This study proposed a feature selection procedure to further refine the classification model of the DNA replication origins. The experimental data demonstrated that as large as 26% improvement in the prediction accuracy may be achieved on the yeast Saccharomyces cerevisiae. Moreover, the prediction accuracies of the DNA replication origins were improved for all the four yeast genomes investigated in this study. AVAILABILITY AND IMPLEMENTATION: The software sefOri version 1.0 was available at http://www.healthinformaticslab.org/supp/resources.php. An online server was also provided for the convenience of the users, and its web link may be found in the above-mentioned web page. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Chenwei Lou, Ruoyao Shi, Wenyang Zhou, Yubo Wang 0006, Lan Huang 0002, Xin Feng 0004, Fengfeng Zhou |
Bioinform. | 8 |
| 2020 | A novel deep learning method for extracting unspecific biomedical relationabstractSummary Biomedical relation extraction is an important research subject in Natural language processing (NLP). Deep learning technology has shown greater value in improving accuracy of relation extraction results recently. Existing methods mostly focus on extracting (1) specific relation from short texts (eg, drug‐drug interaction and protein‐protein interaction) and (2) unspecific relation from full text corpora. However, extracting unspecific relation from short text, which is more and more important in practical use, is rarely studied. In this paper, a new model called MAT‐LSTM is proposed to extract unspecific relation from short text in biomedical literatures. Experiments on two Biocreative benchmark datasets and one BioNLP benchmark datasets were made to measure the validity of the proposed model MAT‐LSTM, and better performance is achieved. The MAT‐LSTM model is also applied practically in extracting unspecific relation contained in the PubMed literatures. The results extracted from PubMed by using the proposed model were verified by experts mostly, indicating the practical value of the MAT‐LSTM model. Tian Bai 0002, Lan Huang 0002, Fuyong Xing |
Concurr. Comput. Pract. Exp. | 4 |
| 2020 | A Survey on Performance Optimization of High-Level Synthesis Tools
Lan Huang 0002, Dalin Li, Kangping Wang, Teng Gao, Adriano Tavares |
J. Comput. Sci. Technol. | 1 |
| 2020 | An Extended Nonstrict Partially Ordered Set-Based Configurable Linear Sorter on FPGAsabstractSorting is essential for many scientific and data processing problems. It is significant to improve the efficiency of sorting. Taking advantage of specialized hardware, parallel sorting, e.g., sorting networks and linear sorters, implements sorting in lower time complexity. However, most of them are designed based on the parallelization of algorithms, lacking consideration of specialized hardware structures. In this article, we propose an extended nonstrict partially ordered set-based configurable linear sorter on field-programmable gate arrays (FPGAs). First, we extend nonstrict partial order to the binary tuple and n-tuple nonstrict partial orders. Then, the linear sorting algorithm is defined based on them, with the consideration of hardware performance. It has 4N/n time complexity varying from 4 to 2 N as the tuple size varies. The number of comparisons reduces to N/2 in binary tuple-based sorting, which is half of the state-of-the-art insertion linear sorting. Finally, we implement the linear sorter on FPGAs. It consists of multiple customizable micro-cores, named sorting units (SUs). The SU packages the storage and comparison of the tuple. All the SUs are connected into a chain with simple communication, which makes the sorter fully configurable in length, bandwidth, and throughput. They also act the same in each clock cycle, so that the achieved frequency of the sorter improves. In our experiment, the sorter achieves at most 660-MHz frequency, 5.6 Gb/s throughput, and 87 times speed-up compared with the quick sort algorithm on general processors. Dalin Li, Lan Huang 0002, Teng Gao, Adriano Tavares, Kangping Wang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2019 | Computational Prediction of Human Body-Fluid ProteinabstractResearch on body fluid proteomes has led to the discoveries of context-dependent proteomics profiles and numerous novel disease biomarkers. Common challenges remain with current proteomics technologies about how to effectively handle the large variety of protein modifications in those fluids. To this end, computational efforts have shown early successes in identifying biomarker proteins in specific human diseases. In this article, we first reviewed published computational methods on this topic and then presented a new database system and a novel deep learning-based method for predicting body-fluid proteins in 16 types of human fluids. The results show that our new system outperforms existing methods and provides a highly promising tool to facilitate clinical proteomic discovery. Dan Shao, Lan Huang 0002, Yan Wang 0028, Xueteng Cui, Yao Wang 0010 |
BIBM | 2 |
| 2019 | An Improved Selection Operator for Multi-objective Optimization
Zhi-hui Zhan, Weineng Chen, Tianlong Gu, Renchu Guan, Lan Huang 0002, Jun Zhang 0003 |
ISNN (1) | 7 |
| 2019 | Resource Allocation Scheme Based on Rate-Requirement for Device-to-Device Downlink CommunicationsabstractThe rate-requirement of device-to-device (D2D) users is associated with the context information of velocity and data size of users to some extent. In this study, an efficient context-aware resource allocation scheme based on rate requirement (RARR) is proposed. This scheme consists of two allocation phases. In the rate-ensuring resource allocation phase, D2D pairs are allocated a certain amount of spectrum resource according to their rate requirement. In the allocation, the interference restricted area is limited to exclude cellular users that bring a negative capacity gain to the communication system. In the residual resource reallocation phase, surplus resources are assigned to D2D pairs according to the system fairness. Simulation results indicate that the proposed RARR scheme efficiently leads to superior performance in terms of system throughput and fairness and exhibits low complexity relative to traditional resource allocation. Xin Wang 0050, Zhihong Qian, Xue Wang 0002, Lan Huang 0002 |
Int. J. Pattern Recognit. Artif. Intell. | 4 |
| 2019 | A novel MEDLINE topic indexing method using image presentation
Lan Huang 0002, Shuyu Guo, Leiguang Gong, Tian Bai 0002 |
J. Vis. Commun. Image Represent. | 2 |
| 2018 | ε-Distance Weighted Support Vector Regression
Ge Ou, Yan Wang 0028, Lan Huang 0002, Wei Pang 0001, George Macleod Coghill |
PAKDD (1) | 3 |
| 2017 | A Novel Diversity Measure for Understanding Movie Ranks in Movie Collaboration Networks
Manqing Ma, Wei Pang 0001, Lan Huang 0002, Zhe Wang 0007 |
PAKDD (1) | 3 |
| 2017 | An improved fruit fly optimization algorithm for solving traveling salesman problemabstractThe traveling salesman problem (TSP), a typical non-deterministic polynomial (NP) hard problem, has been used in many engineering applications. As a new swarm-intelligence optimization algorithm, the fruit fly optimization algorithm (FOA) is used to solve TSP, since it has the advantages of being easy to understand and having a simple implementation. However, it has problems, including a slow convergence rate for the algorithm, easily falling into the local optimum, and an insufficient optimi-zation precision. To address TSP effectively, three improvements are proposed in this paper to improve FOA. First, the vision search process is reinforced in the foraging behavior of fruit flies to improve the convergence rate of FOA. Second, an elimination mechanism is added to FOA to increase the diversity. Third, a reverse operator and a multiplication operator are proposed. They are performed on the solution sequence in the fruit fly’s smell search and vision search processes, respectively. In the experiment, 10 benchmarks selected from TSPLIB are tested. The results show that the improved FOA outperforms other alternatives in terms of the convergence rate and precision. Lan Huang 0002, Gui-chao Wang, Tian Bai 0002, Zhe Wang 0007 |
Frontiers Inf. Technol. Electron. Eng. | 1 |
| 2016 | Partitioning Clustering Based on Support Vector Ranking
Qing Peng, Yan Wang 0028, Ge Ou, Yuan Tian 0016, Lan Huang 0002, Wei Pang 0001 |
ADMA | 5 |
| 2016 | A method for exploring implicit concept relatedness in biomedical knowledge networkabstractBACKGROUND: Biomedical information and knowledge, structural and non-structural, stored in different repositories can be semantically connected to form a hybrid knowledge network. How to compute relatedness between concepts and discover valuable but implicit information or knowledge from it effectively and efficiently is of paramount importance for precision medicine, and a major challenge facing the biomedical research community. RESULTS: In this study, a hybrid biomedical knowledge network is constructed by linking concepts across multiple biomedical ontologies as well as non-structural biomedical knowledge sources. To discover implicit relatedness between concepts in ontologies for which potentially valuable relationships (implicit knowledge) may exist, we developed a Multi-Ontology Relatedness Model (MORM) within the knowledge network, for which a relatedness network (RN) is defined and computed across multiple ontologies using a formal inference mechanism of set-theoretic operations. Semantic constraints are designed and implemented to prune the search space of the relatedness network. CONCLUSIONS: Experiments to test examples of several biomedical applications have been carried out, and the evaluation of the results showed an encouraging potential of the proposed approach to biomedical knowledge discovery. Tian Bai 0002, Leiguang Gong, Yan Wang 0028, Casimir A. Kulikowski, Lan Huang 0002 |
BMC Bioinform. | 6 |
| 2015 | Implicit knowledge discovery in biomedical ontologies: Computing interesting relatednessesabstractOntologies, seen as effective representations for sharing and reusing knowledge, have become increasingly important in biomedicine, usually focusing on taxonomic knowledge specific to a subject. Efforts have been made to uncover implicit knowledge within large biomedical ontologies by exploring semantic similarity and relatedness between concepts. However, much less attention has been paid to another potentially helpful approach: discovering implicit knowledge across multiple ontologies of different types, such as disease ontologies, symptom ontologies, and gene ontologies. In this paper, we propose a unified approach to the problem of ontology based implicit knowledge discovery - a Multi-Ontology Relatedness Model (MORM), which includes the formation of multiple related ontologies, a relatedness network and a formal inference mechanism based on set-theoretic operations. Experiments for biomedical applications have been carried out, and preliminary results show the potential value of the proposed approach for biomedical knowledge discovery. Tian Bai 0002, Leiguang Gong, Casimir A. Kulikowski, Lan Huang 0002 |
BIBM | 4 |
| 2009 | Immune Particle Swarm Optimization for Support Vector Regression on Forest Fire Prediction
Yan Wang 0028, Juexin Wang, Wei Du 0002, Chuncai Wang, Yanchun Liang 0001, Chunguang Zhou, Lan Huang 0002 |
ISNN (2) | 7 |
| 2008 | Research on Data Mining Algorithms for Automotive Customers' Behavior Prediction ProblemabstractThis paper extracts automotive marketing information, constructs data warehouse, adopts an improved ID3 decision tree model and an association rule model to do data mining, and then obtains prediction information of automotive customers' behavior. Experimental and comparative results verify the validity and accuracy of the prediction results. Lan Huang 0002, Chunguang Zhou, Yu-qin Zhou, Zhe Wang 0007 |
ICMLA | 1 |