Leon Wong

dblp:01/3115 · DBLP profile ↗
← Back
33ranked-venue papers
3as first author
23since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 21 · 3 first-author · 13 since 2021Systems, architecture and hardware · 4 · 4 since 2021Software engineering, systems software and programming languages · 3 · 3 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1Theory of computation · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Deep-Learning-Enabled Fast Prediction of Desired Amplitude-Phase Responses in Massively Reconfigurable RF Phase Shifter
Zhilong Lu, Qincheng Qin, Kin-Fai Tong, Leon Wong, Baiyang Liu
ICIC (7)5
2026 Formal Modelling and Analysis of the O-RAN O2 Interface in Alloy: Implications for NTN Deployment
Sean McLaren, Tsutomu Kobayashi, Leon Wong, Paul Harvey 0002
ABZ3
2026 Mosaic: Composite projection pruning for resource-efficient LLMs
abstract
Extensive compute and memory requirements limit the deployment of large language models (LLMs) on any hardware. Compression methods, such as pruning, can reduce model size, which in turn reduces resource requirements. State-of-the-art pruning is based on coarse-grained methods. They are time-consuming and inherently remove critical model parameters, adversely impacting the quality of the pruned model. This paper introduces projection pruning, a novel fine-grained method for pruning LLMs. In addition, LLM projection pruning is enhanced by a new approach we refer to as composite projection pruning — the synergistic combination of unstructured pruning that retains accuracy and structured pruning that reduces model size. We develop Mosaic , a novel system to create and deploy pruned LLMs using composite projection pruning. Mosaic is evaluated using a range of performance and quality metrics on multiple hardware platforms, LLMs, and datasets. Mosaic is 7.19 × faster in producing models than existing approaches. Mosaic models achieve up to 84.2% lower perplexity and 31.4% higher accuracy than models obtained from coarse-grained pruning. Up to 67% faster inference and 68% lower GPU memory use is noted for Mosaic models. • LLMs are difficult to optimize without hardware and software accelerators. • Pruning makes LLMs smaller and faster for Edge devices. • Existing pruning methods rely on accelerators and reduce LLM quality. • Mosaic uses composite pruning to make LLMs resource-efficient without accelerators. • Mosaic has better accuracy and perplexity than existing pruning methods.
Bailey J. Eccles, Leon Wong, Blesson Varghese
Future Gener. Comput. Syst.2
2026 FedFreeze: A dual-phase layer freezing framework for federated learning
abstract
Running Federated Learning (FL) on resource-constrained devices is challenging due to the resources required for training. Layer freezing has been proposed to reduce the computational costs and thus accelerate training. However, we identify that existing layer freezing approaches either learn quickly or learn effectively, but do not balance them. Specifically, aggressive early-stage layer freezing (e.g., AutoFreeze) accelerates training but achieves a lower final accuracy. On the other hand, accuracy-guaranteed layer freezing (e.g., ALF) obtains higher final accuracy but with marginal training time improvement. This article proposes FedFreeze – a dual-phase layer freezing federated learning framework that, for the first time, combines early-stage and accuracy-guaranteed layer freezing into a unified mechanism. FedFreeze designs a novel regularization-based layer freezing strategy on the device to apply early-stage layer freezing even during the initial stages for improving training speedup. In addition, FedFreeze develops a convergence-based layer freezing strategy to achieve a high final accuracy. Experimental results show that the proposed FedFreeze framework achieves up to 1.3 × training speedup while limiting the accuracy drop to no more than 1.67% compared to vanilla FL. In contrast to state-of-the-art early-stage and accuracy-guaranteed layer freezing methods, FedFreeze consistently strikes a better balance between efficiency and accuracy across a wide range of settings, including different hardware platforms (Raspberry Pi and Jetson Nano clusters), datasets (FMNIST, CIFAR-10, CIFAR-100), model architectures (AlexNet, VGG11, ResNet12), and initialization strategies. The results demonstrate that FedFreeze outperforms state-of-the-art layer freezing techniques in accelerating FL training on resource-constrained devices, without incurring significant accuracy loss.
Di Wu 0065, Leon Wong, Blesson Varghese
Future Gener. Comput. Syst.2
2026 FedOptima: Optimizing resource utilization in federated learning
Leon Wong, Blesson Varghese
Future Gener. Comput. Syst.2
2026 scProGraph: A Cell Bagging Strategy for Cell Type Annotation With Gene Interaction-Aware Explainability
Yue-Chao Li, Hai-Ru You, Xuequn Shang 0001, Leon Wong, Zhi-an Huang, Zhu-Hong You
IEEE Trans. Big Data5
2025 Comparative Analysis Between Decentralized and Centralized Network Digital Twins of Kubernetes Clusters
abstract
In the realm of cluster operation, continuously validating and optimizing the configuration requires access to accurate cluster behavioral models. Network Digital Twins (NDTs) have emerged as a paradigm to provide such accurate, live representations of network systems. To capture the live state, NDTs need to anticipate the cluster behavior in a faster than real-time manner. With increasingly complex clusters, classical NDTs relying on detailed handcrafted simulators become too slow to fulfill this task. Leveraging measurements from the actual system demonstrates the potential to create more highlevel, lightweight NDTs that are still fairly accurate. Nonetheless, the degree of abstraction required to create fast and accurate data-driven NDTs is not well understood. To address this, our work investigates the impact of different abstraction levels on modeling accuracy. We develop and compare three Network Digital Twins of a Kubernetes Cluster - a Twin based on a Handcrafted Simulator, a Decentralized Data-driven Twin, abstracting individual system components, and a Centralized Data-driven Twin, abstracting the system as a whole. Our results show that Data-driven Twins improve the performance prediction by 18-53% over the handcrafted one, with the Centralized Twin surpassing the Decentralized Twin in accuracy by 35% and speed by two orders of magnitude.
Razvan-Mihai Ursu, Navidreza Asadi, Johannes Zerwas, Leon Wong, Wolfgang Kellerer
NetSoft4
2024 Rapid Deployment of DNNs for Edge Computing via Structured Pruning at Initialization
abstract
Edge machine learning (ML) enables localized processing of data on devices and is underpinned by deep neural networks (DNNs). However, DNNs cannot be easily run on devices due to their substantial computing, memory and energy requirements for delivering performance that is comparable to cloud-based ML. Therefore, model compression techniques, such as pruning, have been considered. Existing pruning methods are problematic for edge ML since they: (1) Create compressed models that have limited runtime performance benefits (using unstructured pruning) or compromise the final model accuracy (using structured pruning), and (2) Require substantial compute resources and time for identifying a suitable compressed DNN model (using neural architecture search). In this paper, we explore a new avenue, referred to as Pruning-at-Initialization (PaI), using structured pruning to mitigate the above problems. We develop Reconvene, a system for rapidly generating pruned models suited for edge deployments using structured PaI. Reconvene systematically identifies and prunes DNN convolution layers that are least sensitive to structured pruning. Reconvene rapidly creates pruned DNNs within seconds that are up to 16.21× smaller and 2× faster while maintaining the same accuracy as an unstructured PaI counterpart.
Bailey J. Eccles, Leon Wong, Blesson Varghese
CCGrid2
2024 MNESEDA: A prior-guided subgraph representation learning framework for predicting disease-related enhancers
Jinsheng Xu, Weicheng Sun, Weihan Zhang, Yongbin Zeng, Leon Wong, Ping Zhang 0027
Knowl. Based Syst.7
2024 AMDECDA: Attention Mechanism Combined With Data Ensemble Strategy for Predicting CircRNA-Disease Association
abstract
Accumulating evidence from recent research reveals that circRNA is tightly bound to human complex disease and plays an important regulatory role in disease progression. Identifying disease-associated circRNA occupies a key role in the research of disease pathogenesis. In this study, we propose a new model AMDECDA for predicting circRNA-disease association (CDA) by combining attention mechanism and data ensemble strategy. Firstly, we fuse the heterogeneous information including circRNA Gaussian interaction profile (GIP), disease semantics and disease GIP, and then use the attention mechanism of Graph Attention Network (GAT) to focus on the critical information of data, reasonably allocate resources and extract their essential features. Finally, the ensemble deep RVFL network (edRVFL) is utilized to quickly and accurately predict CDA in the non-iterative manner of closed-form solutions. In the five-fold cross-validation experiment on the benchmark data set, AMDECDA achieves an accuracy of 93.10% with a sensitivity of 97.56% in 0.9235 AUC. In comparison with previous models, AMDECDA exhibits highly competitiveness. Furthermore, 26 of the top 30 unknown CDAs of AMDECDA predicted scores are proved by the related literature. These results indicate that AMDECDA can effectively anticipate latent CDA and provide help for further biological wet experiments.
Lei Wang 0121, Leon Wong, Zhu-Hong You, De-Shuang Huang
IEEE Trans. Big Data2
2024 GSLCDA: An Unsupervised Deep Graph Structure Learning Method for Predicting CircRNA-Disease Association
abstract
Growing studies reveal that Circular RNAs (circRNAs) are broadly engaged in physiological processes of cell proliferation, differentiation, aging, apoptosis, and are closely associated with the pathogenesis of numerous diseases. Clarification of the correlation among diseases and circRNAs is of great clinical importance to provide new therapeutic strategies for complex diseases. However, previous circRNA-disease association prediction methods rely excessively on the graph network, and the model performance is dramatically reduced when noisy connections occur in the graph structure. To address this problem, this paper proposes an unsupervised deep graph structure learning method GSLCDA to predict potential CDAs. Concretely, we first integrate circRNA and disease multi-source data to constitute the CDA heterogeneous network. Then the network topology is learned using the graph structure, and the original graph is enhanced in an unsupervised manner by maximize the inter information of the learned and original graphs to uncover their essential features. Finally, graph space sensitive k-nearest neighbor (KNN) algorithm is employed to search for latent CDAs. In the benchmark dataset, GSLCDA obtained 92.67% accuracy with 0.9279 AUC. GSLCDA also exhibits exceptional performance on independent datasets. Furthermore, 14, 12 and 14 of the top 16 circRNAs with the most points GSLCDA prediction scores were confirmed in the relevant literature in the breast cancer, colorectal cancer and lung cancer case studies, respectively. Such results demonstrated that GSLCDA can validly reveal underlying CDA and offer new perspectives for the diagnosis and therapy of complex human diseases.
Lei Wang 0121, Zhengwei Li 0001, Zhu-Hong You, De-Shuang Huang, Leon Wong
IEEE J. Biomed. Health Informatics5
2024 MAGCDA: A Multi-Hop Attention Graph Neural Networks Method for CircRNA-Disease Association Prediction
abstract
With a growing body of evidence establishing circular RNAs (circRNAs) are widely exploited in eukaryotic cells and have a significant contribution in the occurrence and development of many complex human diseases. Disease-associated circRNAs can serve as clinical diagnostic biomarkers and therapeutic targets, providing novel ideas for biopharmaceutical research. However, available computation methods for predicting circRNA-disease associations (CDAs) do not sufficiently consider the contextual information of biological network nodes, making their performance limited. In this work, we propose a multi-hop attention graph neural network-based approach MAGCDA to infer potential CDAs. Specifically, we first construct a multi-source attribute heterogeneous network of circRNAs and diseases, then use a multi-hop strategy of graph nodes to deeply aggregate node context information through attention diffusion, thus enhancing topological structure information and mining data hidden features, and finally use random forest to accurately infer potential CDAs. In the four gold standard data sets, MAGCDA achieved prediction accuracy of 92.58%, 91.42%, 83.46% and 91.12%, respectively. MAGCDA has also presented prominent achievements in ablation experiments and in comparisons with other models. Additionally, 18 and 17 potential circRNAs in top 20 predicted scores for MAGCDA prediction scores were confirmed in case studies of the complex diseases breast cancer and Almozheimer's disease, respectively. These results suggest that MAGCDA can be a practical tool to explore potential disease-associated circRNAs and provide a theoretical basis for disease diagnosis and treatment.
Lei Wang 0121, Zhengwei Li 0001, Zhu-Hong You, De-Shuang Huang, Leon Wong
IEEE J. Biomed. Health Informatics5
2023 Towards Digital Network Twins: Can we Machine Learn Network Function Behaviors?
abstract
Cluster orchestrators such as Kubernetes (K8s) provide many knobs that cloud administrators can tune to conFigure their system. However, different configurations lead to different levels of performance, which additionally depend on the application. Hence, finding exactly the best configuration for a given system can be a difficult task. A particularly innovative approach to evaluate configurations and optimize desired performance metrics is the use of Digital Twins (DT). To achieve good results in short time, the models of the cloud network functions underlying the DT must be minimally complex but highly accurate. Developing such models requires detailed knowledge about the system components and their interactions. We believe that a data-driven paradigm can capture the actual behavior of a network function (NF) deployed in the cluster, while decoupling it from internal feedback loops. In this paper, we analyze the HTTP load balancing function as an example of an NF and explore the data-driven paradigm to learn its behavior in a K8s cluster deployment. We develop, implement, and evaluate two approaches to learn the behavior of a state-of-the-art load balancer and show that Machine Learning has the potential to enhance the way we model NF behaviors.
Razvan-Mihai Ursu, Johannes Zerwas, Patrick Krämer, Navidreza Asadi, Phil Rodgers, Leon Wong, Wolfgang Kellerer
NetSoft6
2023 Enabling Auditable Trust in Autonomous Networks with Ethereum and IPFS
abstract
Operation and management of telecommunication networks are increasingly difficult with the demands and behaviors of users exceeding the capacity of network engineers to keep pace. This has led to increased automation of the network, enabled by various forms of intelligent software. One such proposal from the ITU-T Focus Group on Autonomous Networks (standardization group) is an architecture to achieve self-driven automation (i.e. autonomy) of network operation, whereby technology from different operators and third parties is self-assembled and deployed in production networks. This raises questions and challenges regarding transparency, auditability, and trust while maintaining interoperability.This work presents an initial study of a distributed and decentralized marketplace to bring transparent and auditable trust to the proposed architecture without sacrificing interoperable functionality. We demonstrated this by our proof of concept implementation of both the proposed architecture and marketplace based on the combination of Ethereum and IPFS.
Jaime Fúster De La Fuente, Álvaro Pendás Recondo, Leon Wong, Paul Harvey 0002
NOMS3
2023 GKLOMLI: a link prediction model for inferring miRNA-lncRNA interactions by using Gaussian kernel-based method on network profile and linear optimization algorithm
abstract
BACKGROUND: The limited knowledge of miRNA-lncRNA interactions is considered as an obstruction of revealing the regulatory mechanism. Accumulating evidence on Human diseases indicates that the modulation of gene expression has a great relationship with the interactions between miRNAs and lncRNAs. However, such interaction validation via crosslinking-immunoprecipitation and high-throughput sequencing (CLIP-seq) experiments that inevitably costs too much money and time but with unsatisfactory results. Therefore, more and more computational prediction tools have been developed to offer many reliable candidates for a better design of further bio-experiments. METHODS: In this work, we proposed a novel link prediction model based on Gaussian kernel-based method and linear optimization algorithm for inferring miRNA-lncRNA interactions (GKLOMLI). Given an observed miRNA-lncRNA interaction network, the Gaussian kernel-based method was employed to output two similarity matrixes of miRNAs and lncRNAs. Based on the integrated matrix combined with similarity matrixes and the observed interaction network, a linear optimization-based link prediction model was trained for inferring miRNA-lncRNA interactions. RESULTS: To evaluate the performance of our proposed method, k-fold cross-validation (CV) and leave-one-out CV were implemented, in which each CV experiment was carried out 100 times on a training set generated randomly. The high area under the curves (AUCs) at 0.8623 ± 0.0027 (2-fold CV), 0.9053 ± 0.0017 (5-fold CV), 0.9151 ± 0.0013 (10-fold CV), and 0.9236 (LOO-CV), illustrated the precision and reliability of our proposed method. CONCLUSION: GKLOMLI with high performance is anticipated to be used to reveal underlying interactions between miRNA and their target lncRNAs, and deciphers the potential mechanisms of the complex diseases.
Leon Wong, Lei Wang 0121, Zhu-Hong You, Chang-an Yuan 0001, Mei-Yuan Cao
BMC Bioinform.1
2023 Combining K Nearest Neighbor With Nonnegative Matrix Factorization for Predicting Circrna-Disease Associations
abstract
Accumulating evidences show that circular RNAs (circRNAs) play an important role in regulating gene expression, and involve in many complex human diseases. Identifying associations of circRNA with disease helps to understand the pathogenesis, treatment and diagnosis of complex diseases. Since inferring circRNA-disease associations by biological experiments is costly and time-consuming, there is an urgently need to develop a computational model to identify the association between them. In this paper, we proposed a novel method named KNN-NMF, which combines K nearest neighbors with nonnegative matrix factorization to infer associations between circRNA and disease (KNN-NMF). Frist, we compute the Gaussian Interaction Profile (GIP) kernel similarity of circRNA and disease, the semantic similarity of disease, respectively. Then, the circRNA-disease new interaction profiles are established using weight K nearest neighbors to reduce the false negative association impact on prediction performance. Finally, Nonnegative Matrix Factorization is implemented to predict associations of circRNA with disease. The experiment results indicate that the prediction performance of KNN-NMF outperforms the competing methods under five-fold cross-validation. Moreover, case studies of two common diseases further show that KNN-NMF can identify potential circRNA-disease associations effectively.
Meineng Wang, Xue-Jun Xie, Zhu-Hong You, Leon Wong, Liping Li 0003
IEEE ACM Trans. Comput. Biol. Bioinform.4
2023 PPAEDTI: Personalized Propagation Auto-Encoder Model for Predicting Drug-Target Interactions
abstract
Identifying protein targets for drugs establishes an indispensable knowledge foundation for drug repurposing and drug development. Though expensive and time-consuming, vitro trials are widely employed to discover drug targets, and the existing relevant computational algorithms still cannot satisfy the demand for real application in drug R&D with regards to the prediction accuracy and performance efficiency, which are urgently needed to be improved. To this end, we propose here the PPAEDTI model, which uses the graph personalized propagation technique to predict drug-target interactions from the known interaction network. To evaluate the prediction performance, six benchmark datasets were used for testing with some state-of-the-art methods compared. As a result, using the 5-fold cross-validation, the proposed PPAEDTI model achieves average AUCs>90% on 5 collected datasets. We also manually checked the top-20 prediction list for 2 proteins (hsa:775 and hsa:779) and a kind of drug (D00618), and successfully confirmed 18, 17, and 20 items from the public datasets, respectively. The experimental results indicate that, given known drug-target interactions, the PPAEDTI model can provide accurate predictions for the new ones, which is anticipated to serve as a useful tool for pharmacology research. Using the proposed model that was trained with the collected datasets, we have built a computational platform that is accessible at http://120.77.11.78/PPAEDTI/ and corresponding codes and datasets are also released.
Yue-Chao Li, Zhu-Hong You, Lei Wang 0121, Leon Wong, Lun Hu, Pengwei Hu 0001
IEEE J. Biomed. Health Informatics5
2023 Biomedical Knowledge Graph Embedding With Capsule Network for Multi-Label Drug-Drug Interaction Prediction
abstract
Drug-drug interaction (DDI) plays an important role in drug development and administration. Most of existing network-based computation models regard the DDI prediction as a binary classification problem and generate negative DDI samples randomly, but the binary classification is not in line with the real problem since there are dozens of types of DDI and randomly generating negative samples may introduce false-negative samples since the non-observed facts can be either false or just missing. To address the above limitations, we propose a new framework called KG2ECapsule that explicitly models the multi-relational DDI data based on biomedical knowledge graphs in an end-to-end fashion. It first generates high-quality negative samples based on the average number of tail entities and head entities for each relation to reduce false-negative samples to some extent. KG2ECapsule then refines the representations of entities by recursively propagating the embeddings from the attention-based receptive fields of entities. Empirical results on three biomedical knowledge graphs of different scales show that KG2ECapsule outperforms the state-of-the-art methods consistently in multi-label DDI prediction task and further studies verify the efficacy of both probability-based sampling strategy and non-linear transformation for modeling multi-relational data.
Xiao-Rui Su 0001, Zhu-Hong You, De-Shuang Huang, Lei Wang 0121, Leon Wong, Bo-Wei Zhao
IEEE Trans. Knowl. Data Eng.5
2022 A machine learning framework based on multi-source feature fusion for circRNA-disease association prediction
abstract
Circular RNAs (circRNAs) are involved in the regulatory mechanisms of multiple complex diseases, and the identification of their associations is critical to the diagnosis and treatment of diseases. In recent years, many computational methods have been designed to predict circRNA-disease associations. However, most of the existing methods rely on single correlation data. Here, we propose a machine learning framework for circRNA-disease association prediction, called MLCDA, which effectively fuses multiple sources of heterogeneous information including circRNA sequences and disease ontology. Comprehensive evaluation in the gold standard dataset showed that MLCDA can successfully capture the complex relationships between circRNAs and diseases and accurately predict their potential associations. In addition, the results of case studies on real data show that MLCDA significantly outperforms other existing methods. MLCDA can serve as a useful tool for circRNA-disease association prediction, providing mechanistic insights for disease research and thus facilitating the progress of disease treatment.
Lei Wang 0121, Leon Wong, Zhengwei Li 0001, Xiao-Rui Su 0001, Bo-Wei Zhao, Zhu-Hong You
Briefings Bioinform.2
2022 NSECDA: Natural Semantic Enhancement for CircRNA-Disease Association Prediction
abstract
Increasing evidence suggest that circRNA, as one of the most promising emerging biomarkers, has a very close relationship with diseases. Exploring the relationship between circRNA and diseases can provide novel perspective for diseases diagnosis and pathogenesis. The existing circRNA-disease association (CDA) prediction models, however, generally treat the data attributes equally, do not pay special attention to the attributes with more significant influence, and do not make full use of the correlation and symbiosis between attributes to dig into the latent semantic information of the data. Therefore, in response to the above problems, this paper proposes a natural semantic enhancement method NSECDA to predict CDA. In practical terms, we first recognize the circRNA sequence as a biological language, and analyze its natural semantic properties through the natural language understanding theory; then integrate it with disease attributes, circRNA and disease Gaussian Interaction Profile (GIP) kernel attributes, and use Graph Attention Network (GAT) to focus on the influential attributes, so as to mine the deeply hidden features; finally, the Rotation Forest (RoF) classifier was used to accurately determine CDA. In the gold standard data set CircR2Disease, NSECDA achieved 92.49% accuracy with 0.9225 AUC score. In comparison with the non-natural semantic enhancement model and other classifier models, NSECDA also shows competitive performance. Additionally, 25 of the CDA pairs with unknown associations in the top 30 prediction scores of NSECDA have been proven by newly reported studies. These achievements suggest that NSECDA is an effective model to predict CDA, which can provide credible candidate for subsequent wet experiments, thus significantly reducing the scope of investigations.
Lei Wang 0121, Leon Wong, Zhu-Hong You, De-Shuang Huang, Xiao-Rui Su 0001, Bo-Wei Zhao
IEEE J. Biomed. Health Informatics2
2021 Predicting miRNA-Disease Associations via a New MeSH Headings Representation of Diseases and eXtreme Gradient Boosting
Zhu-Hong You, Lei Wang 0121, Leon Wong, Xiao-Rui Su 0001, Bo-Wei Zhao
ICIC (3)4
2021 Weighted Nonnegative Matrix Factorization Based on Multi-source Fusion Information for Predicting CircRNA-Disease Associations
Meineng Wang, Xue-Jun Xie, Zhu-Hong You, Leon Wong, Liping Li 0003
ICIC (3)4
2021 A Multi-graph Deep Learning Model for Predicting Drug-Disease Associations
Bo-Wei Zhao, Zhu-Hong You, Lun Hu, Leon Wong, Ping Zhang 0027
ICIC (3)4
2020 A Highly Efficient Biomolecular Network Representation Model for Predicting Drug-Disease Associations
Hanjing Jiang, Zhu-Hong You, Lun Hu, Zhen-Hao Guo, Leon Wong
ICIC (3)6
2020 A Gaussian Kernel Similarity-Based Linear Optimization Model for Predicting miRNA-lncRNA Interactions
Leon Wong, Zhu-Hong You, Xi Zhou 0007, Mei-Yuan Cao
ICIC (2)1
2020 A Novel Computational Method for Predicting LncRNA-Disease Associations from Heterogeneous Information Network with SDNE Embedding Model
Ping Zhang 0027, Bo-Wei Zhao, Leon Wong, Zhu-Hong You, Zhen-Hao Guo
ICIC (2)3
2020 Inferring Disease-Associated Piwi-Interacting RNAs via Graph Attention Networks
Kai Zheng 0020, Zhu-Hong You, Lei Wang 0121, Leon Wong
ICIC (2)4
2020 NEMPD: a network embedding-based method for predicting miRNA-disease associations by preserving behavior and attribute information
abstract
BACKGROUND: As an important non-coding RNA, microRNA (miRNA) plays a significant role in a series of life processes and is closely associated with a variety of Human diseases. Hence, identification of potential miRNA-disease associations can make great contributions to the research and treatment of Human diseases. However, to our knowledge, many existing computational methods only utilize the single type of known association information between miRNAs and diseases to predict their potential associations, without focusing on their interactions or associations with other types of molecules. RESULTS: In this paper, we propose a network embedding-based method for predicting miRNA-disease associations by preserving behavior and attribute information. Firstly, a heterogeneous network is constructed by integrating known associations among miRNA, protein and disease, and the network representation method Learning Graph Representations with Global Structural Information (GraRep) is implemented to learn the behavior information of miRNAs and diseases in the network. Then, the behavior information of miRNAs and diseases is combined with the attribute information of them to represent miRNA-disease association pairs. Finally, the prediction model is established based on the Random Forest algorithm. Under the five-fold cross validation, the proposed NEMPD model obtained average 85.41% prediction accuracy with 80.96% sensitivity at the AUC of 91.58%. Furthermore, the performance of NEMPD is also validated by the case studies. Among the top 50 predicted disease-related miRNAs, 48 (breast neoplasms), 47 (colon neoplasms), 47 (lung neoplasms) were confirmed by two other databases. CONCLUSIONS: The proposed NEMPD model has a good performance in predicting the potential associations between miRNAs and diseases, and has great potency in the field of miRNA-disease association prediction in the future.
Zhu-Hong You, Leon Wong
BMC Bioinform.4
2019 A Gated Recurrent Unit Model for Drug Repositioning by Combining Comprehensive Similarity Measures and Gaussian Interaction Profile Kernel
Zhu-Hong You, Liping Li 0003, Lun Hu, Leon Wong
ICIC (2)7
2015 Predicting Protein-Protein Interactions from Amino Acid Sequences Using SaE-ELM Combined with Continuous Wavelet Descriptor and PseAA Composition
Zhu-Hong You, Jianqiang Li 0001, Leon Wong, Shubin Cai
ICIC (2)4
2015 Detection of Protein-Protein Interactions from Amino Acid Sequences Using a Rotation Forest Model with a Novel PR-LPQ Descriptor
Leon Wong, Zhu-Hong You, Shuai Li 0002
ICIC (3)1
2006 High accuracy retrieval with multiple nested ranker
abstract
High precision at the top ranks has become a new focus of research in information retrieval. This paper presents the multiple nested ranker approach that improves the accuracy at the top ranks by iteratively re-ranking the top scoring documents. At each iteration, this approach uses the RankNet learning algorithm to re-rank a subset of the results. This splits the problem into smaller and easier tasks and generates a new distribution of the results to be learned by the algorithm. We evaluate this approach using different settings on a data set labeled with several degrees of relevance. We use the normalized discounted cumulative gain (NDCG) to measure the performance because it depends not only on the position but also on the relevance score of the document in the ranked list. Our experiments show that making the learning algorithm concentrate on the top scoring results improves precision at the top ten documents in terms of the NDCG score.
Irina Matveeva, Christopher J. C. Burges, Timo Burkard, Andy Laucius, Leon Wong
SIGIR5
2002 Combination of statistical and rule-based approaches for spoken language understanding
abstract
A Natural User Interface (NUI), where a user can type or speak a request, is a good complement to the well-known Graphical User Interface (GUI). Accurately extracting user intent from such typed or spoken queries is a very difficult challenge. In this paper we evaluate several techniques to extract user intent from typed sentences in the context of the well-known Airline Travel Information (ATIS) domain, where we want to extract which of the possible tasks the user wants to do and the value of the slots associated to that task. In previous work we showed that a Semantic Context Free Grammar (CFG) semi-automatically derived from labeled data can offer very good results. In this paper we evaluate several statistical pattern recognition techniques including Support Vector Machines (SVM), Naïve Bayes classifiers and task-dependent n-gram language models. These methods can yield a very low task classification error rate. If used in combination with our CFG system, they can also lead to very low slot error rates. 1.
Ye-Yi Wang, Alex Acero, Ciprian Chelba, Brendan J. Frey, Leon Wong
INTERSPEECH5