Man Hon Wong 0001

dblp:w/ManHonWong · DBLP profile ↗
← Back
66ranked-venue papers
8as first author
7since 2021 · last 2025
0009-0000-8553-6004ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 27 · 6 first-authorApplied, interdisciplinary, general and emerging computing · 22 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 12Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 since 2021Human-computer interaction and ubiquitous computing · 3Systems, architecture and hardware · 2Computer networks · 1Security and privacy · 1Theory of computation · 1 · 1 first-author
YearPublicationVenuePosition
2025 MESH - Understanding Videos Like Human: Measuring Hallucinations in Large Video Models
abstract
Large Video Models (LVMs) build on the semantic capabilities of Large Language Models (LLMs) and vision modules by integrating temporal information to better understand dynamic video content. Despite their progress, LVMs are prone to hallucinations-producing inaccurate or irrelevant descriptions. Current benchmarks for video hallucination depend heavily on manual categorization of video content, neglecting the perception-based processes through which humans naturally interpret videos. We introduce MESH, a benchmark designed to evaluate hallucinations in LVMs systematically. MESH uses a Question-Answering framework with binary and multi-choice formats incorporating target and trap instances. It follows a bottom-up approach, evaluating basic objects, coarse-to-fine subject features, and subject-action pairs, aligning with human video understanding. We demonstrate that MESH offers an effective and comprehensive approach for identifying hallucinations in video understanding. Our evaluations show that while LVMs excel at recognizing basic objects and features, their susceptibility to hallucinations increases markedly when handling fine details or aligning multiple actions involving various subjects in longer videos. The benchmark is available at MESH-Benchmark.
Garry Yang, Zizhe Chen, Man Hon Wong 0001, Haoyu Lei, Yongqiang Chen 0002, Zhenguo Li, Kaiwen Zhou 0001, James Cheng
ACM Multimedia3
2024 scCaT: An explainable capsulating architecture for sepsis diagnosis transferring from single-cell RNA sequencing
abstract
Sepsis is a life-threatening condition characterized by an exaggerated immune response to pathogens, leading to organ damage and high mortality rates in the intensive care unit. Although deep learning has achieved impressive performance on prediction and classification tasks in medicine, it requires large amounts of data and lacks explainability, which hinder its application to sepsis diagnosis. We introduce a deep learning framework, called scCaT, which blends the capsulating architecture with Transformer to develop a sepsis diagnostic model using single-cell RNA sequencing data and transfers it to bulk RNA data. The capsulating architecture effectively groups genes into capsules based on biological functions, which provides explainability in encoding gene expressions. The Transformer serves as a decoder to classify sepsis patients and controls. Our model achieves high accuracy with an AUROC of 0.93 on the single-cell test set and an average AUROC of 0.98 on seven bulk RNA cohorts. Additionally, the capsules can recognize different cell types and distinguish sepsis from control samples based on their biological pathways. This study presents a novel approach for learning gene modules and transferring the model to other data types, offering potential benefits in diagnosing rare diseases with limited subjects.
Xubin Zheng, Dian Meng, Wan-Ki Wong, Ka-Ho To, Lei Zhu 0016, Jiafei Wu, Yining Liang, Kwong-Sak Leung, Man Hon Wong 0001, Lixin Cheng
PLoS Comput. Biol.10
2023 bvnGPS: a generalizable diagnostic model for acute bacterial and viral infection using integrative host transcriptomics and pretrained neural networks
abstract
MOTIVATION: The confusion of acute inflammation infected by virus and bacteria or noninfectious inflammation will lead to missing the best therapy occasion resulting in poor prognoses. The diagnostic model based on host gene expression has been widely used to diagnose acute infections, but the clinical usage was hindered by the capability across different samples and cohorts due to the small sample size for signature training and discovery. RESULTS: Here, we construct a large-scale dataset integrating multiple host transcriptomic data and analyze it using a sophisticated strategy which removes batch effect and extracts the common information from different cohorts based on the relative expression alteration of gene pairs. We assemble 2680 samples across 16 cohorts and separately build gene pair signature (GPS) for bacterial, viral, and noninfected patients. The three GPSs are further assembled into an antibiotic decision model (bacterial-viral-noninfected GPS, bvnGPS) using multiclass neural networks, which is able to determine whether a patient is bacterial infected, viral infected, or noninfected. bvnGPS can distinguish bacterial infection with area under the receiver operating characteristic curve (AUC) of 0.953 (95% confidence interval, 0.948-0.958) and viral infection with AUC of 0.956 (0.951-0.961) in the test set (N = 760). In the validation set (N = 147), bvnGPS also shows strong performance by attaining an AUC of 0.988 (0.978-0.998) on bacterial-versus-other and an AUC of 0.994 (0.984-1.000) on viral-versus-other. bvnGPS has the potential to be used in clinical practice and the proposed procedure provides insight into data integration, feature selection and multiclass classification for host transcriptomics data. AVAILABILITY AND IMPLEMENTATION: The codes implementing bvnGPS are available at https://github.com/Ritchiegit/bvnGPS. The construction of iPAGE algorithm and the training of neural network was conducted on Python 3.7 with Scikit-learn 0.24.1 and PyTorch 1.7. The visualization of the results was implemented on R 4.2, Python 3.7, and Matplotlib 3.3.4.
Qizhi Li, Xubin Zheng, Jize Xie, Man Hon Wong 0001, Kwong-Sak Leung, Shuai Li 0010, Qingshan Geng, Lixin Cheng
Bioinform.6
2023 Deciphering associations between gut microbiota and clinical factors using microbial modules
abstract
MOTIVATION: Human gut microbiota plays a vital role in maintaining body health. The dysbiosis of gut microbiota is associated with a variety of diseases. It is critical to uncover the associations between gut microbiota and disease states as well as other intrinsic or environmental factors. However, inferring alterations of individual microbial taxa based on relative abundance data likely leads to false associations and conflicting discoveries in different studies. Moreover, the effects of underlying factors and microbe-microbe interactions could lead to the alteration of larger sets of taxa. It might be more robust to investigate gut microbiota using groups of related taxa instead of the composition of individual taxa. RESULTS: We proposed a novel method to identify underlying microbial modules, i.e. groups of taxa with similar abundance patterns affected by a common latent factor, from longitudinal gut microbiota and applied it to inflammatory bowel disease (IBD). The identified modules demonstrated closer intragroup relationships, indicating potential microbe-microbe interactions and influences of underlying factors. Associations between the modules and several clinical factors were investigated, especially disease states. The IBD-associated modules performed better in stratifying the subjects compared with the relative abundance of individual taxa. The modules were further validated in external cohorts, demonstrating the efficacy of the proposed method in identifying general and robust microbial modules. The study reveals the benefit of considering the ecological effects in gut microbiota analysis and the great promise of linking clinical factors with underlying microbial modules. AVAILABILITY AND IMPLEMENTATION: https://github.com/rwang-z/microbial_module.git.
Xubin Zheng, Fangda Song, Man Hon Wong 0001, Kwong-Sak Leung, Lixin Cheng
Bioinform.4
2022 Improving bulk RNA-seq classification by transferring gene signature from single cells in acute myeloid leukemia
abstract
The advances in single-cell RNA sequencing (scRNA-seq) technologies enable the characterization of transcriptomic profiles at the cellular level and demonstrate great promise in bulk sample analysis thereby offering opportunities to transfer gene signature from scRNA-seq to bulk data. However, the gene expression signatures identified from single cells are typically inapplicable to bulk RNA-seq data due to the profiling differences of distinct sequencing technologies. Here, we propose single-cell pair-wise gene expression (scPAGE), a novel method to develop single-cell gene pair signatures (scGPSs) that were beneficial to bulk RNA-seq classification to transfer knowledge across platforms. PAGE was adopted to tackle the challenge of profiling differences. We applied the method to acute myeloid leukemia (AML) and identified the scGPS from mouse scRNA-seq that allowed discriminating between AML and control cells. The scGPS was validated in bulk RNA-seq datasets and demonstrated better performance (average area under the curve [AUC] = 0.96) than the conventional gene expression strategies (average AUC$\le$ 0.88) suggesting its potential in disclosing the molecular mechanism of AML. The scGPS also outperformed its bulk counterpart, which highlighted the benefit of gene signature transfer. Furthermore, we confirmed the utility of scPAGE in sepsis as an example of other disease scenarios. scPAGE leveraged the advantages of single-cell profiles to enhance the analysis of bulk samples revealing great potential of transferring knowledge from single-cell to bulk transcriptome studies.
Xubin Zheng, Shibiao Wan, Fangda Song, Man Hon Wong 0001, Kwong-Sak Leung, Lixin Cheng
Briefings Bioinform.6
2022 meGPS: a multi-omics signature for hepatocellular carcinoma detection integrating methylome and transcriptome data
abstract
MOTIVATION: Hepatocellular carcinoma (HCC) is a primary malignancy with a poor prognosis. Recently, multi-omics molecular-level measurement enables HCC diagnosis and prognosis prediction, which is crucial for early intervention of personalized therapy to diminish mortality. Here, we introduce a novel strategy utilizing DNA methylation and RNA expression data to achieve a multi-omics gene pair signature (GPS) for HCC discrimination. RESULTS: The immune genes with negative correlations between expression and promoter methylation are enriched in the highly connected cancer-related pathway network, which are considered as the candidates for HCC detection. After that, we separately construct a methylation GPS (mGPS) and an expression GPS (eGPS), and then assemble them as a meGPS with five gene pairs, in which the significant methylation and expression changes occur between HCC tumor and non-tumor groups. Reliable performance has been validated by independent tissue (age, gender and etiology) and blood datasets. This study proposes a procedure for multi-omics GPS identification and develops a novel HCC signature using both methylome and transcriptome data, suggesting potential molecular targets for the detection and therapy of HCC. AVAILABILITY AND IMPLEMENTATION: Models are available at https://github.com/bioinformaticStudy/meGPS.git. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Xubin Zheng, Kwong-Sak Leung, Man Hon Wong 0001, Stephen Kwok-Wing Tsui, Lixin Cheng
Bioinform.4
2022 A Robust and Generalizable Immune-Related Signature for Sepsis Diagnostics
abstract
High-throughput sequencing can detect tens of thousands of genes in parallel, providing opportunities for improving the diagnostic accuracy of multiple diseases including sepsis, which is an aggressive inflammatory response to infection that can cause organ failure and death. Early screening of sepsis is essential in clinic, but no effective diagnostic biomarkers are available yet. Here, we present a novel method, Recurrent Logistic Regression, to identify diagnostic biomarkers for sepsis from the blood transcriptome data. A panel including five immune-related genes, LRRN3, IL2RB, FCER1A, TLR5, and S100A12, are determined as diagnostic biomarkers (LIFTS) for sepsis. LIFTS discriminates patients with sepsis from normal controls in high accuracy (AUROC = 0.9959 on average; IC = [0.9722-1.0]) on nine validation cohorts across three independent platforms, which outperforms existing markers. Our analysis determined an accurate prediction model and reproducible transcriptome biomarkers that can lay a foundation for clinical diagnostic tests and biological mechanistic studies.
Yueran Yang, Yu Zhang 0151, Shuai Li 0010, Xubin Zheng, Man Hon Wong 0001, Kwong-Sak Leung, Lixin Cheng
IEEE ACM Trans. Comput. Biol. Bioinform.5
2020 Drug2vec: A Drug Embedding Method with Drug-Drug Interaction as the Context
Xubin Zheng, Man Hon Wong 0001, Kwong-Sak Leung
EANN3
2019 Classical scoring functions for docking are unable to exploit large volumes of structural and interaction data
abstract
MOTIVATION: Studies have shown that the accuracy of random forest (RF)-based scoring functions (SFs), such as RF-Score-v3, increases with more training samples, whereas that of classical SFs, such as X-Score, does not. Nevertheless, the impact of the similarity between training and test samples on this matter has not been studied in a systematic manner. It is therefore unclear how these SFs would perform when only trained on protein-ligand complexes that are highly dissimilar or highly similar to the test set. It is also unclear whether SFs based on machine learning algorithms other than RF can also improve accuracy with increasing training set size and to what extent they learn from dissimilar or similar training complexes. RESULTS: We present a systematic study to investigate how the accuracy of classical and machine-learning SFs varies with protein-ligand complex similarities between training and test sets. We considered three types of similarity metrics, based on the comparison of either protein structures, protein sequences or ligand structures. Regardless of the similarity metric, we found that incorporating a larger proportion of similar complexes to the training set did not make classical SFs more accurate. In contrast, RF-Score-v3 was able to outperform X-Score even when trained on just 32% of the most dissimilar complexes, showing that its superior performance owes considerably to learning from dissimilar training complexes to those in the test set. In addition, we generated the first SF employing Extreme Gradient Boosting (XGBoost), XGB-Score, and observed that it also improves with training set size while outperforming the rest of SFs. Given the continuous growth of training datasets, the development of machine-learning SFs has become very appealing. AVAILABILITY AND IMPLEMENTATION: https://github.com/HongjianLi/MLSF. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Jiangjun Peng, Pavel Sidorov, Yee Leung, Kwong-Sak Leung, Man Hon Wong 0001, Pedro J. Ballester
Bioinform.6
2019 Predicting associations among drugs, targets and diseases by tensor decomposition for drug repositioning
abstract
BACKGROUND: Development of new drugs is a time-consuming and costly process, and the cost is still increasing in recent years. However, the number of drugs approved by FDA every year per dollar spent on development is declining. Drug repositioning, which aims to find new use of existing drugs, attracts attention of pharmaceutical researchers due to its high efficiency. A variety of computational methods for drug repositioning have been proposed based on machine learning approaches, network-based approaches, matrix decomposition approaches, etc. RESULTS: We propose a novel computational method for drug repositioning. We construct and decompose three-dimensional tensors, which consist of the associations among drugs, targets and diseases, to derive latent factors reflecting the functional patterns of the three kinds of entities. The proposed method outperforms several baseline methods in recovering missing associations. Most of the top predictions are validated by literature search and computational docking. Latent factors are used to cluster the drugs, targets and diseases into functional groups. Topological Data Analysis (TDA) is applied to investigate the properties of the clusters. We find that the latent factors are able to capture the functional patterns and underlying molecular mechanisms of drugs, targets and diseases. In addition, we focus on repurposing drugs for cancer and discover not only new therapeutic use but also adverse effects of the drugs. In the in-depth study of associations among the clusters of drugs, targets and cancer subtypes, we find there exist strong associations between particular clusters. CONCLUSIONS: The proposed method is able to recover missing associations, discover new predictions and uncover functional clusters of drugs, targets and diseases. The clustering of drugs, targets and diseases, as well as the associations among the clusters, provides a new guiding framework for drug repositioning.
Shuai Li 0010, Lixin Cheng, Man Hon Wong 0001, Kwong-Sak Leung
BMC Bioinform.4
2018 Drug-Protein-Disease Association Prediction and Drug Repositioning Based on Tensor Decomposition
Shuai Li 0010, Man Hon Wong 0001, Kwong-Sak Leung
BIBM3
2018 An integrated web-based air pollution decision support system - a prototype
abstract
To efficiently and effectively monitor and mitigate air pollution in the urban environment, it is of paramount importance to integrate into a unified whole air pollutant concentration databases coming from different sources including the ground-based stations, mobile sensors, remote sensing, atmospheric-chemical-transport models and social media for the analysis and unraveling of the complex air pollution processes in space and time. This study constructs and implements for the first time a prototype of the fully integrated air pollution decision support system (APDSS) that put together in an integrated manner all relevant multi-scale, multi-type and multi-source data for decision-making on urban air pollution. The prototype contains the main system that handles the multi-source, multi-type and multi-scale databases, queries, visualization and data mining algorithms and the integrated modules that individually and holistically capitalize on the power of the ground-based stations, ground and aerial mobile sensors, satellite-borne remote-sensing technologies, atmospheric-chemical-transport models and social media. It renders a solid scientific foundation and system development methodology for the study of the spatiotemporal air pollution profiles crucial to the mitigation of urban air pollution. Real-life applications of the prototype are employed to illustrate the functionality of the APDSS.
Yee Leung, Kwong-Sak Leung, Man Hon Wong 0001, Terrence S. T. Mak, Kwan-Yau Cheung, Leung-Yau Lo, Wei Ying Yi, Yuan-Lin Dong
Int. J. Geogr. Inf. Sci.3
2017 Discovering Protein-DNA Binding Cores by Aligned Pattern Clustering
abstract
Understanding binding cores is of fundamental importance in deciphering Protein-DNA (TF-TFBS) binding and gene regulation. Limited by expensive experiments, it is promising to discover them with variations directly from sequence data. Although existing computational methods have produced satisfactory results, they are one-to-one mappings with no site-specific information on residue/nucleotide variations, where these variations in binding cores may impact binding specificity. This study presents a new representation for modeling binding cores by incorporating variations and an algorithm to discover them from only sequence data. Our algorithm takes protein and DNA sequences from TRANSFAC (a Protein-DNA Binding Database) as input; discovers from both sets of sequences conserved regions in Aligned Pattern Clusters (APCs); associates them as Protein-DNA Co-Occurring APCs; ranks the Protein-DNA Co-Occurring APCs according to their co-occurrence, and among the top ones, finds three-dimensional structures to support each binding core candidate. If successful, candidates are verified as binding cores. Otherwise, homology modeling is applied to their close matches in PDB to attain new chemically feasible binding cores. Our algorithm obtains binding cores with higher precision and much faster runtime ( ≥ 1,600x) than that of its contemporaries, discovering candidates that do not co-occur as one-to-one associated patterns in the raw data. AVAILABILITY: http://www.pami.uwaterloo.ca/~ealee/files/tcbbPnDna2015/Release.zip.
Annie En-Shiun Lee, Ho-Yin Sze-To, Man Hon Wong 0001, Kwong-Sak Leung, Terrence Chi-Kong Lau, Andrew K. C. Wong
IEEE ACM Trans. Comput. Biol. Bioinform.3
2016 Correcting the impact of docking pose generation error on binding affinity prediction
abstract
BACKGROUND: Pose generation error is usually quantified as the difference between the geometry of the pose generated by the docking software and that of the same molecule co-crystallised with the considered protein. Surprisingly, the impact of this error on binding affinity prediction is yet to be systematically analysed across diverse protein-ligand complexes. RESULTS: Against commonly-held views, we have found that pose generation error has generally a small impact on the accuracy of binding affinity prediction. This is also true for large pose generation errors and it is not only observed with machine-learning scoring functions, but also with classical scoring functions such as AutoDock Vina. Furthermore, we propose a procedure to correct a substantial part of this error which consists of calibrating the scoring functions with re-docked, rather than co-crystallised, poses. In this way, the relationship between Vina-generated protein-ligand poses and their binding affinities is directly learned. As a result, test set performance after this error-correcting procedure is much closer to that of predicting the binding affinity in the absence of pose generation error (i.e. on crystal structures). We evaluated several strategies, obtaining better results for those using a single docked pose per ligand than those using multiple docked poses per ligand. CONCLUSIONS: Binding affinity prediction is often carried out on the docked pose of a known binder rather than its co-crystallised pose. Our results suggest than pose generation error is in general far less damaging for binding affinity prediction than it is currently believed. Another contribution of our study is the proposal of a procedure that largely corrects for this error. The resulting machine-learning scoring function is freely available at http://istar.cse.cuhk.edu.hk/rf-score-4.tgz and http://ballester.marseille.inserm.fr/rf-score-4.tgz .
Kwong-Sak Leung, Man Hon Wong 0001, Pedro J. Ballester
BMC Bioinform.3
2015 Discovering Binding Cores in Protein-DNA Binding Using Association Rule Mining with Statistical Measures
abstract
Understanding binding cores is of fundamental importance in deciphering Protein-DNA (TF-TFBS) binding and for the deep understanding of gene regulation. Traditionally, binding cores are identified in resolved high-resolution 3D structures. However, it is expensive, labor-intensive and time-consuming to obtain these structures. Hence, it is promising to discover binding cores computationally on a large scale. Previous studies successfully applied association rule mining to discover binding cores from TF-TFBS binding sequence data only. Despite the successful results, there are limitations such as the use of tight support and confidence thresholds, the distortion by statistical bias in counting pattern occurrences, and the lack of a unified scheme to rank TF-TFBS associated patterns. In this study, we proposed an association rule mining algorithm incorporating statistical measures and ranking to address these limitations. Experimental results demonstrated that, even when the threshold on support was lowered to one-tenth of the value used in previous studies, a satisfactory verification ratio was consistently observed under different confidence levels. Moreover, we proposed a novel ranking scheme for TF-TFBS associated patterns based on p-values and co-support values. By comparing with other discovery approaches, the effectiveness of our algorithm was demonstrated. Eighty-four binding cores with PDB support are uniquely identified.
Man Hon Wong 0001, Ho-Yin Sze-To, Leung-Yau Lo, Tak-Ming Chan, Kwong-Sak Leung
IEEE ACM Trans. Comput. Biol. Bioinform.1
2014 Discovering protein-DNA binding cores by aligned pattern clustering
abstract
Understanding binding cores is of fundamental importance in deciphering Protein-DNA (TF-TFBS) binding and gene regulation. Variations (or mutations) in binding cores are ubiquitous and have different levels of effects on the binding specificity. To alleviate expensive experiments, we have developed a new method to discover directly from sequence data binding cores and study the effect due to variations. Although existing computational methods have produced satisfactory TF-TFBS binding cores, they are only one-to-one mappings with no site-specific information on residue/nucleotide variations; and also are largely overlapped. In this study, we propose a new representation for modeling TF-TFBS binding with variants known as TF-TFBS Co-Supportive Aligned Pattern Clusters (APCs), which are more compact, with more details for site-specific variants, and biologically more intuitive for analysis. To achieve this task, we have also developed an algorithm to discover TF-TFBS Co-Supportive APCs to capture binding cores at a higher precision with much faster runtime (≥1600X) comparing to other methods. The variants in TF-TFBS Co-Supportive APCs are also statistically analyzed and demonstrated that they can assist homology modeling to synthesize new biological knowledge.
Annie En-Shiun Lee, Kwong-Sak Leung, Ho-Yin Sze-To, Terrence Chi-Kong Lau, Man Hon Wong 0001, Andrew K. C. Wong
BIBM5
2014 iSyn: WebGL-Based Interactive De Novo Drug Design
abstract
We present iSyn, a WebGL-based tool for interactivede novo drug design. It features an evolutionary algorithm that automatically designs novel ligands with drug-like properties and synthetic feasibility using click chemistry. Isyn interfaces with our popular and fast molecular docking engine idock, remarkably reducing the evaluation and ranking time of drug candidates. Furthermore, inspired by our user friendly and high-performance WebGL visualizer iview, our iSyn also implements a tailor-made interactive visualizer to aid novel drug design. We believe iSyn can supplement the efforts of medicinal chemists in drug discovery research. To illustrate the utility of iSyn in generating novelligands ex nihilo, we designed predicted inhibitors of two important drug targets, which are RNA editing ligase 1(REL1) from T. Brucei, the etiological agent of African sleeping sickness, and cyclin-dependent kinase 2 (CDK2), a positive regulator of eukaryotic cell cycle progression. Results show that iSyn managed to significantly enhance the predicted binding affinity of the best generated ligand by more than 3 orders of magnitude in potency. Isyn is written in C++, Python, HTML5 and JavaScript. It is free and open source, available athttp://istar.cse.cuhk.edu.hk/iSyn.tgz. It has been tested successfully on both Linux and Windows.
Kwong-Sak Leung, Chun Ho Chan, Hei Lun Cheung, Man Hon Wong 0001
IV5
2014 iview: an interactive WebGL visualizer for protein-ligand complex
abstract
BACKGROUND: Visualization of protein-ligand complex plays an important role in elaborating protein-ligand interactions and aiding novel drug design. Most existing web visualizers either rely on slow software rendering, or lack virtual reality support. The vital feature of macromolecular surface construction is also unavailable. RESULTS: We have developed iview, an easy-to-use interactive WebGL visualizer of protein-ligand complex. It exploits hardware acceleration rather than software rendering. It features three special effects in virtual reality settings, namely anaglyph, parallax barrier and oculus rift, resulting in visually appealing identification of intermolecular interactions. It supports four surface representations including Van der Waals surface, solvent excluded surface, solvent accessible surface and molecular surface. Moreover, based on the feature-rich version of iview, we have also developed a neat and tailor-made version specifically for our istar web platform for protein-ligand docking purpose. This demonstrates the excellent portability of iview. CONCLUSIONS: Using innovative 3D techniques, we provide a user friendly visualizer that is not intended to compete with professional visualizers, but to enable easy accessibility and platform independence.
Kwong-Sak Leung, Takanori Nakane, Man Hon Wong 0001
BMC Bioinform.4
2014 Substituting random forest for multiple linear regression improves binding affinity prediction of scoring functions: Cyscore as a case study
abstract
BACKGROUND: State-of-the-art protein-ligand docking methods are generally limited by the traditionally low accuracy of their scoring functions, which are used to predict binding affinity and thus vital for discriminating between active and inactive compounds. Despite intensive research over the years, classical scoring functions have reached a plateau in their predictive performance. These assume a predetermined additive functional form for some sophisticated numerical features, and use standard multivariate linear regression (MLR) on experimental data to derive the coefficients. RESULTS: In this study we show that such a simple functional form is detrimental for the prediction performance of a scoring function, and replacing linear regression by machine learning techniques like random forest (RF) can improve prediction performance. We investigate the conditions of applying RF under various contexts and find that given sufficient training samples RF manages to comprehensively capture the non-linearity between structural features and measured binding affinities. Incorporating more structural features and training with more samples can both boost RF performance. In addition, we analyze the importance of structural features to binding affinity prediction using the RF variable importance tool. Lastly, we use Cyscore, a top performing empirical scoring function, as a baseline for comparison study. CONCLUSIONS: Machine-learning scoring functions are fundamentally different from classical scoring functions because the former circumvents the fixed functional form relating structural features with binding affinities. RF, but not MLR, can effectively exploit more structural features and more training samples, leading to higher prediction performance. The future availability of more X-ray crystal structures will further widen the performance gap between RF-based and MLR-based scoring functions. This further stresses the importance of substituting RF for MLR in scoring function development.
Kwong-Sak Leung, Man Hon Wong 0001, Pedro J. Ballester
BMC Bioinform.3
2013 Modeling Associated Protein-DNA Pattern Discovery with Unified Scores
abstract
Understanding protein-DNA interactions, specifically transcription factor (TF) and transcription factor binding site (TFBS) bindings, is crucial in deciphering gene regulation. The recent associated TF-TFBS pattern discovery combines one-sided motif discovery on both the TF and the TFBS sides. Using sequences only, it identifies the short protein-DNA binding cores available only in high-resolution 3D structures. The discovered patterns lead to promising subtype and disease analysis applications. While the related studies use either association rule mining or existing TFBS annotations, none has proposed any formal unified (both-sided) model to prioritize the top verifiable associated patterns. We propose the unified scores and develop an effective pipeline for associated TF-TFBS pattern discovery. Our stringent instance-level evaluations show that the patterns with the top unified scores match with the binding cores in 3D structures considerably better than the previous works, where up to 90 percent of the top 20 scored patterns are verified. We also introduce extended verification from literature surveys, where the high unified scores correspond to even higher verification percentage. The top scored patterns are confirmed to match the known WRKY binding cores with no available 3D structures and agree well with the top binding affinities of in vivo experiments.
Tak-Ming Chan, Leung-Yau Lo, Ho-Yin Sze-To, Kwong-Sak Leung, Xinshu Xiao, Man Hon Wong 0001
IEEE ACM Trans. Comput. Biol. Bioinform.6
2012 idock: A multithreaded virtual screening tool for flexible ligand docking
abstract
AutoDock Vina is a competitive protein-ligand docking tool well known for its fast execution and high accuracy. Nevertheless, when docking a massive number of ligands, Vina has to be run multiple times, repeating receptor parsing and grid maps building over and over again. There are tremendous requests for revising Vina to reuse precalculated data and incorporate built-in support for virtual screening. Hence we developed idock, inheriting from AutoDock Vina the accurate scoring function and the efficient optimization algorithm, and significantly improving the fundamental implementation and numerical model for even faster execution. idock achieves a speedup of 3.3 in terms of CPU time and a speedup of 7.5 in terms of elapsed time on average. idock is free and open source, available at https://GitHub.com/HongjianLi/idock.
Kwong-Sak Leung, Man Hon Wong 0001
CIBCB3
2012 Efficient Algorithm for Mining Correlated Protein-DNA Binding Cores
Po-Yuen Wong, Tak-Ming Chan, Man Hon Wong 0001, Kwong-Sak Leung
DASFAA (1)3
2012 Predicting Approximate Protein-DNA Binding Cores Using Association Rule Mining
abstract
The studies of protein-DNA bindings between transcription factors (TFs) and transcription factor binding sites (TFBSs) are important bioinformatics topics. High-resolution (length;490) are shown promising in identifying accurate binding cores without using any 3D structures. While the current association rule mining method on this problem addresses exact sequences only, the most recent ad hoc method for approximation does not establish any formal model and is limited by experimentally known patterns. As biological mutations are common, it is desirable to formally extend the exact model into an approximate one. In this paper, we formalize the problem of mining approximate protein-DNA association rules from sequence data and propose a novel efficient algorithm to predict protein-DNA binding cores. Our two-phase algorithm first constructs two compact intermediate structures called frequent sequence tree (FS-Tree) and frequent sequence class tree (FSCTree). Approximate association rules are efficiently generated from the structures and bioinformatics concepts (position weight matrix and information content) are further employed to prune meaningless rules. Experimental results on real data show the performance and applicability of the proposed algorithm.
Po-Yuen Wong, Tak-Ming Chan, Man Hon Wong 0001, Kwong-Sak Leung
ICDE3
2012 A novel web-based system for tropical cyclone analysis and prediction
abstract
A web-based system is developed for the analysis and prediction of tropical cyclones, particularly their landfalls and recurvatures. To facilitate accessibility to the system, its development is based on Google Maps application programming interface (API), Java and client/server architecture. In addition to the construction of a powerful query system for the multi-source, multi-scale and multi-level tropical cyclone database, data mining approach and dynamic modelling approach have been implemented and integrated for effective and efficient analysis, prediction and visualization of tropical cyclone movements. The system can be accessed worldwide by researchers, professionals and the general public. It is thus a powerful system for research, real-life application and knowledge dissemination. Its extensibility and user-friendliness pave the road for further development and enable more in-depth analysis and real-time operation.
Yee Leung, Man Hon Wong 0001, Ka-Chun Wong, Wei Zhang 0048, Kwong-Sak Leung
Int. J. Geogr. Inf. Sci.2
2011 Interactive Drug Design in Virtual Reality
abstract
Discovering new drugs for emerging diseases has been a challenging task. There are numerous drug design techniques including fragment-based and diversity-oriented methods but their accuracies and efficiencies are low. By incorporating visualisation, biomedical experts can interact with the process to produce drug-like ligands more efficiently. The paper presents an interactive drug design algorithm which generates lead candidates against a protein. A set of drug candidates, created by an in house fragment-based method and docked on the target protein, are visualised in the virtual reality settings. Biomedical experts can investigate and select some of the ligands for further processing, aided with distance and bonding information. It also assists the user to drag and rotate the ligand to the binding site they find suitable. The algorithm runs iteratively and improves the quality of lead candidates every step. The paper compares the quality of resulting ligands between interactive and automatic approaches.
Ching-Man Tse, Kwong-Sak Leung, Kin-Hong Lee, Man Hon Wong 0001
IV5
2011 Discovering approximate-associated sequence patterns for protein-DNA interactions
abstract
MOTIVATION: The bindings between transcription factors (TFs) and transcription factor binding sites (TFBSs) are fundamental protein-DNA interactions in transcriptional regulation. Extensive efforts have been made to better understand the protein-DNA interactions. Recent mining on exact TF-TFBS-associated sequence patterns (rules) has shown great potentials and achieved very promising results. However, exact rules cannot handle variations in real data, resulting in limited informative rules. In this article, we generalize the exact rules to approximate ones for both TFs and TFBSs, which are essential for biological variations. RESULTS: A progressive approach is proposed to address the approximation to alleviate the computational requirements. Firstly, similar TFBSs are grouped from the available TF-TFBS data (TRANSFAC database). Secondly, approximate and highly conserved binding cores are discovered from TF sequences corresponding to each TFBS group. A customized algorithm is developed for the specific objective. We discover the approximate TF-TFBS rules by associating the grouped TFBS consensuses and TF cores. The rules discovered are evaluated by matching (verifying with) the actual protein-DNA binding pairs from Protein Data Bank (PDB) 3D structures. The approximate results exhibit many more verified rules and up to 300% better verification ratios than the exact ones. The customized algorithm achieves over 73% better verification ratios than traditional methods. Approximate rules (64-79%) are shown statistically significant. Detailed variation analysis and conservation verification on NCBI records demonstrate that the approximate rules reveal both the flexible and specific protein-DNA interactions accurately. The approximate TF-TFBS rules discovered show great generalized capability of exploring more informative binding rules.
Tak-Ming Chan, Ka-Chun Wong, Kin-Hong Lee, Man Hon Wong 0001, Terrence Chi-Kong Lau, Stephen Kwok-Wing Tsui, Kwong-Sak Leung
Bioinform.4
2011 Boundary-based lower-bound functions for dynamic time warping and their indexing
Man Hon Wong 0001
Inf. Sci.2
2011 Generalizing and learning protein-DNA binding sequence representations by an evolutionary algorithm
Ka-Chun Wong, Chengbin Peng 0001, Man Hon Wong 0001, Kwong-Sak Leung
Soft Comput.3
2011 Fast and Simultaneous Data Aggregation Over Multiple Regions in Wireless Sensor Networks
abstract
As the applications of wireless sensor networks continue to expand, it is important to support fast and simultaneous data aggregation over multiple regions for advanced data analysis. In this paper, we propose a solution by using a novel distributed data structure called distributed data cube (DDC). A DDC maintains a set of special forms of aggregate values (prefix sum, prefix average, prefix max, and prefix min) in distributed sensor nodes. We will first present fast algorithms to build a DDC within a sharp time bound. Then, we will present efficient distributed query-processing algorithms to handle aggregate queries by using a DDC. For a query region withnsensor nodes, our algorithms can return withinO(√n) time. Finally, extensive simulation studies confirm that a DDC can be built very quickly, which is consistent with the theoretical time bound. The network traffic injected while constructing a DDC is acceptable and also scalable as the network size grows. Query processing on a DDC is fast and energy efficient in terms of the time units needed and the number of messages incurred.
Dan Wu 0005, Man Hon Wong 0001
IEEE Trans. Syst. Man Cybern. Part C2
2010 Effect of Spatial Locality on an Evolutionary Algorithm for Multimodal Optimization
Ka-Chun Wong, Kwong-Sak Leung, Man Hon Wong 0001
EvoApplications (1)3
2010 Protein structure prediction on a lattice model via multimodal optimization techniques
abstract
This paper considers the protein structure prediction problem as a multimodal optimization problem. In particular, de novo protein structure prediction problems on the 3D Hydrophobic-Polar (HP) lattice model are tackled by evolutionary algorithms using multimodal optimization techniques. In addition, a new mutation approach and performance metric are proposed for the problem. The experimental results indicate that the proposed algorithms are more effective than the state-of-the-arts algorithms, even though they are simple.
Ka-Chun Wong, Kwong-Sak Leung, Man Hon Wong 0001
GECCO3
2010 Quantifying Private Information with Human Factors in Data Publishing for Binary Sensitive Attributes
abstract
Human factors play a key role in many problems, especially those involve human. In this paper, we address the problem of quantifying private information with the consideration of human factors in data publishing for binary sensitive attributes. We first propose five axioms to capture the properties of private information in data publishing. In particular, the fifth axiom considers human factors in the view of social psychology. We believe that a function that quantifies private information has to satisfy all these five axioms. Therefore, we propose an expected gain model, which allows users to tune a weighting factor to reflect the importance of human factors in their applications. The proposed model has been analyzed and compared with some existing models, such as information gain. Moreover, experiments have been performed on real applications to study the practicality of the proposed model.
Chi Hong Cheong, Dan Wu 0005, Man Hon Wong 0001
Int. J. Uncertain. Fuzziness Knowl. Based Syst.3
2009 An evolutionary algorithm with species-specific explosion for multimodal optimization
abstract
This paper presents an evolutionary algorithm, which we call Evolutionary Algorithm with Species-specific Explosion (EASE), for multimodal optimization. EASE is built on the Species Conserving Genetic Algorithm (SCGA), and the design is improved in several ways. In particular, it not only identifies species seeds, but also exploits the species seeds to create multiple mutated copies in order to further converge to the respective optimum for each species. Experiments were conducted to compare EASE and SCGA on four benchmark functions. Cross-comparison with recent rival techniques on another five benchmark functions was also reported. The results reveal that EASE has a competitive edge over the other algorithms tested.
Ka-Chun Wong, Kwong-Sak Leung, Man Hon Wong 0001
GECCO3
2009 Supporting asynchronous update for distributed data cubes
Dan Wu 0005, Chi Hong Cheong, Man Hon Wong 0001
J. Netw. Comput. Appl.3
2008 N-SAMSAM : A simple and faster algorithm for solving approximate matching in DNA sequences
abstract
This work proposes a novel algorithm to do approximate matching in a database consisting of multiple sequences. We apply Agrep algorithm in an indexing structure, the r-cut numerical substring array (r-NSA). The structure basically indexes all the substrings of length r. The advantage of using the r-NSA is two-fold: (1) The space requirement of the r-NSA is much smaller than that of the other existing indexing structures, such as the generalized suffix tree. (2) We propose an algorithm to apply Agrep in the r-NSA, in which the substrings are processed sequentially. Since the common substrings are processed only once, the cost of our algorithm is smaller than that of the full scanning search by Agrep. Consequently, the matching time of our algorithm is also reduced. We design experiments to validate and compare the performance of our algorithm against the full scanning search by Agrep. We define the speed-up of our algorithm as the time required by the full scanning search by Agrep over that of our algorithm. We use eight sets of real DNA sequences in our experiments, and the results show that our algorithm achieves significant speed-up. We also investigate the speed-up of difference data sets, and analyze their differences in detail.
Bing Ni, Man Hon Wong 0001, Kwong-Sak Leung
IEEE Congress on Evolutionary Computation2
2008 Efficient Online Subsequence Searching in Data Streams under Dynamic Time Warping Distance
abstract
Data streams of real numbers are generated naturally in many applications. The technology of online subsequence searching in data streams becomes more and more important for monitoring and mining stream data. Due to its capability of handling temporal distortions in sequences, dynamic time warping (DTW) distance is a widely used similarity measure for time-series pattern matching. Unfortunately, because of the high computational complexity of DTW, no one has proposed efficient methods for online subsequence searching under DTW distance, especially over high speed data streams. In this paper, we observe that some important properties of DTW can be used to eliminate a lot of redundant computations. Based on these properties, an efficient batch filtering method for online subsequence searching in data streams is proposed. The experimental results show that when no global path constraint is used, the proposed method outperforms the best known method up to 25 times in terms of throughput. When global path constraint is considered, the proposed method can still outperform the rival method under most of the settings of the global path constraint, although our method does not exploit any information about the constraint.
Man Hon Wong 0001
ICDE2
2008 A Precise Termination Condition of the Probabilistic Packet Marking Algorithm
abstract
The probabilistic packet marking (PPM) algorithm is a promising way to discover the Internet map or an attack graph that the attack packets traversed during a distributed denial-of-service attack. However, the PPM algorithm is not perfect, as its termination condition is not well defined in the literature. More importantly, without a proper termination condition, the attack graph constructed by the PPM algorithm would be wrong. In this work, we provide a precise termination condition for the PPM algorithm and name the new algorithm the rectified PPM (RPPM) algorithm. The most significant merit of the RPPM algorithm is that when the algorithm terminates, the algorithm guarantees that the constructed attack graph is correct, with a specified level of confidence. We carry out simulations on the RPPM algorithm and show that the RPPM algorithm can guarantee the correctness of the constructed attack graph under 1) different probabilities that a router marks the attack packets and 2) different structures of the network graph. The RPPM algorithm provides an autonomous way for the original PPM algorithm to determine its termination, and it is a promising means of enhancing the reliability of the PPM algorithm.
Tsz-Yeung Wong, Man Hon Wong 0001, John C. S. Lui
IEEE Trans. Dependable Secur. Comput.2
2007 Boundary-Based Lower-Bound Functions for Dynamic Time Warping and Their Indexing
abstract
Lower-bound functions are crucial for indexing time-series data under dynamic time warping (DTW) distance. In this paper, we propose a unified framework to explain the existing lower-bound functions. Based on the framework, we further propose a group of lower-bound functions for DTW and investigate their performances through extensive experiments. Experimental results show that the new lower-bound functions are better than the existing one in most cases. An index structure based on the new lower-bound functions is also implemented.
Man Hon Wong 0001
ICDE2
2006 An Efficient Distributed Algorithm to Identify and Traceback DDoS Traffic
abstract
Distributed denial-of-service attack is one of the most pressing security problems that the Internet community needs to address. Two major requirements for effective traceback are (i) to quickly and accurately locate potential attackers and (ii) to filter attack packets so that a host can resume the normal service to legitimate clients. Most of the existing IP traceback techniques focus on tracking the location of attackers after-the-fact. In this work, we provide an efficient methodology for locating potential attackers who employ the flood-based attack. We propose a distributed algorithm so that a set of routers can correctly (in a distributed sense) gather statistics in a coordinated fashion and that a victim site can deduce the local traffic intensities of all these participating routers. We prove the correctness of our distributed algorithm, and given the collected statistics, we provide a method for the victim site to locate attackers who sent out dominating flows of packets. The proposed distributed traceback methodology can also complement and leverage on the existing ICMP traceback so that a more efficient and accurate traceback can be obtained. We carry out simulations to illustrate that the proposed methodology can locate the attackers in a short period of time. Moreover, the applications as well as the limitations of the proposed methodology are covered. We believe this work also provides the theoretical foundation on how to correctly and accurately perform distributed measurement and traffic estimation on the Internet.
Tsz-Yeung Wong, K. T. Law, John C. S. Lui, Man Hon Wong 0001
Comput. J.4
2006 A geometrical solution to time series searching invariant to shifting and scaling
Man Hon Wong 0001, Kelvin Kam Wing Chu
Knowl. Inf. Syst.2
2006 Stream segregation algorithm for pattern matching in polyphonic music databases
Wai Man Szeto, Man Hon Wong 0001
Multim. Tools Appl.2
2005 Aggregate Sum Retrieval in Sensor Network by Distributed Prefix Sum Data Cube
abstract
Several data aggregation algorithms for sensor networks have been proposed. They are capable of returning the aggregate value of a single set of sensors. However, when data aggregates of several sets of sensors are needed at the same time, the only solution these techniques provide is to build multiple distributed data structures or gossip groups in these sets of sensors. Hence in a sensor network with N sensors, we may need 2/sup N/ distributed data structures or gossip groups in order to get the aggregates of all possible sets of sensors. In this paper, we propose a novel and data-centric technique for the fast retrieval of aggregate sums from multiple regions in a sensor network, using only one single distributed data structure. Our idea is to construct a distributed data cube in the sensor network. The distributed data cube construction algorithm we propose makes use of the inclusion-exclusion principle and it can build a distributed prefix sum data cube in a sensor network in O(N) worst case time. With the distributed data cube, data aggregate queries on any rectangular regions in the sensor network can be answered in just a constant number of operations.
Lok Hang Lee, Man Hon Wong 0001
AINA2
2005 A segment-wise time warping method for time scaling searching
Man Hon Wong 0001
Inf. Sci.2
2003 A Stream Segregation Algorithm for Polyphonic Music Databases
abstract
Most of the existing algorithms for music information retrieval are based on string matching. However, some searching results are perceptually insignificant in the sense that they cannot really be heard, owing to negligence of how people perceive music. When listening to music, it is perceived in groupings of musical notes called streams. Stream-crossing musical patterns are perceptually insignificant and should be pruned out from the final results. Stream segregation should be added as a pre-processing or post-processing step in existing retrieval systems in order to improve the quality of retrieval results. The key ideas are: (a) representation of music in the form of events, (b) formulation of the inter-event and the intercluster distance functions based on the findings in auditory psychology, and (c) application of the distance functions in the adapted single-link clustering algorithm without input of number of clusters. Experiments are performed on real music data to verify our proposed method.
Wai Man Szeto, Man Hon Wong 0001
IDEAS2
2003 Efficient Subsequence Matching for Sequences Databases under Time Warping
abstract
It has been found that the technique of searching for similar patterns among time series data is very important in a wide range of scientific and business applications. Most of the research works use Euclidean distance as their similarity metric. However, dynamic time warping (DTW) is a more robust distance measure than Euclidean distance in many situations, where sequences may have different lengths or have patterns which are out of phase in the time axis. Unfortunately, DTW does not satisfy the triangle inequality, so spatial indexing techniques cannot be applied. In this paper, we present a method that supports dynamic time warping for subsequence matching within a collection of sequences. Our method takes full advantage of the "sliding window" approach and can handle queries of arbitrary length.
Teddy Siu Fung Wong, Man Hon Wong 0001
IDEAS2
2001 Efficient and Robust Feature Extraction and Pattern Matching of Time Series by a Lattice Structure
abstract
The efficiency of searching scaling-invariant and shifting-invariant shapes in a set of massive time series data can be improved if searching is performed on an approximated sequence which involves less data but contains all the significant features. However, commonly used smoothing techniques, such as moving averages and best-fitting polylines, usually miss important peaks and troughs and deform the time series. In addition, these techniques are not robust, as they often requires users to supply a set of smoothing parameters which has direct effect on the resultant approximation pattern. To address these problems, an algorithm to construct a lattice structure as an underlying framework for pattern matching is proposed in this paper. As inputs, the algorithm takes a time series and users' requirements of level of detail. The algorithm then identifies all the important peaks and troughs (known as controlm points) in the time series and classifies the points into appropriate layers of the lattice structure. The control points in each layer of the structure form an approximation pattern an yet preserve the overall shape of the original series with approximation error lies within certain bound. The lower the layer, the more precise the approximation pattern is. Putting in another way, the algorithm takes different levels of data smoothing into account. Also, the lattice structure can be indexed to further improve the performance of pattern matching.
Wan Po Man Polly, Man Hon Wong 0001
CIKM2
2000 An experimental study of semantics-based concurrency control protocols
Hang Kwong Mak, Man Hon Wong 0001
Data Knowl. Eng.2
2000 Diamond Quorum Consensus for High Capacity and Efficiency in a Replicated Database System
Ada Wai-Chee Fu, Yat Sheung Wong, Man Hon Wong 0001
Distributed Parallel Databases3
1999 Interactive Data Analysis on Numeric-Data
abstract
Data mining has been a hot topic in computer science. Many researchers have been putting lots of effort into how to extract explicit knowledge from large databases. Among the problems in data mining, finding useful patterns in large databases has attracted lots of interest in recent years. However, like other data mining algorithms, most of the proposed clustering algorithms suffer from the same demerit: lack of user interaction and exploration. In this paper, a new algorithm called IDAN (Interactive Data Analysis on Numeric-data) is being introduced. IDAN is good in discovering clustering patterns from numeric data. This algorithm is incremental and provides more user interaction in the mining process. At the same time, it allows the user to explore the rules or clusters found when integrated with a visualizer.
Hong Ki Chu, Man Hon Wong 0001
IDEAS2
1999 Fast Time-Series Searching with Scaling and Shifting
abstract
Recently, it has been found that the technique of searching for similar patterns among time series data is very important in a wide range of scientific and business applications.In this paper, we first propose a definition of similarity based on scaling and shifting transformations.Sequence A is defined to be similar to sequence B if suitable scaling and shifting transformations can be found to transform A to B. Then, we present a geometrical view of the problem so that the scaling factor and the shifting offset can be determined.Moreover, sequence searching based on tree-based indexing structure can be performed.Finally, some technical aspects are discussed and some experiments are performed on real data (stock price movement) to measure the performance of our algorithm.
Kelvin Kam Wing Chu, Man Hon Wong 0001
PODS2
1999 Distributed Caching and Broadcast in a Wireless Mobile Computing Environment
abstract
In a mobile computing system, the wireless communication bandwidth is a scarce resource that needs to be managed carefully. In this paper, we investigate the use of distributed caching as an approach to reduce the wireless bandwidth consumption for data access. We find that conventional caching techniques cannot fully utilize the dissemination feature of the wireless channel. We thus propose a novel distributed caching protocol that can minimize the overall system bandwidth consumption at the cost of central processor unit processing time at the server side. This protocol allows the base station to select data items into a broadcast set, based on a performance gain parameter called the bandwidth gain, and then send the broadcast set to all the mobile computers within the server's cell. We show that, in general, this selection process is NP-hard and therefore, we propose a heuristic algorithm that can attain a near-optimal performance. We also propose an analytical model for the protocol and derive closed-form performance measures, such as the bandwidth utilization and the expected response time of data access by mobile computers. Experiments show that our distributed caching protocol can greatly reduce the bandwidth consumption so that the wireless network environment can accommodate more users, and at the same time vastly improve the expected response time for data access by mobile computers.
Cedric C. F. Fong, John C. S. Lui, Man Hon Wong 0001
Comput. J.3
1998 Transient Performance Analysis for Location Update Protocols in Cellular Networks
abstract
Currently, many location update protocols have been proposed for mobile terminal tracking in cellular networks. These protocols, such as the time-based and the distance-based protocols, try to minimize the system overhead for locating a mobile user in the network. The contribution of this paper to develop a novel mathematical model to analyze these protocols and at the same time, make a general model to capture many important movement features. In this paper, we propose to use the transient analysis method to predict the location of a mobile user and the accuracy of the prediction of the location of a user can be enhanced. By applying the transient analysis to evaluate the time-based and the distance-based protocol, it gives us more insights in the performance of these protocols under different movement configurations. Last but not least, our model is general enough such that it provides a general framework to analyze other location update protocols in cellular networks.
Cedric C. F. Fong, John C. S. Lui, Man Hon Wong 0001, Edmundo de Souza e Silva
MASCOTS3
1998 An Efficient Hash-Based Algorithm for Sequence Data Searching
abstract
In real life, data collected day by day often appear in sequences and this type of data is called sequence data. The technique of searching for similar patterns among sequence data is very important in many applications. We first point out that there are some deficiencies in the existing definitions of sequence similarity. We then introduce a definition of sequence similarity based on the shape of sequences. The definition is also extended to handle sequence matching with linear scaling in both amplitude and time dimensions. A fast sequence searching algorithm based on extendable hashing is also proposed. The algorithm can match linearly scaled sequences and guarantee that no qualified data subsequence is falsely rejected. Several experiments are performed on real data (stock price movement) and synthetic data to measure the performance of the algorithm in different aspects.
Kelvin Kam Wing Chu, Sze Kin Lam, Man Hon Wong 0001
Comput. J.3
1998 A Fast Projection Algorithm for Sequence Data Searching
Sze Kin Lam, Man Hon Wong 0001
Data Knowl. Eng.2
1997 Quantifying Complexity and Performance Gains of Distributed Caching in a Wireless Mobile Computing Environment
abstract
In a mobile computing system, the wireless communication bandwidth is a scarce resource that needs to be managed carefully. In this paper, we investigate the use of distributed caching as an approach to reduce the wireless bandwidth consumption for data access. We find that conventional caching techniques cannot fully utilize the dissemination feature of the wireless channel. We thus propose a novel distributed caching protocol that can minimize the overall system bandwidth consumption at the cost of CPU processing time at the server side. This protocol allows the server to select data items into a broadcast set, based on a performance gain parameter called the bandwidth gain, and then send the broadcast set to all the mobile computers within the server's cell. We show that in general, this selection process is NP-hard, and therefore we propose a heuristic algorithm that can attain a near-optimal performance. We also propose an analytical model for the protocol and derive closed-form performance measures, such as the bandwidth utilization and the expected response time of data access by mobile computers. Experiments show that our distributed caching protocol can greatly reduce the bandwidth consumption so that the wireless network environment can accommodate more users and, at the same time, vastly improve the expected response time for data access by mobile computers.
Cedric C. F. Fong, John C. S. Lui, Man Hon Wong 0001
ICDE3
1997 Bounded Inconsistency for Type-Specific Concurrency Control
Man Hon Wong 0001, Divyakant Agrawal, Hang Kwong Mak
Distributed Parallel Databases1
1996 Recovery for Transaction Failures in Object-Based Databases
abstract
A set of recoverability theory is derived in this paper for an object-based database. Instead of considering serializability and recoverability as two orthogonal concepts, we simply keep serializability as the only correctness criterion and require serializability to be maintained even when failures of transactions may occur. Based on this fundamental notion of correctness, the definition of recoverability is derived. The recoverability theory derived in this way is a generalization of the traditional recoverability theory in the read/write model. In addition, we find that the set of strict histories depends on the strength of the inverse operations being used to cancel the effects of aborted operations. At one extreme, when the strongest inverse operations are used, the set of strict histories is the same as the set of avoid cascading aborts histories. At the other extreme, when the weakest inverse operations are used, the set of strict histories is the same as the set of rigorous his...
Man Hon Wong 0001
PODS1
1995 Trading Operation Consistency for Concurrency
Hang Kwong Mak, Man Hon Wong 0001
DASFAA2
1995 On Distributed Object Checkpointing and Recovery
abstract
Recoveryby checkpointing on distributed shared memory systems is investigated in this paper.The no-
Manhoi Choy, Hong Va Leong, Man Hon Wong 0001
PODC3
1995 Context-Specific Synchronization for Atomic Data Types in Object-Based Databases
Man Hon Wong 0001, Divyakant Agrawal
Theor. Comput. Sci.1
1993 Context-Based Synchronisation: An Approach beyond Semantics for Concurrency Control
abstract
The expressiveness of various object-oriented languages is investigated with respect to their ability to create new objects. We focus on database method schemas (dms), a model capturing the data manipulation capabilities of a large class of deterministic methods in object-oriented databases. The results clarify the impact of various language constructs on object creation. Several new constructs based on expanded notions of deep equality are introduced. In particular, we provide a tractable construct which yields a language complete with respect to object creation. The new construct is also relevant to query complexity. For example, it allows expressing in polynomial time some queries, like counting, requiring exponential space in dms alone.
Man Hon Wong 0001, Divyakant Agrawal
PODS1
1992 Context-Specific Synchronization for Atomic Data Types
Man Hon Wong 0001, Divyakant Agrawal
ICDT1
1992 Tolerating Bounded Inconsistency for Increasing Concurrency in Database Systems
abstract
Recently, the scope of databases has been extended to many non-standard applications, and serializability is found to be too restrictive for such applications. In general, two approaches are adopted to address this problem. The first approach considers placing more structure on data objects to exploit type specific properties while keeping serializability as the correctness criterion. The other approach uses explicit semantics of transactions and databases to permit interleaved executions of transactions that are non-serializable. In this paper, we attempt to bridge the gap between the two approaches by using the notion of serializability with bounded inconsistency. Users are free to specifiy the maximum level of inconsistency that can be allowed in the executions of operations dynamically. In particular, if no inconsistency is allowed in the execution of any operation, the protocol will be reduced to a standard strict two phase locking protocol based on type-specific semantics of data objects. Bounded inconsistency can be applied to many areas which do not require exact values of the data such as for gathering information for statistical purpose, for making high level decisions and reasoning in expert systems which can tolerate uncertainty in input data.
Man Hon Wong 0001, Divyakant Agrawal
PODS1
1992 Fuzzy concepts in an object oriented expert system shell
abstract
Fuzzy logic is one of the methods to model the vagueness and imprecision of human knowledge. Some rule-based expert system shells have been successfully developed and have demonstrated the power of fuzzy logic in dealing with inexact reasoning and rule inferences. However, using rules for knowledge representation is not structured enough. In addition, knowledge cannot be easily represented in an abstracted (hierarchical) from. In this article the introduction of fuzzy concepts into object oriented knowledge representation (OOKR), which is a structured knowledge representation scheme, is presented. A framework for handling all the possible fuzzy concepts in OOKR at both the dynamic and static levels is proposed. In order to handle the inheritance mechanism and to model the relations among classes, instances, and attributes, some new fuzzy concepts and operations are introduced. These concepts and operations are developed from the semantic meaning rather than by an ad hoc approach. A prototype of the expert system shell. System FX-I, has been successfully developed based on the above framework, showing the feasibility of handling inexact knowledge in a structural way.
Kwong-Sak Leung, Man Hon Wong 0001
Int. J. Intell. Syst.2
1990 A fuzzy database-query language
Man Hon Wong 0001, Kwong-Sak Leung
Inf. Syst.1
1989 A Fuzzy Expert Database System
Kwong-Sak Leung, Man Hon Wong 0001
Data Knowl. Eng.2