Kee Yuan Ngiam

dblp:192/7337 · DBLP profile ↗
← Back
17ranked-venue papers
0as first author
10since 2021 · last 2026
0000-0001-5676-2520ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 9 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 7 since 2021Artificial intelligence and machine learning · 4Graphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2026 Unsupervised Adaptive Path Optimization for Knowledge Graph Reasoning in Multimodal Medical Diagnosis
Xiaoqing Li 0005, Wenbin Feng, Yu Lu 0001, Judice Koh, Ellie Choi, Jianli Chen, Jinhong He, Kee Yuan Ngiam
ICIC (15)8
2026 Knowledge Graph-Enhanced Medical Large Language Models: Challenges, Methods, and Future Directions
Xiaoqing Li 0005, Yu Lu 0001, Kee Yuan Ngiam
ICIC (15)4
2025 Evaluating Image Matching With Robust Estimators: Bridging Natural and Surgical Domains to Enhance Scene Understanding
abstract
State-of-the-art image matching methods have shown strong generalization across natural image datasets, but their effectiveness in complex surgical environments remains underexplored. Surgical scenes introduce unique challenges, including homogeneous tissue textures, variable lighting, and frequent occlusions, which can degrade the reliability of keypoint correspondences essential for downstream vision tasks such as camera pose estimation and structure-from-motion. In this study, we systematically evaluate leading image matching methods within laparoscopic surgical settings, emphasizing performance under resource-constrained conditions. We present an optimized evaluation pipeline that incorporates robust estimators to enhance correspondence filtering and assess their impact on pose estimation accuracy. Our approach also examines the influence of fine-tuning individual pipeline components, particularly robust estimators, on overall system performance. Mean Reprojection error is refined by thresholding the nearest ground truth projections, enabling a more precise characterization of matching accuracy. Across five robust estimators, FM_8PTS consistently demonstrates superior resilience to outliers. Our results establish RoMa as the leading model for balancing pose estimation accuracy, reprojection performance, and computational efficiency, making it suitable for real-time surgical applications. By providing the first systematic benchmark and actionable insights for optimizing image matching pipelines in surgical domains, this work sets a new standard and paves the way for more reliable, efficient, and clinically applicable image-guided tools in minimally invasive surgery.
Ying Zhen Tan, Haojie Cheng, Kian Wei Ng, Kee Yuan Ngiam, Eng Tat Khoo
IEEE J. Biomed. Health Informatics4
2024 FHRDiff: Leveraging Diffusion Models for Conditional Fetal Heart Rate Signal Generation
abstract
Accurate analysis of Fetal Heart Rate (FHR) signal is often impeded by challenges such as data scarcity and label imbalance, which affect the reliability and robustness of deep learning models. To address these challenges, this study introduces FHRDiff, a novel diffusion model conditioned on Phase-Rectified Signal Averaging (PRSA) spectrograms for the creation of synthetic FHR signals. Our model integrates time encoding, condition generation from PRSA spectrograms, and residual blocks with dilated convolutions to effectively manage temporal dynamics and long-range dependencies. Extensive qualitative and quantitative experiments on FHR signal synthesis demonstrate the feasibility and effectiveness of FHRDiff. Compared to Generative Adversarial Networks (GANs) and image-based diffusion model, our method achieves the highest signal fidelity and distribution similarity, with key measures including 0.067 maximum mean deviation (MMD), 0.492 percent root mean square difference (PRD), 1.763 relative entropy (RE), and 0.160 Frechet distance (FD). Expert validation confirms the model’s capacity to accurately generate data for normal and abnormal FHR signals based on the paired condition. In addition, an ablation study was conducted to highlight that the spectrogram paired condition can guide the diffusion model to produce synthetics FHR signals with greater diversity compared to unconditional models. The results emphasize diffusion models’ potential for broad application in biomedical time series analysis, such as generation, imputation and noise removal.
Xiaoqing Li 0005, Yu Lu 0001, Kee Yuan Ngiam, Zichang Yu, Mohammad Shaheryar Furqan
BIBM3
2022 SurvMaximin: Robust federated approach to transporting survival risk prediction models
Harrison G. Zhang, Xin Xiong 0006, Chuan Hong, Griffin M. Weber, Gabriel A. Brat, Clara-Lea Bonzel, Yuan Luo 0001, Rui Duan 0004, Nathan P. Palmer, Meghan Hutch, Alba Gutiérrez-Sacristán, Riccardo Bellazzi, Luca Chiovato, Kelly Cho, Arianna Dagliati, Hossein Estiri, Noelia García-Barrio, Romain Griffier, David A. Hanauer, Yuk-Lam Ho, John H. Holmes, Mark S. Keller, Jeffrey G. Klann, Sehi L'Yi, Sara Lozano-Zahonero, Sarah E. Maidlow, Adeline Makoudjou, Alberto Malovini, Bertrand Moal, Jason H. Moore, Michele Morris, Danielle L. Mowery, Shawn N. Murphy, Antoine Neuraz, Kee Yuan Ngiam, Gilbert S. Omenn, Lav P. Patel, Miguel Pedrera-Jiménez, Andrea Prunotto, Malarkodi J. Samayamuthu, Fernando J. Sanz Vidorreta, Emily Schriver, Petra Schubert, Pablo Serrano-Balazote, Andrew M. South, Amelia L. M. Tan, Byorn W. L. Tan, Valentina Tibollo, Patric Tippmann, Shyam Visweswaran, Zongqi Xia, William Yuan, Daniela Zöller, Isaac S. Kohane, Paul Avillach, Zijian Guo 0003, Tianxi Cai
J. Biomed. Informatics36
2021 MLCask: Efficient Management of Component Evolution in Collaborative Data Analytics Pipelines
abstract
With the ever-increasing adoption of machine learning for data analytics, maintaining a machine learning pipeline is becoming more complex as both the datasets and trained models evolve with time. In a collaborative environment, the changes and updates due to pipeline evolution often cause cumbersome coordination and maintenance work, raising the costs and making it hard to use. Existing solutions, unfortunately, do not address the version evolution problem, especially in a collaborative environment where non-linear version control semantics are necessary to isolate operations made by different user roles. The lack of version control semantics also incurs unnecessary storage consumption and lowers efficiency due to data duplication and repeated data pre-processing, which are avoidable.In this paper, we identify two main challenges that arise during the deployment of machine learning pipelines, and address them with the design of versioning for an end-to-end analytics system MLCask. The system supports multiple user roles with the ability to perform Git-like branching and merging operations in the context of the machine learning pipelines. We define and accelerate the metric-driven merge operation by pruning the pipeline search tree using reusable history records and pipeline compatibility information. Further, we design and implement the prioritized pipeline search, which gives preference to the pipelines that probably yield better performance. The effectiveness of MLCask is evaluated through an extensive study over several real-world deployment cases. The performance evaluation shows that the proposed merge operation is up to 7.8x faster and saves up to 11.9x storage space than the baseline method that does not utilize history records.
Zhaojing Luo, Sai Ho Yeung, Meihui Zhang 0001, Kaiping Zheng, Lei Zhu 0015, Gang Chen 0001, Feiyi Fan, Qian Lin 0002, Kee Yuan Ngiam, Beng Chin Ooi
ICDE9
2021 PACE: Learning Effective Task Decomposition for Human-in-the-loop Healthcare Delivery
abstract
Human-in-the-loop data analysis involves both machine learning models and humans in analytic tasks. In healthcare applications, human-in-the-loop data analysis is crucial in that the model can handle "easy" tasks and hand over "hard" ones to medical experts for assistance and medical judgment, where easy tasks are the ones for which the model can provide high accuracy and hard tasks vice versa. In this process, how to decompose tasks in an effective manner is an important stage. To achieve task decomposition, classification with a reject option is a solution. However, existing studies either directly implement a reject option or dive into the theoretical details of the rejection mechanism. Different from such studies, we aim to optimize general classifiers with a reject option and hence, optimize task decomposition for healthcare applications.
Kaiping Zheng, Gang Chen 0001, Melanie Herschel, Kee Yuan Ngiam, Beng Chin Ooi, Jinyang Gao
SIGMOD Conference4
2021 Validation of an internationally derived patient severity phenotype to support COVID-19 analytics from electronic health record data
abstract
OBJECTIVE: The Consortium for Clinical Characterization of COVID-19 by EHR (4CE) is an international collaboration addressing coronavirus disease 2019 (COVID-19) with federated analyses of electronic health record (EHR) data. We sought to develop and validate a computable phenotype for COVID-19 severity. MATERIALS AND METHODS: Twelve 4CE sites participated. First, we developed an EHR-based severity phenotype consisting of 6 code classes, and we validated it on patient hospitalization data from the 12 4CE clinical sites against the outcomes of intensive care unit (ICU) admission and/or death. We also piloted an alternative machine learning approach and compared selected predictors of severity with the 4CE phenotype at 1 site. RESULTS: The full 4CE severity phenotype had pooled sensitivity of 0.73 and specificity 0.83 for the combined outcome of ICU admission and/or death. The sensitivity of individual code categories for acuity had high variability-up to 0.65 across sites. At one pilot site, the expert-derived phenotype had mean area under the curve of 0.903 (95% confidence interval, 0.886-0.921), compared with an area under the curve of 0.956 (95% confidence interval, 0.952-0.959) for the machine learning approach. Billing codes were poor proxies of ICU admission, with as low as 49% precision and recall compared with chart review. DISCUSSION: We developed a severity phenotype using 6 code classes that proved resilient to coding variability across international institutions. In contrast, machine learning approaches may overfit hospital-specific orders. Manual chart review revealed discrepancies even in the gold-standard outcomes, possibly owing to heterogeneous pandemic conditions. CONCLUSIONS: We developed an EHR-based severity phenotype for COVID-19 in hospitalized patients and validated it at 12 international sites.
Jeffrey G. Klann, Hossein Estiri, Griffin M. Weber, Bertrand Moal, Paul Avillach, Chuan Hong, Amelia L. M. Tan, Brett K. Beaulieu-Jones, Victor M. Castro, Thomas Maulhardt, Alon Geva, Alberto Malovini, Andrew M. South, Shyam Visweswaran, Michele Morris, Malarkodi J. Samayamuthu, Gilbert S. Omenn, Kee Yuan Ngiam, Kenneth D. Mandl, Martin Boeker, Karen L. Olson, Danielle L. Mowery, Robert W. Follett, David A. Hanauer, Riccardo Bellazzi, Jason H. Moore, Ne-Hooi Will Loh, Douglas S. Bell, Kavishwar B. Wagholikar, Luca Chiovato, Valentina Tibollo, Siegbert Rieg, Anthony L. L. J. Li, Vianney Jouhet, Emily Schriver, Zongqi Xia, Meghan Hutch, Yuan Luo 0001, Isaac S. Kohane, Gabriel A. Brat, Shawn N. Murphy
J. Am. Medical Informatics Assoc.18
2021 Mobile health applications for older adults: a systematic review of interface and persuasive feature design
abstract
OBJECTIVE: Mobile-based interventions have the potential to promote healthy aging among older adults. However, the adoption and use of mobile health applications are often low due to inappropriate designs. The aim of this systematic review is to identify, synthesize, and report interface and persuasive feature design recommendations of mobile health applications for elderly users to facilitate adoption and improve health-related outcomes. MATERIALS AND METHODS: We searched PubMed, Embase, PsycINFO, CINAHL, and Scopus databases to identify studies that discussed and evaluated elderly-friendly interface and persuasive feature designs of mobile health applications using an elderly cohort. RESULTS: We included 74 studies in our analysis. Our analysis revealed a total of 9 elderly-friendly interface design recommendations: 3 recommendations were targeted at perceptual capabilities of elderly users, 2 at motor coordination problems, and 4 at cognitive and memory deterioration. We also compiled and reported 5 categories of persuasive features: reminders, social features, game elements, personalized interventions, and health education. DISCUSSION: Only 5 studies included design elements that were based on theories. Moreover, the majority of the included studies evaluated the application as a whole without examining end-user perceptions and the effectiveness of each single design feature. Finally, most studies had methodological limitations, and better research designs are needed to quantify the effectiveness of the application designs rigorously. CONCLUSIONS: This review synthesizes elderly-friendly interface and persuasive feature design recommendations for mobile health applications from the existing literature and provides recommendations for future research in this area and guidelines for designers.
Na Liu 0005, Jiamin Yin, Sharon Swee-Lin Tan, Kee Yuan Ngiam, Hock-Hai Teo
J. Am. Medical Informatics Assoc.4
2021 Improving Data Analytics with Fast and Adaptive Regularization
abstract
Deep Learning and Machine Learning models have recently been shown to be effective in many real world applications. While these models achieve increasingly better predictive performance, their structures have also become much more complex. A common and difficult problem for complex models is overfitting. Regularization is used to penalize the complexity of the model in order to avoid overfitting. However, in most learning frameworks, regularization function is usually set with some hyper-parameters where the best setting is difficult to find. In this paper, we propose an adaptive regularization method, as part of a large end-to-end healthcare data analytics software stack, which effectively addresses the above difficulty. First, we propose a general adaptive regularization method based on Gaussian Mixture (GM) to learn the best regularization function according to the observed parameters. Second, we develop an effective update algorithm which integrates Expectation Maximization (EM) with Stochastic Gradient Descent (SGD). Third, we design a lazy update and sparse update algorithm to reduce the computational cost by 4x and 20x, respectively. The overall regularization framework is fast, adaptive, and easy-to-use. We validate the effectiveness of our regularization method through an extensive experimental study over 14 standard benchmark datasets and three kinds of deep learning/machine learning models. The results illustrate that our proposed adaptive regularization method achieves significant improvement over state-of-the-art regularization methods.
Zhaojing Luo, Shaofeng Cai, Gang Chen 0001, Jinyang Gao, Wang-Chien Lee, Kee Yuan Ngiam, Meihui Zhang 0001
IEEE Trans. Knowl. Data Eng.6
2020 TRACER: A Framework for Facilitating Accurate and Interpretable Analytics for High Stakes Applications
abstract
In high stakes applications such as healthcare and finance analytics, the interpretability of predictive models is required and necessary for domain practitioners to trust the predictions. Traditional machine learning models, e.g., logistic regression (LR), are easy to interpret in nature. However, many of these models aggregate time-series data without considering the temporal correlations and variations. Therefore, their performance cannot match up to recurrent neural network (RNN) based models, which are nonetheless difficult to interpret. In this paper, we propose a general framework TRACER to facilitate accurate and interpretable predictions, with a novel model TITV devised for healthcare analytics and other high stakes applications such as financial investment and risk management. Different from LR and other existing RNN-based models, TITV is designed to capture both the time-invariant and the time-variant feature importance using a feature-wise transformation subnetwork and a self-attention subnetwork, for the feature influence shared over the entire time series and the time-related importance respectively. Healthcare analytics is adopted as a driving use case, and we note that the proposed TRACER is also applicable to other domains, e.g., fintech. We evaluate the accuracy of TRACER extensively in two real-world hospital datasets, and our doctors/clinicians further validate the interpretability of TRACER in both the patient level and the feature level. Besides, TRACER is also validated in a critical financial application. The experimental results confirm that TRACER facilitates both accurate and interpretable analytics for high stakes applications.
Kaiping Zheng, Shaofeng Cai, Horng Ruey Chua, Wei Wang 0059, Kee Yuan Ngiam, Beng Chin Ooi
SIGMOD Conference5
2018 Adaptive Lightweight Regularization Tool for Complex Analytics
abstract
Deep Learning and Machine Learning models have recently been shown to be effective in many real world applications. While these models achieve increasingly better predictive performance, their structures have also become much more complex. A common and difficult problem for complex models is overfitting. Regularization is used to penalize the complexity of the model in order to avoid overfitting. However, in most learning frameworks, regularization function is usually set as some hyper parameters, and therefore the best setting is difficult to find. In this paper, we propose an adaptive regularization method, as part of a large end-to-end healthcare data analytics software stack, which effectively addresses the above difficulty. First, we propose a general adaptive regularization method based on Gaussian Mixture (GM) to learn the best regularization function according to the observed parameters. Second, we develop an effective update algorithm which integrates Expectation Maximization (EM) with Stochastic Gradient Descent (SGD). Third, we design a lazy update algorithm to reduce the computational cost by 4x. The overall regularization framework is fast, adaptive and easy-to-use. We validate the effectiveness of our regularization method through an extensive experimental study over 13 standard benchmark datasets and three kinds of deep learning/machine learning models. The results illustrate that our proposed adaptive regularization method achieves significant improvement over state-of-the-art regularization methods.
Zhaojing Luo, Shaofeng Cai, Jinyang Gao, Meihui Zhang 0001, Kee Yuan Ngiam, Gang Chen 0001, Wang-Chien Lee
ICDE5
2018 Medical Concept Embedding with Time-Aware Attention
abstract
Embeddings of medical concepts such as medication, procedure and diagnosis codes in Electronic Medical Records (EMRs) are central to healthcare analytics. Previous work on medical concept embedding takes medical concepts and EMRs as words and documents respectively. Nevertheless, such models miss out the temporal nature of EMR data. On the one hand, two consecutive medical concepts do not indicate they are temporally close, but the correlations between them can be revealed by the time gap. On the other hand, the temporal scopes of medical concepts often vary greatly (e.g., common cold and diabetes). In this paper, we propose to incorporate the temporal information to embed medical codes. Based on the Continuous Bag-of-Words model, we employ the attention mechanism to learn a ``soft'' time-aware context window for each medical concept. Experiments on public and proprietary datasets through clustering and nearest neighbour search tasks demonstrate the effectiveness of our model, showing that it outperforms five state-of-the-art baselines.
Xiangrui Cai, Jinyang Gao, Kee Yuan Ngiam, Beng Chin Ooi, Ying Zhang 0015, Xiaojie Yuan
IJCAI3
2018 Fine-grained Concept Linking using Neural Networks in Healthcare
abstract
To unlock the wealth of the healthcare data, we often need to link the real-world text snippets to the referred medical concepts described by the canonical descriptions. However, existing healthcare concept linking methods, such as dictionary-based and simple machine learning methods, are not effective due to the word discrepancy between the text snippet and the canonical concept description, and the overlapping concept meaning among the fine-grained concepts. To address these challenges, we propose a Neural Concept Linking (NCL) approach for accurate concept linking using systematically integrated neural networks. We call the novel neural network architecture as the COMposite AttentIonal encode-Decode neural network (COM-AID). COM-AID performs an encode-decode process that encodes a concept into a vector and decodes the vector into a text snippet with the help of two devised contexts. On the one hand, it injects the textual context into the neural network through the attention mechanism, so that the word discrepancy can be overcome from the semantic perspective. On the other hand, it incorporates the structural context into the neural network through the attention mechanism, so that minor concept meaning differences can be enlarged and effectively differentiated. Empirical studies on two real-world datasets confirm that the NCL produces accurate concept linking results and significantly outperforms state-of-the-art techniques.
Meihui Zhang 0001, Gang Chen 0001, Ju Fan, Kee Yuan Ngiam, Beng Chin Ooi
SIGMOD Conference5
2017 Capturing Feature-Level Irregularity in Disease Progression Modeling
abstract
Disease progression modeling (DPM) analyzes patients' electronic medical records (EMR) to predict the health state of patients, which facilitates accurate prognosis, early detection and treatment of chronic diseases. However, EMR are irregular because patients visit hospital irregularly based on the need of treatment. For each visit, they are typically given different diagnoses, prescribed various medications and lab tests. Consequently, EMR exhibit irregularity at the feature level. To handle this issue, we propose a model based on the Gated Recurrent Unit by decaying the effect of previous records using fine-grained feature-level time span information, and learn the decaying parameters for different features to take into account their different behaviours like decaying speeds under irregularity. Extensive experimental results in both an Alzheimer's disease dataset and a chronic kidney disease dataset demonstrate that our proposed model of capturing feature-level irregularity can effectively improve the accuracy of DPM.
Kaiping Zheng, Wei Wang 0059, Jinyang Gao, Kee Yuan Ngiam, Beng Chin Ooi, James Wei Luen Yip
CIKM4
2017 Resolving the Bias in Electronic Medical Records
abstract
Electronic Medical Records (EMR) are the most fundamental resources used in healthcare data analytics. Since people visit hospital more frequently when they feel sick and doctors prescribe lab examinations when they feel necessary, we argue that there could be a strong bias in EMR observations compared with the hidden conditions of patients. Directly using such EMR for analytical tasks without considering the bias may lead to misinterpretation. To this end, we propose a general method to resolve the bias by transforming EMR to regular patient hidden condition series using a Hidden Markov Model (HMM) variant. Compared with the biased EMR series with irregular time stamps, the unbiased regular time series is much easier to be processed by most analytical models and yields better results. Extensive experimental results demonstrate that our bias resolving method imputes missing data more accurately than baselines and improves the performance of the state-of-the-art methods on typical medical data analytics.
Kaiping Zheng, Jinyang Gao, Kee Yuan Ngiam, Beng Chin Ooi, James Wei Luen Yip
KDD3
2016 Towards longitudinal analysis of a population's electronic health records using factor graphs
abstract
In this feasibility study, we demonstrate the use of a factor-graph-based probabilistic graphical model approach to process longitudinal data derived from a population's electronic health records (EHR). Processing of EHR allows for fore-casting patient-specific health complications and inference of population-level statistics on several epidemiological factors. As a case-study, we provide preliminary results and demonstrate feasibility of our approach by processing the EHR of a diabetic cohort in Singapore. Our model passes the feasibility test as we are able to forecast a series of health complications of a new patient based on the factor functions inferred from EHR of 100 diabetic patients spanning 10-years. This forecast gives both the caregivers and the patient a better view of the patient's health in the coming years and increases patient's motivation to stay healthy and conform to medication plan. Furthermore, our approach informs commonly occurring health complications in the population that warrant hospital readmissions, which helps a physician/clinician in decide when to intervene to avoid complications in order to improve the patient's quality of life and minimize the cost of care.
Arjun P. Athreya, Kee Yuan Ngiam, Zhaojing Luo, E. Shyong Tai, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer
BDCAT2