Nan Liu 0003

dblp:86/4643-3 · DBLP profile ↗
← Back
33ranked-venue papers
9as first author
18since 2021 · last 2026
0000-0003-3610-4883ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 20 · 4 first-author · 14 since 2021Artificial intelligence and machine learning · 10 · 3 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 first-authorDatabases, data management, data science and information retrieval · 2Human-computer interaction and ubiquitous computing · 2 · 2 first-authorSystems, architecture and hardware · 1
YearPublicationVenuePosition
2026 The detectability paradox: bilingual medical report generation with open-weight models and the limits of human oversight
abstract
OBJECTIVES: The automation of medical report generation using large language models (LLMs) could significantly reduce physicians' documentation burden while enhancing healthcare efficiency. However, the misuse of generative artificial intelligence in medical reporting can lead to important safety risks for patients. We addressed 2 questions: (1) What is the quality of medical reports generated by LLMs in English and French? and (2) Can we distinguish between human-written and LLM-generated medical reports? MATERIALS AND METHODS: We evaluated the quality of reports generated by several multilingual, open-weight LLMs using text similarity metrics on 4212 medical reports in English and French across multiple specialties. A bilingual expert panel of certified physicians (n = 4) and medical residents (n = 5) scored accuracy, fluency, and completeness of generated reports using a 1-5 Likert scale. Experts also completed a Turing-like test, blindly identifying reports as human or machine-generated. RESULTS: Phi-4 achieved the best overall performance (ROUGE-1: 0.70, BERTScore: 0.83). Expert evaluation confirmed high-quality reports in both languages (overall 4.6/5.0). Medical experts performed better than chance but struggled to differentiate human versus machine reports (accuracy: 0.60). Automatic classifiers showed strong performance (accuracy: 0.98). DISCUSSION: The high quality of LLM-generated reports supports their potential to enhance healthcare efficiency in multilingual settings. However, the discrepancy between human detection difficulty and automated detection success reveals inherent limitations in relying solely on human oversight for quality assurance and misuse prevention. CONCLUSIONS: Deployment of LLMs for medical reporting requires combining automated detection tools with human expertise to ensure patient safety. Dataset and code: https://github.com/ds4dh/medical_report_generation.
Hossein Rouhizadeh, Abiram Sandralegar, Anthony Yazdani, Weibo Feng, Oren Schreier, Yonnou Ahn-Kim, Assiya Sirbal, Valentino Pirelli, Rui Yang 0016, Lukas Sveikata, Elena Tessitore, Nan Liu 0003, Philippe Bijlenga, Douglas Teodoro
J. Am. Medical Informatics Assoc.12
2025 MMLU-ProX: A Multilingual Benchmark for Advanced Large Language Model Evaluation
abstract
Weihao Xuan, Rui Yang, Heli Qi, Qingcheng Zeng, Yunze Xiao, Aosong Feng, Dairui Liu, Yun Xing, Junjue Wang, Fan Gao, Jinghui Lu, Yuang Jiang, Huitao Li, Xin Li, Kunyu Yu, Ruihai Dong, Shangding Gu, Yuekang Li, Xiaofei Xie, Felix Juefei-Xu, Foutse Khomh, Osamu Yoshie, Qingyu Chen, Douglas Teodoro, Nan Liu, Randy Goebel, Lei Ma, Edison Marrese-Taylor, Shijian Lu, Yusuke Iwasawa, Yutaka Matsuo, Irene Li. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Weihao Xuan, Rui Yang 0016, Heli Qi, Qingcheng Zeng, Yunze Xiao, Aosong Feng, Dairui Liu, Yun Xing 0001, Jinghui Lu, Yuang Jiang, Huitao Li, Xin Li 0079, Kunyu Yu, Ruihai Dong, Shangding Gu, Yuekang Li, Xiaofei Xie, Felix Juefei-Xu, Foutse Khomh, Osamu Yoshie, Qingyu Chen 0001, Douglas Teodoro, Nan Liu 0003, Randy Goebel, Lei Ma 0003, Edison Marrese-Taylor, Shijian Lu, Yusuke Iwasawa, Yutaka Matsuo, Irene Li
EMNLP25
2025 Reporting guideline for chatbot health advice studies: The CHART statement
abstract
The Chatbot Assessment Reporting Tool (CHART) is a reporting guideline developed to provide reporting recommendations for studies evaluating the performance of generative artificial intelligence (AI)-driven chatbots when summarizing clinical evidence and providing health advice, referred to as Chatbot Health Advice (CHA) studies. CHART was developed in several phases after performing a comprehensive systematic review to identify variation in the conduct, reporting and methodology in CHA studies. Findings from the review were used to develop a draft checklist that was revised through an international, multidisciplinary modified asynchronous Delphi consensus process of 531 stakeholders, three synchronous panel consensus meetings of 48 stakeholders, and subsequent pilot testing of the checklist. CHART includes 12 items and 39 subitems to promote transparent and comprehensive reporting of CHA studies. These include Title (subitem 1a), Abstract/Summary (subitem 1b), Background (subitems 2ab), Model Identifiers (subitem 3ab), Model Details (subitems 4abc), Prompt Engineering (subitems 5ab), Query Strategy (subitems 6abcd), Performance Evaluation (subitems 7ab), Sample Size (subitem 8), Data Analysis (subitem 9a), Results (subitems 10abc), Discussion (subitems 11abc), Disclosures (subitem 12a), Funding (subitem 12b), Ethics (subitem 12c), Protocol (subitem 12d), and Data Availability (subitem 12e). The CHART checklist and corresponding methodological diagram were designed to support key stakeholders including clinicians, researchers, editors, peer reviewers, and readers in reporting, understanding, and interpreting the findings of CHA studies.
Bright Huo, Gary S. Collins, David Chartash, Arun Thirunavukarasu, Annette Flanagin, Alfonso Iorio, Giovanni E. Cacciamani, Nan Liu 0003, Piyush Mathur, An-Wen Chan, Christine Laine, Daniela Pacella, Michael Berkwits, Stavros A. Antoniou, Jennifer C. Camaradou, Carolyn Canfield, Michael Mittelman, Timothy Feeney, Elizabeth Loder, Riaz Agha, Ashirbani Saha, Julio Mayol, Anthony Sunjaya, Hugh Harvey, Jeremy Y. Ng, Tyler McKechnie, Yung Lee, Nipun Verma, Gregor Stiglic, Melissa McCradden, Karim Ramji, Vanessa Boudreau, Monica Ortenzi, Joerg Meerpohl, Per Olav Vandvik, Thomas Agoritsas, Diana Samuel, Helen Frankish, Xiaomei Yao, Stacy Loeb, Cynthia Lokker, Eliseo Guallar, Gordon Henry Guyatt
Artif. Intell. Medicine9
2025 Enabling inclusive systematic reviews: incorporating preprint articles with large language model-driven evaluations
abstract
OBJECTIVES: Systematic reviews in comparative effectiveness research require timely evidence synthesis. With the rapid advancement of medical research, preprint articles play an increasingly important role in accelerating knowledge dissemination. However, as preprint articles are not peer-reviewed before publication, their quality varies significantly, posing challenges for evidence inclusion in systematic reviews. MATERIALS AND METHODS: We developed AutoConfidenceScore (automated confidence score assessment), an advanced framework for predicting preprint publication, which reduces reliance on manual curation and expands the range of predictors, including three key advancements: (1) automated data extraction using natural language processing techniques, (2) semantic embeddings of titles and abstracts, and (3) large language model (LLM)-driven evaluation scores. Additionally, we employed two prediction models: a random forest classifier for binary outcome and a survival cure model that predicts both binary outcome and publication risk over time. RESULTS: The random forest classifier achieved an area under the receiver operating characteristic curve (AUROC) of 0.747 using all features. The survival cure model achieved an AUROC of 0.731 for binary outcome prediction and a concordance index of 0.667 for time-to-publication risk. DISCUSSION: Our study advances the framework for preprint publication prediction through automated data extraction and multiple feature integration. By combining semantic embeddings with LLM-driven evaluations, AutoConfidenceScore significantly enhances predictive performance while reducing manual annotation burden. CONCLUSION: AutoConfidenceScore has the potential to facilitate incorporation of preprint articles during the appraisal phase of systematic reviews, supporting researchers in more effective utilization of preprint resources.
Rui Yang 0016, Jiayi Tong, Nan Liu 0003, Christopher J. Lindsell, Michael J. Pencina, Yong Chen 0016, Chuan Hong
J. Am. Medical Informatics Assoc.7
2025 FedIMPUTE: Privacy-preserving missing value imputation for multi-site heterogeneous electronic health records
Siqi Li 0004, Mengying Yan, Ruizhi Yuan, Molei Liu, Nan Liu 0003, Chuan Hong
J. Biomed. Informatics5
2024 Clinical domain knowledge-derived template improves post hoc AI explanations in pneumothorax classification
Chuan Hong, Pengtao Jiang, Gangming Zhao, Nguyen Tuan Anh Tran, Xinxing Xu, Yet Yen Yan, Nan Liu 0003
J. Biomed. Informatics8
2023 Handling missing values in healthcare data: A systematic review of deep learning-based imputation techniques
Mingxuan Liu 0005, Siqi Li 0004, Marcus Eng Hock Ong, Yilin Ning, Feng Xie 0004, Seyed Ehsan Saffari, Yuqing Shang, Victor Volovici, Bibhas Chakraborty, Nan Liu 0003
Artif. Intell. Medicine11
2023 Federated and distributed learning applications for electronic health records and structured medical data: a scoping review
abstract
OBJECTIVES: Federated learning (FL) has gained popularity in clinical research in recent years to facilitate privacy-preserving collaboration. Structured data, one of the most prevalent forms of clinical data, has experienced significant growth in volume concurrently, notably with the widespread adoption of electronic health records in clinical practice. This review examines FL applications on structured medical data, identifies contemporary limitations, and discusses potential innovations. MATERIALS AND METHODS: We searched 5 databases, SCOPUS, MEDLINE, Web of Science, Embase, and CINAHL, to identify articles that applied FL to structured medical data and reported results following the PRISMA guidelines. Each selected publication was evaluated from 3 primary perspectives, including data quality, modeling strategies, and FL frameworks. RESULTS: Out of the 1193 papers screened, 34 met the inclusion criteria, with each article consisting of one or more studies that used FL to handle structured clinical/medical data. Of these, 24 utilized data acquired from electronic health records, with clinical predictions and association studies being the most common clinical research tasks that FL was applied to. Only one article exclusively explored the vertical FL setting, while the remaining 33 explored the horizontal FL setting, with only 14 discussing comparisons between single-site (local) and FL (global) analysis. CONCLUSIONS: The existing FL applications on structured medical data lack sufficient evaluations of clinically meaningful benefits, particularly when compared to single-site analyses. Therefore, it is crucial for future FL applications to prioritize clinical motivations and develop designs and methodologies that can effectively support and aid clinical practice and research.
Siqi Li 0004, Pinyan Liu, Gustavo G. Nascimento, Fabio Renato Manzolli Leite, Bibhas Chakraborty, Chuan Hong, Yilin Ning, Feng Xie 0004, Zhen Ling Teo, Daniel S. W. Ting, Hamed Haddadi 0001, Marcus Eng Hock Ong, Marco Aurélio Peres, Nan Liu 0003
J. Am. Medical Informatics Assoc.15
2023 A scoping review of the clinical application of machine learning in data-driven population segmentation analysis
abstract
OBJECTIVE: Data-driven population segmentation is commonly used in clinical settings to separate the heterogeneous population into multiple relatively homogenous groups with similar healthcare features. In recent years, machine learning (ML) based segmentation algorithms have garnered interest for their potential to speed up and improve algorithm development across many phenotypes and healthcare situations. This study evaluates ML-based segmentation with respect to (1) the populations applied, (2) the segmentation details, and (3) the outcome evaluations. MATERIALS AND METHODS: MEDLINE, Embase, Web of Science, and Scopus were used following the PRISMA-ScR criteria. Peer-reviewed studies in the English language that used data-driven population segmentation analysis on structured data from January 2000 to October 2022 were included. RESULTS: We identified 6077 articles and included 79 for the final analysis. Data-driven population segmentation analysis was employed in various clinical settings. K-means clustering is the most prevalent unsupervised ML paradigm. The most common settings were healthcare institutions. The most common targeted population was the general population. DISCUSSION: Although all the studies did internal validation, only 11 papers (13.9%) did external validation, and 23 papers (29.1%) conducted methods comparison. The existing papers discussed little validating the robustness of ML modeling. CONCLUSION: Existing ML applications on population segmentation need more evaluations regarding giving tailored, efficient integrated healthcare solutions compared to traditional segmentation analysis. Future ML applications in the field should emphasize methods' comparisons and external validation and investigate approaches to evaluate individual consistency using different methods.
Pinyan Liu, Nan Liu 0003, Marco Aurélio Peres
J. Am. Medical Informatics Assoc.3
2023 FedScore: A privacy-preserving framework for federated scoring system development
Siqi Li 0004, Yilin Ning, Marcus Eng Hock Ong, Bibhas Chakraborty, Chuan Hong, Feng Xie 0004, Mingxuan Liu 0005, Daniel M. Buckland, Yong Chen 0016, Nan Liu 0003
J. Biomed. Informatics11
2023 Epicasting: An Ensemble Wavelet Neural Network for forecasting epidemics
Madhurima Panja, Tanujit Chakraborty, Uttam Kumar 0001, Nan Liu 0003
Neural Networks4
2022 A Novel Interpretable Machine Learning System to Generate Clinical Risk Scores: An Application for Predicting Early Mortality or Unplanned Readmission in A Retrospective Cohort Study
Yilin Ning, Siqi Li 0004, Marcus Eng Hock Ong, Feng Xie 0004, Bibhas Chakraborty, Daniel S. W. Ting, Nan Liu 0003
AMIA7
2022 AutoScore-Ordinal: An Interpretable Machine Learning Framework for Generating Scoring Models for Ordinal Outcomes
Seyed Ehsan Saffari, Yilin Ning, Feng Xie 0004, Bibhas Chakraborty, Victor Volovici, Roger Vaughan, Marcus Eng Hock Ong, Nan Liu 0003
AMIA8
2022 Benchmarking Emergency Department Triage Prediction Models with Machine Learning and Large Public Electronic Health Records
Feng Xie 0004, Jun Zhou 0014, Jin Wee Lee, Mingrui Tan, Siqi Li 0004, Logasan S/O Rajnthern, Marcel Lucas Chee, Bibhas Chakraborty, An-Kwok Ian Wong, Alon Dagan, Marcus Eng Hock Ong, Nan Liu 0003
AMIA13
2022 AutoScore-Survival: Developing interpretable machine learning-based time-to-event scores with right-censored survival data
Feng Xie 0004, Yilin Ning, Benjamin Goldstein 0001, Marcus Eng Hock Ong, Nan Liu 0003, Bibhas Chakraborty
J. Biomed. Informatics6
2022 Deep learning for temporal data representation in electronic health records: A systematic review of challenges and methodologies
Feng Xie 0004, Yilin Ning, Marcus Eng Hock Ong, Mengling Feng, Wynne Hsu, Bibhas Chakraborty, Nan Liu 0003
J. Biomed. Informatics8
2022 AutoScore-Imbalance: An interpretable machine learning tool for development of clinical scores with rare events data
Feng Xie 0004, Marcus Eng Hock Ong, Yilin Ning, Marcel Lucas Chee, Seyed Ehsan Saffari, Hairil Rizal Abdullah, Benjamin Goldstein 0001, Bibhas Chakraborty, Nan Liu 0003
J. Biomed. Informatics10
2021 Development and Validation of a Survival Score for the Emergency Department in Singapore
Feng Xie 0004, Bibhas Chakraborty, Nan Liu 0003, Marcus Eng Hock Ong
AMIA3
2019 Serial Heart Rate Variability Measures for Risk Prediction of Septic Patients in the Emergency Department
Calvin Chiew, Han Wang 0001, Marcus Eng Hock Ong, Ting Hway Wong, Zhixiong Koh, Nan Liu 0003, Mengling Feng
AMIA6
2018 Development of a Radiology Decision Support System for the Classification of MRI Brain Scans
abstract
Previous studies revealed that the ordering of Magnetic resonance imaging (MRI) brain scans following American College of Radiology (ACR) guidelines showed a higher percentage of brain abnormalities compared to scans that do not. As the process of manually labelling patient orders obtained from a local tertiary hospital in accordance to ACR guidelines is intensive and time consuming, this study aims to develop predictive machine learning models; Logistic Regression (LR), Support Vector Machine (SVM), Random Forest (RF) and XGBoost (XGB), to automate the classification process through text mining methods and derive insights that are useful for future clinical decision-making and resource optimization. Using 1,924 observations as the labelled training data, RF and XGB were found to be the best performing robust models with ROC values of 0.9459 and 0.9508 respectively on the validation set (481 observations). Further exploration into the interpretability of black-box algorithms using the model agnostic LIME (Local Interpretable Model-Agnostic Explanations) framework was used to generate further insights for decisions made using a separate XGB model with respect to individual patients. The LIME framework is a significant first step towards the development of a comprehensive decision support system for patient-level decisions in the ordering of MRI scans.
Alwin Yaoxian Zhang, Sean Shao Wei Lam, Nan Liu 0003, Ling Ling Chan, Phua Hwee Tang
BDCAT3
2017 Extreme learning machine based mutual information estimation with application to time-series change-points detection
Beom-Seok Oh, Lei Sun 0006, Chung Soo Ahn, Yong Kiang Yeo, Nan Liu 0003, Zhiping Lin 0001
Neurocomputing6
2015 Effects of two new features of approximate entropy and sample entropy on cardiac arrest prediction
abstract
Sixteen conventional heart beat variability (HRV) parameters and eight vital signs have shown promise in the prediction of cardiac arrest within 72 hours. Besides these 24 parameters, we proposed adding two new features for cardiac arrest prediction, which are approximate entropy (ApEn) and sample entropy (SpEn). ApEn and SpEn are nonlinear HRV parameters capable of characterizing heart conditions. These two entropies were derived from electrocardiography recordings and combined with the existing 24 features to form feature combinations. The experiments were conducted by using linear kernel Support Vector Machine classification technique to investigate the effects of using ApEn, SpEn together with 24 parameters on cardiac arrest prediction. The dimensionality reduction approach, Principal Component Analysis, was applied to suppress the dimensionality. Results reveal that the prediction performance of adding ApEn and SpEn to the 24 parameters is improved significantly compared to using the 24 parameters only. Dimension reduction has additional positive effects on improving the prediction results.
Yumeng Gao, Zhiping Lin 0001, Tongtong Zhang, Nan Liu 0003, Tianchi Liu 0001, Wee Ser, Zhixiong Koh, Marcus Eng Hock Ong
ISCAS4
2015 Ensemble of subset online sequential extreme learning machine for class imbalance and concept drift
Bilal Mirza, Zhiping Lin 0001, Nan Liu 0003
Neurocomputing3
2014 Risk Scoring for Prediction of Acute Cardiac Complications from Imbalanced Clinical Data
abstract
Fast and accurate risk stratification is essential in the emergency department (ED) as it allows clinicians to identify chest pain patients who are at high risk of cardiac complications and require intensive monitoring and early intervention. In this paper, we present a novel intelligent scoring system using heart rate variability, 12-lead electrocardiogram (ECG), and vital signs where a hybrid sampling-based ensemble learning strategy is proposed to handle data imbalance. The experiments were conducted on a dataset consisting of 564 chest pain patients recruited at the ED of a tertiary hospital. The proposed ensemble-based scoring system was compared with established scoring methods such as the modified early warning score and the thrombolysis in myocardial infarction score, and showed its effectiveness in predicting acute cardiac complications within 72 h in terms of the receiver operation characteristic analysis.
Nan Liu 0003, Zhixiong Koh, Eric Chern-Pin Chua, Licia Mei-Ling Tan, Zhiping Lin 0001, Bilal Mirza, Marcus Eng Hock Ong
IEEE J. Biomed. Health Informatics1
2012 Voting based extreme learning machine
Jiuwen Cao, Zhiping Lin 0001, Guang-Bin Huang, Nan Liu 0003
Inf. Sci.4
2012 An Intelligent Scoring System and Its Application to Cardiac Arrest Prediction
abstract
Traditional risk score prediction is based on vital signs and clinical assessment. In this paper, we present an intelligent scoring system for the prediction of cardiac arrest within 72 h. The patient population is represented by a set of feature vectors, from which risk scores are derived based on geometric distance calculation and support vector machine. Each feature vector is a combination of heart rate variability (HRV) parameters and vital signs. Performance evaluation is conducted on the leave-one-out cross-validation framework, and receiver operating characteristic, sensitivity, specificity, positive predictive value, and negative predictive value are reported. Experimental results reveal that the proposed scoring system not only achieves satisfactory performance on determining the risk of cardiac arrest within 72 h but also has the ability to generate continuous risk scores rather than a simple binary decision by a traditional classifier. Furthermore, the proposed scoring system works well for both balanced and imbalanced datasets, and the combination of HRV parameters and vital signs shows superiority in prediction to using HRV parameters only or vital signs only.
Nan Liu 0003, Zhiping Lin 0001, Jiuwen Cao, Zhixiong Koh, Tongtong Zhang, Guang-Bin Huang, Wee Ser, Marcus Eng Hock Ong
IEEE Trans. Inf. Technol. Biomed.1
2010 Ensemble Based Extreme Learning Machine
abstract
Extreme learning machine (ELM) was proposed as a new class of learning algorithm for single-hidden layer feedforward neural network (SLFN). To achieve good generalization performance, ELM minimizes training error on the entire training data set, therefore it might suffer from overfitting as the learning model will approximate all training samples well. In this letter, an ensemble based ELM (EN-ELM) algorithm is proposed where ensemble learning and cross-validation are embedded into the training phase so as to alleviate the overtraining problem and enhance the predictive stability. Experimental results on several benchmark databases demonstrate that EN-ELM is robust and efficient for classification.
Nan Liu 0003, Han Wang 0001
IEEE Signal Process. Lett.1
2009 Modeling Images With Multiple Trace Transforms for Pattern Analysis
abstract
Taking advantage of the various available trace transforms generated from a single image, the multiple trace feature (MTF) is proposed as a new image representation. In the process of MTF construction, genetic algorithms (GAs) play a key role as an information fusion tool. The systematic evaluations on a combo face data set comprising ORL, Yale, and UMIST databases reveal that MTF presents high discriminative power in terms of outperforming features extracted from principal component analysis (PCA) and linear discriminant analysis (LDA). In addition, the proposed Bagging-based extension of fitness guides GAs achieving more fitting features for classification.
Nan Liu 0003, Han Wang 0001
IEEE Signal Process. Lett.1
2008 Feature selection in frequency domain and its application to face recognition
abstract
Face recognition system usually consists of components of feature extraction and pattern classification. However, not all of extracted facial features contribute to the classification phase positively because of the variations of illumination and poses in face images. In this paper, a three-step feature selection algorithm is proposed in which discrete cosine transform (DCT) and genetic algorithms (GAs) as well as dimensionality reduction methods are utilized to create a combined framework of feature acquisition. In details, the face images are first transformed to frequency domain through DCT. Then GAs are used to seek for optimal features in the redundant DCT coefficients where the generalization performance guides the searching process. The last step is to reduce the dimension of selected features. In experiments, two face databases are used to evaluate the effectiveness of the proposed method. In addition, an entropy-based improvement is also proposed. The experimental results present the superiority of selected frequency features.
Nan Liu 0003, Han Wang 0001
IJCNN1
2007 Extraction of hybrid trace features with evolutionary computation for face recognition
abstract
The Hybrid Trace Features (HTF), a new face representation, is proposed for face authentication system. Trace transforms of multiple Trace functionals are used to construct the HTF, and Genetic Algorithms is implemented as the data fusion tool. In addition, rotation based Hybrid Trace Features (r-HTF) is also introduced as facial features. The systemic evaluations on Cambridge ORL face database reveal that HTF and r-HTF present high discriminatory power and outperform the features extracted by Principal Component Analysis (PCA), Linear Discriminant Analysis (LDA) and Kernel PCA in the task of classification.
Nan Liu 0003, Han Wang 0001
IEEE Congress on Evolutionary Computation1
2007 Classification of transformed face images with majority voting
abstract
Trace transform is an alternative representation of original image, and various transforms can be obtained by implementing different Trace functionals. Taking the advantages of diversified functionals, ensemble Trace transforms based face recognition system is proposed in which majority voting plays the role of decision making. In addition, weighted majority voting and genetic algorithms are combined to achieve better generalization performance. Experiments on combo face database consisting of ORL, Yale, UMIST data sets reveal the power of proposed framework in pattern analysis.
Nan Liu 0003, Han Wang 0001
SMC1
2006 Face Recognition with Weighted Kernel Principal Component Analysis
abstract
Principal component analysis (PCA) is one of the most traditional linear dimensionality reduction algorithms. Kernel principal component analysis (kernel PCA), generalization of PCA, is a nonlinear feature extraction method. However, both PCA and kernel PCA are lack of class information in their feature subspace. In this paper, weighted kernel principal component analysis (WKPCA) is proposed for feature extraction with the application of face recognition. Weights that represent inter-class relationships are incorporated into kernel matrix. Images in training and testing set are projected onto the subspace obtained from weighted kernel matrix. The feature extraction procedure is in a framework of genetic algorithms (GAs) with the fitness as classification accuracy on cross-validation data from training set. The experimental results of WKPCA are compared with PCA and kernel PCA on a combo database (ORL, Yale, UMIST databases), and show that proposed WKPCA algorithm performs best in face recognition
Nan Liu 0003, Han Wang 0001, Weiyun Yau
ICARCV1
2005 Feature extraction using evolutionary weighted principal component analysis
abstract
Principal component analysis (PCA) and Fisher's linear discriminant (FLD) are two commonly used feature extraction techniques. Based on them, an evolutionary weighted principal component analysis (EWPCA) is proposed. Similar to FLD, the proposed EWPCA maximizes the ratio of between-class scatter to that of within-class scatter, while keeps even smaller reconstruction error than that of traditional PCA. Genetic algorithms (GAs) are chosen as the searching method to select optimal weights for the EWPCA. In the face recognition application, Evolutionary Eigenface obtained by performing EWPCA, is used as the representation of original face images. Our experimental results prove that EWPCA outperforms both PCA and FLD. Besides, Evolutionary Cosineface is also proposed, which creates better classification performance than most reported approaches on ORLface database.
Nan Liu 0003, Han Wang 0001
SMC1