Tayo Obafemi-Ajayi

dblp:34/3945 · DBLP profile ↗
← Back
28ranked-venue papers
7as first author
10since 2021 · last 2025
0000-0002-0155-9733ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 16 · 2 first-author · 8 since 2021Artificial intelligence and machine learning · 8 · 3 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-authorHuman-computer interaction and ubiquitous computing · 2 · 1 first-author
YearPublicationVenuePosition
2025 Balanced Benchmarking of Zero-Shot and RAG Approaches for Biomedical Term Normalization
abstract
Normalization of medical concepts to an ontology is a key aspect of the natural language processing of biomedical text. It enables the mapping of medical expressions to standardized ontology terms and their identifiers, thereby enhancing the interoperability and computability of medical concepts. Although large language models (LLMs) can identify and standardize medical terms, they may struggle to accurately map ontology terms to their corresponding ontology identifiers. These challenges arise from the stochastic nature of LLMs, their limited exposure to uncommon ontology identifiers during training, and their lack of an integrated lookup mechanism. We generated test sets of synthetic terms to assess normalization performance by both zero-shot prompted and retrieval-augmented generation (RAG) prompted methods across two ontologies (Human Phenotype Ontology and Gene Ontology) and three LLMs (GPT-4o, LLaMA 3.3 70B, and Phi-4). To ensure a calibrated and fair evaluation of normalization, the test set was balanced along two axes: (1) term prevalence in biomedical literature, as estimated by PubMed Central frequency counts, and (2) semantic proximity to ontology terms, as assessed by cosine similarity of BioBERT embeddings. Our results demonstrate that RAG consistently outperforms zero-shot prompting, particularly on low-prevalence terms that are infrequently encountered in the biomedical literature. This highlights the value of RAG in compensating for gaps in model exposure to uncommon medical concepts. We demonstrate that a synthetic test set can be a valuable tool for evaluating biomedical term normalization across LLMs.
Thanh Son Do, Daniel B. Hier, Tayo Obafemi-Ajayi
CIBCB3
2025 Ethics vs. Regulation: Converging Frameworks for Trustworthy Human-Centered AI in Biomedical Research
abstract
The accelerating impact of AI in biomedical research is driving significant advances in precision medicine. As these systems increasingly shape health outcomes, the imperative to develop trustworthy, reliable, and ethically grounded AI becomes more pressing, particularly in addressing concerns related to data integrity, patient safety, and equitable outcomes. While the potential of AI to transform biomedical research is clear, its responsible integration depends on more than technological capability. Ensuring that these systems are aligned with societal values requires a dual commitment: the operationalization of ethical principles throughout the AI life cycle and the establishment of robust regulatory mechanisms. Ethics provides the normative vision for fairness, accountability, and human dignity, whereas regulation translates these ideals into enforceable standards. This paper explores the convergence of these domains as a necessary foundation for developing trustworthy human-centered AI in biomedical contexts. We provide practical guidance for AI developers and researchers on integrating proactive governance and translating ethical principles into actionable strategies to support equitable and responsible innovation.
Tayo Obafemi-Ajayi, Tiffani J. Bright, Emily F. Wong, Donald C. Wunsch II, Joan Peckham, Jason H. Moore
IJCNN1
2025 Guest Editorial: Deep Medicine and AI for Health
María Fernanda Cabrera-Umpiérrez, Tayo Obafemi-Ajayi, Ahmed Metwally 0002, Bobak Mortazavi
IEEE J. Biomed. Health Informatics2
2025 Guest Editorial: Precision Health: AI Tailored to Individuals
Edward Sazonov, Bobak Mortazavi, Tayo Obafemi-Ajayi, Hassan Ghasemzadeh 0001, María Fernanda Cabrera-Umpiérrez, May D. Wang
IEEE J. Biomed. Health Informatics3
2024 Towards Explainability of Dimension Reduction Plots of Unsupervised Learning Model Outcomes
abstract
Dimension reduction methods are used to visualize the output of unsupervised learning models when applied to complex data. These techniques improve interpretability by transforming a high-dimension space to a lower-dimension space (usually 2D or 3D). The results are typically viewed as 2D scatter plots, and class centroids may be added to increase interpretability. Although useful, the relationship of these class centroids to the underlying feature space remains opaque. The innovative aspect of this work is to create a strong link between the dimension-reduced space and the underlying high-dimension feature space by adding selected feature centroids to the 2D scatter plots. This approach simultaneously visualizes the centers for the classes and the features on the same 2D scatter plot. Since classes are often imbalanced, we provide a method to balance class sizes. We present an automated framework that performs a grid search to find the optimal dimension reduction parameters, balances the class sizes, uses an ensemble approach to find the most important features, and adds class centroids and selected feature centroids to 2D dimension-reduced plots. This is especially useful when applied to complex, feature-rich biomedical data, as addition of feature centroids to 2D scatter plots serve as landmarks for the previously featureless dimension-reduced space. The utility of this approach is demonstrated by its application to seven classes of neurogenetic diseases with 31 defining phenotypic features.
Tony E. Astuhuaman Davila, Daniel B. Hier, Tayo Obafemi-Ajayi
CIBCB3
2024 Evaluation of Transfer Learning Models on Traumatic Brain Injury Severity Classification
abstract
After traumatic brain injury (TBI), clinicians use the Glasgow Coma Scale (GCS) to classify patients by severity and radiologists use the Rotterdam score and the Marshall score to classify CT scans by severity. This work investigates a viable efficient low cost transfer learning model to use MRI images to predict GCS, Rotterdam, and Marshall severity class after TBI. The enhanced transfer learning model architecture integrates multiple fine-tuning steps in which a few layers of the pre-trained model is unfrozen in each iteration so that the neural network is able to learn layer by layer further aspects of the image for better classification. By conducting a thorough evaluation across multiple convolutional neural network (CNN) architectures, this work ascertains the sensitivity of CNN models in detecting anatomical changes presented in MRI images that are predictive of severity class after TBI. These models have potential to predict outcomes after TBI. We utilize both quantitative metrics and qualitative analysis to validate the clinical relevance of the models in predicting severity class on admission and outcome at 6 months. The residual network models (ResNet152 in particular) outperformed the other models in predicting initial severity class.
Ngoc Do, Daniel B. Hier, Tayo Obafemi-Ajayi
CIBCB3
2023 An Explainable Deep Learning Model for Prediction of Severity of Alzheimer's Disease
abstract
Deep Convolutional Neural Networks (CNNs) have become the go-to method for medical imaging classification on various imaging modalities for binary and multiclass problems. Deep CNNs extract spatial features from image data hierarchically, with deeper layers learning more relevant features for the classification application. Despite the high predictive accuracy, usability lags in practical applications due to the black-box model perception. Model explainability and interpretability are essential for successfully integrating artificial intelligence into healthcare practice. This work addresses the challenge of an explainable deep learning model for the prediction of the severity of Alzheimer’s disease (AD). AD diagnosis and prognosis heavily rely on neuroimaging information, particularly magnetic resonance imaging (MRI). We present a deep learning model framework that integrates a local data-driven interpretation method that explains the relationship between the predicted AD severity from the CNN and the input MR brain image. The deep explainer uses SHapley Additive exPlanation values to quantity the contribution of different brain regions utilized by the CNN to predict outcomes. We conduct a comparative analysis of three high-performing CNN models: DenseNet121, DenseNet169, and Inception-ResNet-v2. The framework shows high sensitivity and specificity in the test sample of subjects with varying levels of AD severity. We also correlated five key AD neurocognitive assessment outcome measures and the APOE genotype biomarker with model misclassifications to facilitate a better understanding of model performance.
Godwin Ekuma, Daniel B. Hier, Tayo Obafemi-Ajayi
CIBCB3
2022 Heterogeneity in Blood Biomarker Trajectories After Mild TBI Revealed by Unsupervised Learning
abstract
Concussions, also known as mild traumatic brain injury (mTBI), are a growing health challenge. Approximately four million concussions are diagnosed annually in the United States. Concussion is a heterogeneous disorder in causation, symptoms, and outcome making precision medicine approaches to this disorder important. Persistent disabling symptoms sometimes delay recovery in a difficult to predict subset of mTBI patients. Despite abundant data, clinicians need better tools to assess and predict recovery. Data-driven decision support holds promise for accurate clinical prediction tools for mTBI due to its ability to identify hidden correlations in complex datasets. We apply a Locality-Sensitive Hashing model enhanced by varied statistical methods to cluster blood biomarker level trajectories acquired over multiple time points. Additional features derived from demographics, injury context, neurocognitive assessment, and postural stability assessment are extracted using an autoencoder to augment the model. The data, obtained from FITBIR, consisted of 301 concussed subjects (athletes and cadets). Clustering identified 11 different biomarker trajectories. Two of the trajectories (rising GFAP and rising NF-L) were associated with a greater risk of loss of consciousness or post-traumatic amnesia at onset. The ability to cluster blood biomarker trajectories enhances the possibilities for precision medicine approaches to mTBI.
Lien A. Bui, Dacosta Yeboah, Louis Steinmeister, Sima Azizi, Daniel B. Hier, Donald C. Wunsch II, Gayla R. Olbricht, Tayo Obafemi-Ajayi
IEEE ACM Trans. Comput. Biol. Bioinform.8
2021 A deep learning model to predict traumatic brain injury severity and outcome from MR images
abstract
For many neurological disorders, including traumatic brain injury (TBI), neuroimaging information plays a crucial role determining diagnosis and prognosis. TBI is a heterogeneous disorder that can result in lasting physical, emotional and cognitive impairments. Magnetic Resonance Imaging (MRI) is a non-invasive technique that uses radio waves to reveal fine details of brain anatomy and pathology. Although MRIs are interpreted by radiologists, advances are being made in the use of deep learning for MRI interpretation. This work evaluates a deep learning model based on a residual learning convolutional neural network that predicts TBI severity from MR images. The model achieved a high sensitivity and specificity on the test sample of subjects with varying levels of TBI severity. Six outcome measures were available on TBI subjects at 6 and 12 months. Group comparisons of outcomes between subjects correctly classified by the model with subjects misclassified suggested that the neural network may be able to identify latent predictive information from the MR images not incorporated in the ground truth labels. The residual learning model shows promise in the classification of MR images from subjects with TBI.
Dacosta Yeboah, Daniel B. Hier, Gayla R. Olbricht, Tayo Obafemi-Ajayi
CIBCB5
2021 Noise Quality and Super-Turing Computation in Recurrent Neural Networks
Emmett Redd, Tayo Obafemi-Ajayi
ICANN (4)2
2019 Comparative Analysis of Feature Selection Methods to Identify Biomarkers in a Stroke-Related Dataset
abstract
This paper applies machine learning feature selection techniques to the REGARDS stroke-related dataset to identify health-related biomarkers. A data-driven methodological framework is presented to evaluate multiple feature selection methods. In applying the framework, three classifiers are chosen in conjunction with two wrappers, and their performance with diverse classification targets such as Current Smoker, Current Alcohol Use, and Deceased is evaluated. The performance across logistic regression, random forest and naïve Bayes classifier methods, as quantified by the ROC Area Under Curve metric and selected features, was similar. However, significant differences were observed in running time. Performance of the selected features was also evaluated based on the accuracy of a prediction model generated using a multi-layer perceptron (MLP) classifier.
Thomas Clifford, Justin Bruce, Tayo Obafemi-Ajayi, John Matta
CIBCB3
2019 Multi-objective Optimization Approach to find Biclusters in Gene Expression Data
abstract
Gene expression levels of organisms are measured by DNA microarrays. Finding biclusters in gene expression matrices provides invaluable information about effects of disease at the genetic level. These biclusters could identify which genes are up-regulated/down-regulated under certain conditions. This paper investigates a methodology for evolutionary-based biclustering using the NSGA-II algorithm. It also presents an improvement to the recovery and relevance external validation metrics as well as a new method for synthetic data generation for biclustering. Results obtained demonstrate its effectiveness in discovering useful biclusters on varied synthetic data when applied with the average Spearman's rho measure as the fitness function.
Jeffrey Dale, Junya Zhao, Tayo Obafemi-Ajayi
CIBCB3
2019 Genotype Combinations Linked to Phenotype Subgroups in Autism Spectrum Disorders
abstract
This paper investigates a computational model that allows for systematic comparison of phenotype data with genotype (Single Nucleotide Polymorphisms (SNPs)) data based on machine learning techniques to identify discriminant genotype markers associated with the phenotypic subgroups. The proposed discriminant SNP identifier model is empirically evaluated using Autism Spectrum Disorder (ASD) simplex sample. Six phenotype markers were selected to cluster the sample in a hexagonal lattice format yielding five multidimensional subgroups based on extremities of the phenotype markers. The SNP selection model includes random subspace selection of SNPs in conjunction with feature selection algorithms to determine which set of SNPs were discriminant among these five subgroups. This yielded a set of SNPs that attained a mean ROC performance of 95% using a Support Vector Machine prediction model. Biological analysis of these SNPs and associated genes across the subgroups is presented to examine their potential clinical significance.
Junya Zhao, Thy Nguyen, Jonathan Kopel, Perry B. Koob, Donald A. Adieroh, Tayo Obafemi-Ajayi
CIBCB6
2019 Recurrent Network and Multi-arm Bandit Methods for Multi-task Learning without Task Specification
abstract
This paper addresses the problem of multi-task learning (MTL) in settings where the task assignment is not known. We propose two mechanisms for the problem of inference of task's parameter without task specification: parameter adaptation and parameter selection methods. In parameter adaptation, the model's parameter is iteratively updated using a recurrent neural network (RNN) learner as the mechanism to adapt to different tasks. For the parameter selection model, a parameter matrix is learned beforehand with the task known apriori. During testing, a bandit algorithm is utilized to determine the appropriate parameter vector for the model on the fly. We explored two different scenarios in MTL without task specification, continuous learning and reset learning. In continuous learning, the model has to adjust its parameter continuously to a number of different task without knowing when task changes. Whereas in reset learning, the parameter is reset to an initial value to aid transition to different tasks. Results on three real benchmark datasets demonstrate the comparative performance of both models with respect to multiple RNN configurations, MTL algorithms and bandit selection policies.
Thy Nguyen, Tayo Obafemi-Ajayi
IJCNN2
2019 Stochastic Resonance Enables BPP/log* Complexity and Universal Approximation in Analog Recurrent Neural Networks
abstract
Stochastic resonance (SR) is a natural process that without limit increases the precision of signal measurements in biological and physical sciences. Most artificial neural networks (NNs) are implemented on digital computers of fixed-precision. A NN accessing universal approximation and a computational complexity class more powerful that of a Turing machine needs analog signals utilizing SR’s limitless precision increase. This paper links an analog recurrent (AR) NN theorem, SR, BPP/log* (a physically realizable, super-Turing computation class), and universal approximation so NNs following them can be made computationally more powerful. An optical neural network mimicking chaos indicates super-Turing computation has been achieved. Additional tests are needed which can verify superTuring computation, show its superiority, and demonstrate its practical benefits. Truly powerful cognitively inspired computation needs to access the combination of ARNNs, SR, super-Turing mathematical complexity, and universal approximation.
Emmett Redd, Arthur Steven Younger, Tayo Obafemi-Ajayi
IJCNN3
2018 Random Subspace Projection for Predicting Biogeographical Ancestry
Tanjin Taher Toma, Tayo Obafemi-Ajayi, Jeremy M. Dawson, Donald A. Adjeroh
BIBM2
2018 Analysis of grapevine gene expression data using node-based resilience clustering
abstract
Powdery mildew is the most economically important disease of cultivated grapevines worldwide. In the agricultural community, there is a great need for better understanding of the complex genetic basis of powdery mildew (PM) resistance by delineating possible gene biomarkers associated with the plants' defense mechanisms. Machine learning techniques can be applied to analysis of gene expression data to aid knowledge discovery of disease fighting genes. In this work, we apply a data-driven computational model, utilizing a graph-based clustering algorithm - Node-Based Resilience Clustering (NBR- Clust), to analyze grapevine gene expression data to identify possible gene biomarkers associated with powdery mildew disease defense mechanisms. We investigated two graph representations (geometric and kNN) on the mean differences of PM inoculated vs. mock inoculated gene expression values of Cabernet and Norton (PM disease resistant) species across 6 time points. By applying the contrarian approach, we hypothesized that smaller sized clusters will contain genes that do not follow general patterns, hence, could display distinct expression patterns of PM- induced transcripts across the time points that may insinuate biological relevance. We compared the smaller clusters obtained in Norton in contrast with the ones from Cabernet in terms of the genes that clustered in common between both (intersection of sets) as well as the differences of the sets. The results obtained demonstrate the usefulness of the geometric graphs for this domain application in contrast to the kNN graphs. Some genes that belong to biologically relevant pathways were identified that displayed differences in patterns across the time points between Norton and Cabernet species.
Jeffrey Dale, John Matta, Susanne Howard, Gunes Ercal, Wenping Qiu, Tayo Obafemi-Ajayi
CIBCB6
2018 Ensemble validation paradigm for intelligent data analysis in autism spectrum disorders
abstract
Cluster analysis is an important exploratory tool for a broad range of applications including data analysis of biomedical datasets to uncover meaningful subgroups such as in autism spectrum disorder (ASD). For a given clustering algorithm, multiple results can be obtained on the same dataset by varying the algorithm parameters. In biomedical applications, discovering meaningful subgroups, not just the optimal number of clusters, is expedient. It is imperative to develop quality measures capable of identifying optimal partitions for a given dataset. In this paper, we apply varied clustering methods to subgroup an ASD simplex sample based on relevant phenotype features that may uncover meaningful subtypes. We present a detailed cluster validation analysis using an ensemble validation paradigm and visualization techniques. We present a rigorous clinical/behavioral analysis of the top highly ranked results. The evaluation demonstrated that both configurations yielded similar clinical significance results: 2-subgroups configuration with distinct clinical profile.
Thy Nguyen, Kerri Nowell, Kimberly E. Bodner, Tayo Obafemi-Ajayi
CIBCB4
2018 Performance Evaluation and Enhancement of Biclustering Algorithms
abstract
In gene expression data analysis, biclustering has proven to be an effective method of finding local patterns among subsets of genes and conditions. The task of evaluating the quality of a bicluster when ground truth is not known is challenging. In this analysis, we empirically evaluate and compare the performance of eight popular biclustering algorithms across 119 synthetic datasets that span a wide range of possible bicluster structures and patterns. We also present a method of enhancing performance (relevance score) of the biclustering algorithms to increase confidence in the significance of the biclusters returned based on four internal validation measures. The experimental results demonstrate that the Average Spearman’s Rho evaluation measure is the most effective criteria to improve bicluster relevance with the proposed performance enhancement method, while maintaining a relatively low loss in recovery scores.
Jeffrey Dale, America Nishimoto, Tayo Obafemi-Ajayi
ICPRAM3
2018 Meta-Learning Related Tasks with Recurrent Networks: Optimization and Generalization
abstract
There have been recent interest in meta-learning systems: i.e., networks that are trained to learn across multiple tasks. This paper focuses on optimization and generalization of a meta-learning system based on recurrent networks. The optimization investigates the influence of diverse structures and parameters on its performance. We demonstrate the generalization (robustness) of our meta-learning system to learn across multiple tasks including tasks unseen during the metatraining phase. We introduce a meta-cost function (Mean Squared Fair Error) that enhances the performance of the system by not penalizing it during transitions to learning a new task. Evaluation results are presented for Boolean and quadratic functions datasets. The best performance is obtained using a Long Short-Term Memory (LSTM) topology without a forget gate and with a clipped memory cell. The results demonstrate i) the impact of different LSTM architectures, parameters, and error functions on the meta-learning process; ii) that the mean squared fair error function does improve performance for best learning; and iii) the robustness of our meta-learning framework as it generalizes well when tested on tasks unseen during meta-training. Comparison between No-Forget-Gate LSTM and Gated Recurrent Unit also suggest that absence of a memory cell tends to degrade performance.
Thy Nguyen, Arthur Steven Younger, Emmett Redd, Tayo Obafemi-Ajayi
IJCNN4
2017 Genetic variant analysis of boys with Autism: A pilot study on linking facial phenotype to genotype
abstract
This work examines the validity of facial phenotypes as Autism Spectrum Disorders (ASD) biomarkers in boys with essential autism. A family-based association analysis framework is presented that uses previously identified facially-delineated (FD) clusters to examine relationship between FD clusters and known ASD genes. The hypothesis is that there are certain genetic variants, single nucleotide polymorphisms (SNP), specific to the FD clusters. Although statistical significance was not established, the results identified some candidate SNPs unique to each of the FD clusters that could indicate an underlying etiological difference. Further, recommendations are provided for larger-scale studies that could utilize the analysis framework presented.
Tayo Obafemi-Ajayi, Luke Settles, Yuqing Su, Cynthia Germeroth, Gayla R. Olbricht, Donald C. Wunsch II, T. Nicole Takahashi, Judith H. Miles
BIBM1
2017 Building the K-12 engineering pipeline: An assessment of where we stand
abstract
This paper is a survey-based assessment of the pre-engineering activities focused on building the engineering pipeline in the Ozarks region of Missouri. We assess the impact of diverse outreach programs including Project Lead the Way, Science Olympiad, the Ozarks sySTEAMic Coalition (O-STEAM) and other Science, Technology, Engineering and Math (STEM) programs/outreach events that have sought to engage the local schools and public at large. Evaluation was based on analysis of survey data collected from middle school and high school students (6-12th graders) in the local schools as well as students currently enrolled in the engineering program at Missouri State University (a cooperative joint program with Missouri University of Science and Technology). The broad and diverse spectrum of students feedback analyzed provided an insight into which activities are effectively building the pipeline and what things we might considering doing differently to strengthen the engineering pipeline. Though this survey is focused on the Springfield Public schools of the Ozarks Missouri region, the methodology and results are replicable to other regions in the US.
Theresa Odun-Ayo, Tayo Obafemi-Ajayi
FIE2
2016 Robust Graph-Theoretic Clustering Approaches Using Node-Based Resilience Measures
abstract
This paper examines a schema for graph-theoretic clustering using node-based resilience measures. Node-based resilience measures optimize an objective based on a critical set of nodes whose removal causes some severity of disconnection in the network. Beyond presenting a general framework for the usage of node based resilience measures for variations of clustering problems, we emphasize the unique potential of such methods to accomplish the following properties: (i) clustering a graph in one step without knowing the number of clusters a priori, and (ii) removing noise from noisy data. We first present results of clustering experiments using a β-parametrized generalization of vertex attack tolerance, showing high clustering accuracy for both real datasets and equal density synthetic data sets, as well as successful removal of noise nodes. It is shown that arbitrarily increasing β increases the number of noise nodes removed in some cases, and that internal validation measures can be used to determine the correct number of clusters in a class of datasets. Further results are presented using five different resilience measures with a general node-based resilience clustering technique. In a subset of cases a resilience measure, such as integrity, is able to cluster to high accuracy in one step, giving the correct clustering while also determining the correct number of clusters. Integrity is also shown to be promising with respect to noise removal, removing up to 80% of noise on some datasets.
John Matta, Tayo Obafemi-Ajayi, Jeffrey Borwey, Donald C. Wunsch II, Gunes Ercal
ICDM2
2015 Sorting the phenotypic heterogeneity of autism spectrum disorders: A hierarchical clustering model
abstract
Autism spectrum disorder (ASD) is characterized by notable phenotypic heterogeneity, which is often viewed as an obstacle to the study of its etiology, diagnosis, treatment, and prognosis. Heterogeneity in ASD is multidimensional and complex including variability in phenotype as well as clinical, physiologic, and pathologic parameters. We apply a hierarchical clustering model suited to dealing with datasets of mixed data types to stratify children with ASD into more homogeneous subgroups in line with the Diagnostic and Statistical Manual of Mental Disorders (DSM)-5 model. The results of this cluster analysis will provide a better understanding the complex issue of ASD phenotypic heterogeneity and identify subgroups useful for further ASD genetic studies. Our goal is to provide insight into viable phenotypic and genotypic markers that would guide further cluster analysis of ASD genetic data. We suggest that analyzing the clusters in a hierarchical structure is a well-suited and meaningful model to unravel the complex heterogeneity of this disorder.
Tayo Obafemi-Ajayi, Dao Lam, T. Nicole Takahashi, Stephen Kanne, Donald C. Wunsch II
CIBCB1
2012 Cluster-K+: Network topology for searching replicated data in p2p systems
Tayo Obafemi-Ajayi, Sanjiv Kapoor, Ophir Frieder
Inf. Process. Manag.1
2012 Character-Based Automated Human Perception Quality Assessment in Document Images
abstract
Large degradations in document images impede their readability and deteriorate the performance of automated document processing systems. Document image quality (IQ) metrics have been defined through optical character recognition (OCR) accuracy. Such metrics, however, do not always correlate with human perception of IQ. When enhancing document images with the goal of improving readability, e.g., in historical documents where OCR performance is low and/or where it is necessary to preserve the original context, it is important to understand human perception of quality. The goal of this paper is to design a system that enables the learning and estimation of human perception of document IQ. Such a metric can be used to compare existing document enhancement methods and guide automated document enhancement. Moreover, the proposed methodology is designed as a general framework that can be applied in a wide range of applications.
Tayo Obafemi-Ajayi, Gady Agam
IEEE Trans. Syst. Man Cybern. Part A1
2010 Historical document enhancement using LUT classification
Tayo Obafemi-Ajayi, Gady Agam, Ophir Frieder
Int. J. Document Anal. Recognit.1
2008 Efficient MRF approach to document image enhancement
abstract
Markov random field (MRF) based approaches have been shown to perform well in a wide range of applications. Due to the iterative nature of the algorithm, the computational cost of such applications is normally high. In the context of document image analysis, where numerous documents have to be processed, this computational cost may become prohibitive. We describe a novel approach to document image enhancement using MRF.We show that by using domain specific knowledge, we are able to substantially improve computational performance by an order of magnitude. Moreover, in contrast to known techniques where patch initialization is arbitrary, in the proposed approach patch initialization is data consistent and so results in improved effectiveness. Experimental results comparing the proposed approach to known techniques using historical documents from the Frieder Collection are provided.
Tayo Obafemi-Ajayi, Gady Agam, Ophir Frieder
ICPR1