VLDB 2026 Research / reviewers in the wild / expert
Donald A. Adjeroh
dblp:30/1579 · also Don Adjeroh, Donald Asogu Adjeroh
· DBLP profile ↗
95ranked-venue papers
12as first author
30since 2021 · last 2026
0000-0002-7982-4744ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 40 · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 35 · 10 first-author · 9 since 2021Artificial intelligence and machine learning · 17 · 2 first-author · 9 since 2021Databases, data management, data science and information retrieval · 11 · 4 first-author · 1 since 2021Theory of computation · 6Computer networks · 4Human-computer interaction and ubiquitous computing · 2Security and privacy · 1Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | iStructTab: Structured Feature Sequencing for Multimodal Learning of Image and Tabular Data
Al-Zadid Sultan Bin Habib, Md Younus Ahamed, Prashnna Gyawali, Gianfranco Doretto, Donald A. Adjeroh |
ICPR (8) | 5 |
| 2025 | Graph Neural Networks for LncRNA Subcellular LocalizationabstractLong non-coding RNAs (lncRNAs) play diverse roles across cellular compartments, yet predicting subcellular localization from the RNA sequence remains challenging. We present a new framework that casts each lncRNA sequence as a graph and trains graph neural networks (GNNs) to perform binary localization for a given lncRNA and cell line. We consider five approaches for graph construction: (1) normalized Pointwise Mutual Information (nPMI) with gapped co-occurrence and reversecomplement (RC) pairing, (2) positive PMI (PPMI) withinwindow, (3) directed de Bruijn path with RC links, (4) traditional directed de Bruijn graph without RC link, and (5) positional bin profile graphs. For the graph neural network models, we considered four GNN families, namely: GCN, GAT, GraphSAGE, and GIN. Nodes are$k$-mers, while edges encode statistical or positional relations between nodes, with normalization that controls for sequence length. Combined, these resulted in 20 different graph structure-GNN model combinations. We then integrated these 20 models to build a stacking classifier ensemble for subcellular lncRNA localization. Further, we evaluate inexact (Hamming-1)$k$-mer bucketing under this framework using graph-level SMOTE to handle data imbalance problems. On held-out tests, our best configurations improve overall accuracy compared to single-model baselines, while avoiding problem of fold-leakage by constructing graphs on per sequence basis. Weijun Yi, Gang-Qing Hu, Donald A. Adjeroh |
BIBM | 3 |
| 2025 | ItpCtrl-AI: End-to-end interpretable and controllable artificial intelligence by modeling radiologists' intentions
Trong-Thang Pham, Jacob Brecheisen, Carol C. Wu, Hien Van Nguyen, Zhigang Deng 0001, Donald A. Adjeroh, Gianfranco Doretto, Arabinda Choudhary, T. Hoang Ngan Le |
Artif. Intell. Medicine | 6 |
| 2024 | FG-CXR: A Radiologist-Aligned Gaze Dataset for Enhancing Interpretability in Chest X-Ray Report Generation
Trong-Thang Pham, Ngoc-Vuong Ho, Nhat-Tan Bui, Thinh Phan, Brijesh Patel 0001, Donald A. Adjeroh, Gianfranco Doretto, Anh Nguyen 0003, Carol C. Wu, T. Hoang Ngan Le |
ACCV (6) | 6 |
| 2024 | A Framework for Evaluating Model Trustworthiness in Classification of Very High Resolution Histopathology ImagesabstractIn computer vision, one approach to explaining a deep learning model’s decision is to show regions of visual evidence upon which the model makes a decision. Typically, this evidence is represented in the form of a saliency map which conveys how much an image region is contributing to the model’s decision. For a model to be trustworthy, it is expected that this saliency region should provide relevant information. In this work, we use model "trustworthiness" or "rationale" to describe how much relevant information the model is using to determine the image class. For medical images, this information connects to biological relevance. For very high resolution histopathology image applications, such as gigapixel whole-slide image classification, where patch-based multiple-instance based learning approach is taken to determine the patch label, this biological relevance has to be determined both at the patch and the image level. In this work, we present a novel patch-based model trustworthiness evaluation framework for very high resolution histopathology images. Our trustworthiness framework takes two approaches: spatial overlap based and feature based evaluation. For the overlap based approach, we check overlap with the annotation provided with the database to see if they have biological relevance, since for tumor positive patches only high probability regions from within the annotated regions are likely to be relevant. For feature based approach, we train an interpretability model using the sub-patches of the training set, extract features and cluster them. Then based on the distance from these clusters we determine if there is any biological rationale behind the prediction. Finally, we propose four patch-level and four image-level rationale metrics that evaluate the biological relevance of the information used by the classifier to decide on the patch class. Our experiment using the CAMELYON16 dataset shows the efficacy of this approach for model trustworthiness evaluation and explainability. Mohammad Iqbal Nouyed, Gianfranco Doretto, Donald A. Adjeroh |
BIBM | 3 |
| 2024 | A Machine Learning Approach to Motif Finding for LncRNA Sub-cellular Localization
Weijun Yi, Jason R. Miller, Gang-Qing Hu, Donald A. Adjeroh |
BIBM | 4 |
| 2024 | Multi-label Classification using Self-Supervised Learning: Addressing Class Inter-Dependency and Data ImbalanceabstractDeveloping multi-label classification models under significant class imbalance, and when annotating data requires expert-level knowledge remains a major challenge. Additionally, interdependency and correlation among labels are common in multi-label problems. In this work, we introduce a novel framework to address these challenges. Our approach extracts robust discriminative features from unlabeled data through self-supervised contrastive learning and uses an adaptive data augmentation mechanism (ACBA) to balance the dataset. Independent binary classifiers are trained for each class, using a new custom Focal Weighted Cross-Entropy (FWCE) loss function to focus on hard-to-classify examples. A correlation learning module then refines predictions by integrating statistical and domain-specific knowledge. Finally, a meta-learner, employing a Gated Recurrent Unit (GRU) and multi-head attention, identifies complex relationships between classes, even for those that rarely occur together. We used the detection of thoracic diseases using chest X-rays, a domain with a major class imbalance and highly associated labels, to validate our approach. Our findings demonstrate the potential of our method to apply to other medical and non-medical imaging scenarios with similar multi-label classification problems. Ghazaleh Mirzaee, Gianfranco Doretto, Donald A. Adjeroh |
ICMLA | 3 |
| 2024 | TabSeq: A Framework for Deep Learning on Tabular Data via Sequential Ordering
Al-Zadid Sultan Bin Habib, Kesheng Wang, Mary-Anne Hartley, Gianfranco Doretto, Donald A. Adjeroh |
ICPR (4) | 5 |
| 2024 | Efficient Classification of Histopathology Images Using Highly Imbalanced Data
Mohammad Iqbal Nouyed, Mary-Anne Hartley, Gianfranco Doretto, Donald A. Adjeroh |
ICPR (2) | 4 |
| 2024 | ZEETAD: Adapting Pretrained Vision-Language Model for Zero-Shot End-to-End Temporal Action DetectionabstractTemporal action detection (TAD) involves the localization and classification of action instances within untrimmed videos. While standard TAD follows fully supervised learning with closed-set setting on large training data, recent zero-shot TAD methods showcase the promising open-set setting by leveraging large-scale contrastive visual-language (ViL) pretrained models. However, existing zero-shot TAD methods have limitations on how to properly construct the strong relationship between two interdependent tasks of localization and classification and adapt ViL model to video understanding. In this work, we present ZEE-TAD, featuring two modules: dual-localization and zero-shot proposal classification. The former is a Transformer-based module that detects action events while selectively collecting crucial semantic embeddings for later recognition. The latter one, CLIP-based module, generates semantic embeddings from text and frame inputs for each temporal unit. Additionally, we enhance discriminative capability on unseen classes by minimally updating the frozen CLIP encoder with lightweight adapters. Extensive experiments on THUMOS14 and ActivityNet-1.3 datasets demonstrate our approach’s superior performance in zero-shot TAD and effective knowledge transfer from ViL models to unseen action categories. Code is available at https: //github.com/UARK-AICV/ZEETAD. Thinh Phan, Viet-Khoa Vo-Ho, Duy Le 0004, Gianfranco Doretto, Donald A. Adjeroh, T. Hoang Ngan Le |
WACV | 5 |
| 2024 | Machine learning on alignment features for parent-of-origin classification of simulated hybrid RNA-seqabstractBACKGROUND: Parent-of-origin allele-specific gene expression (ASE) can be detected in interspecies hybrids by virtue of RNA sequence variants between the parental haplotypes. ASE is detectable by differential expression analysis (DEA) applied to the counts of RNA-seq read pairs aligned to parental references, but aligners do not always choose the correct parental reference. RESULTS: We used public data for species that are known to hybridize. We measured our ability to assign RNA-seq read pairs to their proper transcriptome or genome references. We tested software packages that assign each read pair to a reference position and found that they often favored the incorrect species reference. To address this problem, we introduce a post process that extracts alignment features and trains a random forest classifier to choose the better alignment. On each simulated hybrid dataset tested, our machine-learning post-processor achieved higher accuracy than the aligner by itself at choosing the correct parent-of-origin per RNA-seq read pair. CONCLUSIONS: For the parent-of-origin classification of RNA-seq, machine learning can improve the accuracy of alignment-based methods. This approach could be useful for enhancing ASE detection in interspecies hybrids, though RNA-seq from real hybrids may present challenges not captured by our simulations. We believe this is the first application of machine learning to this problem domain. Jason R. Miller, Donald A. Adjeroh |
BMC Bioinform. | 2 |
| 2024 | Online continual decoding of streaming EEG signal with a balanced and informative memory buffer
Tiehang Duan, Zhenyi Wang 0001, Fang Li 0011, Gianfranco Doretto, Donald A. Adjeroh, Yiyi Yin, Cui Tao |
Neural Networks | 5 |
| 2023 | Replay with Stochastic Neural Transformation for Online Continual EEG ClassificationabstractBrain computer interface (BCI) systems used for clinical assistance purposes such as wheelchair control require decoding of streaming brain signals i.e. electroencephalography (EEG) signals over a long period of time with subject shift in the middle. Numerous challenges arise during this online continual brain signal decoding process: 1) the EEG decoder needs to deal with streaming EEG signals from sequentially arriving subjects, with no data available beforehand for large-scale pretraining; 2) the EEG decoder should avoid catastrophic forgetting on previous subjects after learning on a new subject; 3) the EEG decoder should perform well on noisy signals with high variance across subjects. We proposed a principled replay-based approach for this general decoding scenario, forming a bi-level optimization framework with stochastic neural transformation for dynamic memory evolution, making them representative in feature space and encouraging the model to generalize well. The evolved signal segments are stored and replayed during later decoding stages to achieve optimal model performance on all previous subjects. The stochastic neural transformation performed in inner sup of bi-level optimization significantly enhances the diversity of stored signal segments and improves model robustness during online continual decoding. We perform detailed theoretical analysis on model’s generalization ability in addition to the empirical evaluations. We construct multiple new benchmarks to mimic real-world online sequential EEG decoding scenarios with underlying subject shifts. The extensive evaluation of the proposed approach shows it outperforms related strong baselines by a large margin. Tiehang Duan, Zhenyi Wang 0001, Gianfranco Doretto, Fang Li 0011, Cui Tao, Donald A. Adjeroh |
BIBM | 6 |
| 2023 | Heteroskedasticity as a Signature of Association for Age-Related GenesabstractHuman aging is a process controlled by both genetics and environment. Many studies have been conducted to identify a subset of genes related to aging from the human genome. Biologists implicitly categorize age-related genes into genes that cause aging and genes that are influenced by aging, which resulted in both causal inference and inference of associations studies. While inference of association is better explored, causal inference and computational causal inference, remains less explored. In this work, we are primarily motivated to tackle the problem of identifying genes associated with aging, while having a brief look into genes with probable causal relations, both from a computational perspective. Specifically, we form a set of hypotheses and accordingly, introduce a data-tailored framework for inference. First we perform linear modeling on the expression values of age-related genes, and then examine the presence of heteroskedastic properties in the residual of the model. We evaluate this framework and our results suggest that, 1) presence of heteroskedasticity in these residuals is a potential signature of association for age-related genes, and 2) consistent heteroskedasticity along the human life span could imply some sort of causality. To our knowledge, along with identifying age-associated genes, this is the first work to propose a framework for computational causal inference on age-related genes, using a dataset of human dermal fibroblast gene expression data. Hence the results of our simple, yet effective approach can be used not only to assess future age-related genes, but also as a possible criterion to select new associative or potential causal genes with respect to aging. Salman Mohamadi, Donald A. Adjeroh |
BIBM | 2 |
| 2023 | LncRNA Subcellular Localization Signals - Are the Two Ends Equal? A Machine Learning Analysis Across Multiple Cell LinesabstractIn this work, we studied the question of whether the two ends of long non-coding ribonucleic acids (lncRNAs) (i.e., the 5′ end and 3′end) carry similar information about subcellular localization of lncRNAs. We considered this problem from three viewpoints using machine learning models: (1) consideration of the classification performance of the machine learning models using features from defined regions (or segments) along the sequence, (2) correlation-based analysis using models built on regions/segments along the LncRNA sequence, and (3) analysis of the relative positions of predicted lncRNA localization motifs along the LncRNA sequence. Our results and observations suggest that the 5′ region of the lncRNA sequences (the prefixes) tend to carry more localization signals when compared with the 3′ region (the suffixes) of the sequences. These could have implications on how we use machine learning models for improved analysis of lncRNA subcellular localization. Weijun Yi, Jason R. Miller, Gang-Qing Hu, Donald A. Adjeroh |
BIBM | 4 |
| 2023 | Distributionally Robust Cross Subject EEG DecodingabstractRecently, deep learning has shown to be effective for Electroencephalography (EEG) decoding tasks. Yet, its performance can be negatively influenced by two key factors: 1) the high variance and different types of corruption that are inherent in the signal, 2) the EEG datasets are usually relatively small given the acquisition cost, annotation cost and amount of effort needed. Data augmentation approaches for alleviation of this problem have been empirically studied, with augmentation operations on spatial domain, time domain or frequency domain handcrafted based on expertise of domain knowledge. In this work, we propose a principled approach to perform dynamic evolution on the data for improvement of decoding robustness. The approach is based on distributionally robust optimization and achieves robustness by optimizing on a family of evolved data distributions instead of the single training data distribution. We derived a general data evolution framework based on Wasserstein gradient flow (WGF) and provides two different forms of evolution within the framework. Intuitively, the evolution process helps the EEG decoder to learn more robust and diverse features. It is worth mentioning that the proposed approach can be readily integrated with other data augmentation approaches for further improvements. We performed extensive experiments on the proposed approach and tested its performance on different types of corrupted EEG signals. The model significantly outperforms competitive baselines on challenging decoding scenarios. Tiehang Duan, Zhenyi Wang 0001, Gianfranco Doretto, Fang Li 0011, Cui Tao, Donald A. Adjeroh |
ECAI | 6 |
| 2023 | More Synergy, Less Redundancy: Exploiting Joint Mutual Information for Self-Supervised LearningabstractSelf-supervised learning (SSL) is now a serious competitor for supervised learning, even though it does not require data annotation. Several baselines have attempted to make SSL models exploit information about data distribution, and less dependent on the augmentation effect. However, there is no clear consensus on whether maximizing or minimizing the mutual information between representations of augmentation views practically contribute to improvement or degradation in performance of SSL models. This paper is a fundamental work where, we investigate the role of mutual information in SSL, and reformulate the problem of SSL in the context of a new perspective on mutual information. To this end, we consider joint mutual information from the perspective of partial information decomposition (PID) as a key step in reliable multivariate information measurement. PID enables us to decompose joint mutual information into three important components, namely, unique information, redundant information and synergistic information. Our framework aims for minimizing the redundant information between views and the desired target representation while maximizing the synergistic information at the same time. Our experiments lead to a re-calibration of two redundancy reduction baselines, and a proposal for a new SSL training protocol. Experimental results on multiple datasets and two downstream tasks show the effectiveness of this framework. Salman Mohamadi, Gianfranco Doretto, Donald A. Adjeroh |
ICIP | 3 |
| 2023 | A Framework for Token-Based Scene-Classification of Remote Sensing ImagesabstractRemote sensing images often come in very high resolutions. For instance, large images of resolution 1000×1000 are quite common in remote sensing applications. Thus, one key challenge is how to effectively analysis such very high resolution images using low computational resources, while maintain a reasonable performance as appropriate to the specific problem domain, such as scene classification, object segmentation, or object detection. In this work, we propose a framework for addressing this challenge, with focus on the specific problem of scene-classification for remote-sensed aerial imagery. Various approaches have been proposed for scene-based classification of remote sensing imagery [1] , [2] , [3] . For a survey of approaches and challenges, see [4] . See also [5] . Mohammad Iqbal Nouyed, Gianfranco Doretto, Donald A. Adjeroh |
IGARSS | 3 |
| 2023 | FUSSL: Fuzzy Uncertain Self Supervised LearningabstractSelf supervised learning (SSL) has become a very successful technique to harness the power of unlabeled data, with no annotation effort. A number of developed approaches are evolving with the goal of outperforming supervised alternatives, which have been relatively successful. Similar to some other disciplines in deep representation learning, one main issue in SSL is robustness of the approaches under different settings. In this paper, for the first time, we recognise the fundamental limits of SSL coming from the use of a single-supervisory signal. To address this limitation, we leverage the power of uncertainty representation to devise a robust and general standard hierarchical learning/training protocol for any SSL baseline, regardless of their assumptions and approaches. Essentially, using the information bottleneck principle, we decompose feature learning into a two-stage training procedure, each with a distinct supervision signal. This double supervision approach is captured in two key steps: 1) invariance enforcement to data augmentation, and 2) fuzzy pseudo labeling (both hard and soft annotation). This simple, yet, effective protocol which enables cross-class/cluster feature learning, is instantiated via an initial training of an ensemble of models through invariance enforcement to data augmentation as first training phase, and then assigning fuzzy labels to the original samples for the second training phase. We consider multiple alternative scenarios with double supervision and evaluate the effectiveness of our approach on recent baselines, covering four different SSL paradigms, including geometrical, contrastive, non-contrastive, and hard/soft whitening (redundancy reduction) baselines. We performed extensive experiments under multiple settings to show that the proposed training protocol consistently improves the performance of the former baselines, independent of their respective underlying principles. Salman Mohamadi, Gianfranco Doretto, Donald A. Adjeroh |
WACV | 3 |
| 2023 | sCL-ST: Supervised Contrastive Learning With Semantic Transformations for Multiple Lead ECG Arrhythmia ClassificationabstractThe automatic classification of electrocardiogram (ECG) signals has played an important role in cardiovascular diseases diagnosis and prediction. With recent advancements in deep neural networks (DNNs), particularly Convolutional Neural Networks (CNNs), learning deep features automatically from the original data is becoming an effective and widespread approach in a variety of intelligent tasks including biomedical and health informatics. However, most of the existing approaches are trained on either 1D CNNs or 2D CNNs, and they suffer from the limitations of random phenomena (i.e. random initial weights). Furthermore, the ability to train such DNNs in a supervised manner in healthcare is often limited due to the scarcity of labeled training data. To address the problems of weight initialization and limited annotated data, in this work, we leverage recent self-supervised learning technique, namely, contrastive learning, and present supervised contrastive learning (sCL). Different from existing self-supervised contrastive learning approaches, which often generate false negatives because of random selection of negative anchors, our contrastive learning makes use of labeled data to pull the same class closer together and push different classes far apart to avoid potential false negatives. Furthermore, unlike other kinds of signals (e.g. speech, image, video), ECG signal is sensitive to changes, and inappropriate transformation could directly affect diagnosis results. To deal with this issue, we present two semantic transformations, i.e. semantic split-join and semantic weighted peaks noise smoothing. The proposed deep neural network sCL-ST with supervised contrastive learning and semantic transformations is trained as an end-to-end framework for the multi-label classification of 12-lead ECGs. Our sCL-ST network contains two sub-networks i.e. pre-text task and down-stream task. Our experimental results have been evaluated on 12-lead PhysioNet 2020 dataset and shown that our proposed network outperforms the state-of-the-art existing approaches. Sang Truong, Brijesh Patel 0001, Donald A. Adjeroh, T. Hoang Ngan Le |
IEEE J. Biomed. Health Informatics | 4 |
| 2022 | Deep Active Ensemble Sampling for Image Classification
Salman Mohamadi, Gianfranco Doretto, Donald A. Adjeroh |
ACCV (7) | 3 |
| 2022 | Efficient Classification of Very High Resolution Histopathological ImagesabstractOver the years, deep learning approaches have shown significant improvement in various image understanding tasks. However, analysis of high resolution images still remains a major challenge. Apart from the huge computational resources required for such images, the large image sizes make it difficult to extract effective contextual information needed for important tasks, such as classification, segmentation, or clustering of such images. In this work, we address the challenge of high resolution image classification u sing a new discriminative patch selection approach. We embed our patch selection approach inside a novel classification framework, supporting potential use of different pre-trained learning models. We show results on a high resolution image dataset, namely, gigapixel whole slide tissue images for cancer tumors. We demonstrate the performance of the proposed approaches using comparative analysis with state-of-the art methods on this dataset. Mohammad Iqbal Nouyed, Gianfranco Doretto, Donald A. Adjeroh |
BIBM | 3 |
| 2022 | Identification of all-against-all protein-protein interactions based on deep hash learningabstractBACKGROUND: Protein-protein interaction (PPI) is vital for life processes, disease treatment, and drug discovery. The computational prediction of PPI is relatively inexpensive and efficient when compared to traditional wet-lab experiments. Given a new protein, one may wish to find whether the protein has any PPI relationship with other existing proteins. Current computational PPI prediction methods usually compare the new protein to existing proteins one by one in a pairwise manner. This is time consuming. RESULTS: In this work, we propose a more efficient model, called deep hash learning protein-and-protein interaction (DHL-PPI), to predict all-against-all PPI relationships in a database of proteins. First, DHL-PPI encodes a protein sequence into a binary hash code based on deep features extracted from the protein sequences using deep learning techniques. This encoding scheme enables us to turn the PPI discrimination problem into a much simpler searching problem. The binary hash code for a protein sequence can be regarded as a number. Thus, in the pre-screening stage of DHL-PPI, the string matching problem of comparing a protein sequence against a database with M proteins can be transformed into a much more simpler problem: to find a number inside a sorted array of length M. This pre-screening process narrows down the search to a much smaller set of candidate proteins for further confirmation. As a final step, DHL-PPI uses the Hamming distance to verify the final PPI relationship. CONCLUSIONS: The experimental results confirmed that DHL-PPI is feasible and effective. Using a dataset with strictly negative PPI examples of four species, DHL-PPI is shown to be superior or competitive when compared to the other state-of-the-art methods in terms of precision, recall or F1 score. Furthermore, in the prediction stage, the proposed DHL-PPI reduced the time complexity from [Formula: see text] to [Formula: see text] for performing an all-against-all PPI prediction for a database with M proteins. With the proposed approach, a protein database can be preprocessed and stored for later search using the proposed encoding scheme. This can provide a more efficient way to cope with the rapidly increasing volume of protein datasets. Yue Jiang 0001, Donald A. Adjeroh, Zhidong Liu |
BMC Bioinform. | 4 |
| 2022 | Deep Learning for Adverse Event Detection From Web SearchabstractAdverse event detection is critical for many real-world applications including timely identification of product defects, disasters, and major socio-political incidents. In the health context, adverse drug events account for countless hospitalizations and deaths annually. Since users often begin their information seeking and reporting with online searches, examination of search query logs has emerged as an important detection channel. However, search context - including query intent and heterogeneity in user behaviors – is extremely important for extracting information from search queries, and yet the challenge of measuring and analyzing these aspects has precluded their use in prior studies. We propose DeepSAVE, a novel deep learning framework for detecting adverse events based on user search query logs. DeepSAVE uses an enriched variational autoencoder encompassing a novel query embedding and user modeling module that work in concert to address the context challenge associated with search-based detection of adverse events. Evaluation results on three large real-world event datasets show that DeepSAVE outperforms existing detection methods as well as comparison deep learning auto encoders. Ablation analysis reveals that each component of DeepSAVE significantly contributes to its overall performance. Collectively, the results demonstrate the viability of the proposed architecture for detecting adverse events from search query logs. Ahmed Abbasi, Brent Kitchens, Donald A. Adjeroh, Daniel Dajun Zeng |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2021 | A Deep Learning Model for 16S rRNA Classification with Taxonomic Tree EmbeddingabstractBacterial sequence classification is an important, but difficult, challenge in microbiome research. Existing deep learning methods of bacterial sequence classification usually use a discriminator to perform the required classification. This often requires many model parameters, resulting in poor generalizability, and thus limit applicability of the model. In this paper, we propose a novel method, Bacterial Sequence Classification with Taxonomic Tree Embedding (BSC-TTE) to classify bacterial sequences. First, it encodes every taxon in the taxonomic tree into vectors. Second, it applies a deep learning framework to encode the bacterial sequences into corresponding vectors while involving the taxa vectors as part of the training. Finally, it discriminates a bacterial sequence by calculating the cosine distance between its sequence vector and a genera vector. Compared to the general deep learning discriminator, this classification method is not only simpler, but also faster with better adaptability. Experimental results demonstrate that the proposed BSC-TTE approach provides an improved classification accuracy over other state-of-the-art methods. Yue Jiang 0001, Donald A. Adjeroh |
BIBM | 3 |
| 2021 | Detecting Drug-Drug Interactions using Protein Sequence-Structure Similarity NetworksabstractAdverse drug events represent a key challenge in public health, especially with respect to drug safety profiling and drug surveillance. Drug-drug interactions represent one of the most popular types of adverse drug events. Most computational approaches to this problem have used different types of data, such as drug chemical structure, information about protein targets, side effects, pathways, etc to predict potential interactions between drugs. In this work, we study the question of whether using just genetic information about the drugs can provide significant information about the potential safety profile for a given drug. We propose a novel neural network model to predict adverse drug events using only data about the protein sequence and protein structure associated with the drug targets. We compare the results with those from the state-of-the-art methods on this problem. Our results show that the proposed method is quite competitive, at times outperforming the state-of-the-art. Saminur Islam, Ahmed Abbasi, Nitin Agarwal 0001, Wanhong Zheng, Gianfranco Doretto, Donald A. Adjeroh |
BIBM | 6 |
| 2021 | An Information-Theoretic Framework for Identifying Age-Related Genes Using Human Dermal Fibroblast Transcriptome DataabstractInvestigation of age-related genes is of great importance for multiple purposes, for instance, improving our understanding of the mechanism of ageing, increasing life expectancy, age prediction, and other healthcare applications. In this work, starting with a set of 27,142 genes, we develop an information-theoretic framework for identifying genes that are associated with aging by applying unsupervised and semi-supervised learning techniques on human dermal fibroblast gene expression data. First, we use unsupervised learning and apply information-theoretic measures to identify key features for effective representation of gene expression values in the transcriptome data. Using the identified features, we perform clustering on the data. Finally, we apply semi-supervised learning on the clusters using different distance measures to identify novel genes that are potentially associated with aging. Performance assessment for both unsupervised and semi-supervised methods show the effectiveness of the framework. Salman Mohamadi, Donald A. Adjeroh |
BIBM | 2 |
| 2021 | Human Age Estimation from Gene Expression Data using Artificial Neural NetworksabstractThe study of signatures of aging in terms of genomic biomarkers can be uniquely helpful in understanding the mechanisms of aging and developing models to accurately predict the age. Prior studies have employed gene expression and DNA methylation data aiming at accurate prediction of age. In this line, we propose a new framework for human age estimation using information from human dermal fibroblast gene expression data. First, we propose a new spatial representation as well as a data augmentation approach for gene expression data. Next in order to predict the age, we design an architecture of neural network and apply it to this new representation of the original and augmented data, as an ensemble classification approach. Our experimental results suggest the superiority of the proposed framework over state-of-the-art age estimation methods using DNA methylation and gene expression data. Salman Mohamadi, Nasser M. Nasrabadi, Gianfranco Doretto, Donald A. Adjeroh |
BIBM | 4 |
| 2021 | A Deep Learning Approach to LncRNA Subcellular Localization Using Inexact q-mersabstractLong non-coding Ribonucleic Acids (lncRNAs) can be localized to different cellular components, such as the nucleus, exosome, cytoplasm, ribosome, etc. Their biological functions can be influenced by the region of the cell where they are located. Many of these lncRNAs are associated with different challenging diseases. Thus, it is crucial to study their subcellular localization. However, compared to the massive number of lncRNAs, only relatively few have annotations in terms of their subcellular localization. Conventional computational methods use q-mer profiles from lncRNA sequences and train machine learning models, such as support vector machines and logistic regression with the profiles. These methods focus on the exact q-mer. Given possible sequence mutations and other uncertainties in genomic sequences and their role in biological function, a consideration of these changes might improve our ability to model lncRNAs and their localization. We hypothesize that considering these changes may improve our ability to predict subcellular localization of lncRNAs. To test this hypothesis, we propose a deep learning model with inexact q-mers for the localization of lncRNAs in the cell. The proposed method can obtain a high overall accuracy of 94.7%, an average of 91.3% on a benchmark dataset, using 8-mers with mismatches. In comparison, the exact 8-mer result was 89.8%. The proposed approach outperformed existing state-of-art lncRNA localization predictors on two different datasets. Our results, therefore, support the hypothesis that deep learning models using inexact q-mers can improve the performance of computational lncRNA localization algorithms. Weijun Yi, Donald A. Adjeroh |
BIBM | 2 |
| 2021 | Deep learning for biological age estimationabstractModern machine learning techniques (such as deep learning) offer immense opportunities in the field of human biological aging research. Aging is a complex process, experienced by all living organisms. While traditional machine learning and data mining approaches are still popular in aging research, they typically need feature engineering or feature extraction for robust performance. Explicit feature engineering represents a major challenge, as it requires significant domain knowledge. The latest advances in deep learning provide a paradigm shift in eliciting meaningful knowledge from complex data without performing explicit feature engineering. In this article, we review the recent literature on applying deep learning in biological age estimation. We consider the current data modalities that have been used to study aging and the deep learning architectures that have been applied. We identify four broad classes of measures to quantify the performance of algorithms for biological age estimation and based on these evaluate the current approaches. The paper concludes with a brief discussion on possible future directions in biological aging research using deep learning. This study has significant potentials for improving our understanding of the health status of individuals, for instance, based on their physical activities, blood samples and body shapes. Thus, the results of the study could have implications in different health care settings, from palliative care to public health. Syed Ashiqur Rahman, Peter Giacobbi, Lee Pyles, Charles J. Mullett, Gianfranco Doretto, Donald A. Adjeroh |
Briefings Bioinform. | 6 |
| 2020 | J2*: A New Method for Alignment-free Sequence Similarity MeasurementabstractAlignment-free sequence comparison methods can compute the similarity of a huge number of sequences much faster than traditional sequence alignment methods. Here, a new non-parametric alignment-free sequence comparison algorithm, J*2is proposed to measure sequence similarity based on the suffix tree data structure. Compared against other state-of-the-art alignment-free methods, namely, Dz2, D2, D2sh, D*2, WFV, DV , Shi, CPF, DMk, K2 and K*2, J*2has three main advantages: (1) it has the fastest running time in theory and in practice. J*2reduces the time for k-words search from O(N2) to O(N). Our experimental results confirm that it is the fastest among the 11 popular approaches. (2) J*2is easy to use: unlike the other alignment-free methods that often need to choose a suitable parameter k, there is no parameter selection for J*2. (3) J*2does not have any particular requirement for data distribution. Unlike the the parametric methods (such as the D2-family) that require certain distribution for the data, J*2has no demand for a specific input distribution. The improved running time from J*2will be very useful in this era of big data, especially, with the increasing data volume of genome sequences. Yue Jiang 0001, Donald A. Adjeroh, Bing-Hua Jiang |
BIBM | 2 |
| 2020 | Exploring Neural Network Models for LncRNA Sequence IdentificationabstractDistinguishing long non-coding RNA from protein-coding RNA is important to molecular and cellular biology. The problem can be addressed with machine learning in general and with artificial neural networks in particular. We explore the effects of various network design choices on the accuracy of human LncRNA identification. Perceptron-based neural network models were found to be almost as accurate as more complex recurrent neural networks, and K-mer representations of the data seemed to assist both. Size selection of training data affected results. These explorations could assist in neural network design for RNA analysis. Jason R. Miller, Donald A. Adjeroh |
BIBM | 2 |
| 2020 | A New Framework for Spatial Modeling and Synthesis of Genomic SequencesabstractThis paper provides a framework for statistical modeling of genomic sequences. Such a framework can be used a the basis for the synthesize similar sequences. The synthesized sequences could then be used to make for further inference about the genomic sequences. We start by converting the sequence of nucleotides from the genome into a decimal sequence via Huffman coding. Using the HodrickPrescott filter (HP filter) this decimal sequence is decomposed into two components, namely, trend and cyclic. Next, the ARIMA-GARCH statistical modeling approach is applied on the trend component exhibiting heteroskedasticity. The autoregressive integrated moving average (ARIMA) is used to capture the linear characteristics of the sequence, while the generalized autoregressive conditional heteroskedasticity (GARCH) is applied to model the statistical nonlinearity of the genome sequence. This modeling approach allows us to synthesize a given genomic sequence based on its statistical charatceristics. Finally, the probability distribution function (PDF) of a given sequence is estimated using a Gaussian mixture model, and based on the estimated PDF, we determine a new PDF representing sequences that statistically counteract the original sequence. We applied the proposed framework on several genes, as well as on the HIV nucleotide sequence. The corresponding results show some promise. Salman Mohamadi, Donald A. Adjeroh, Behnoush Behi, Hamidreza Amindavar |
BIBM | 2 |
| 2020 | Adversarial Latent AutoencodersabstractAutoencoder networks are unsupervised approaches aiming at combining generative and representational properties by learning simultaneously an encoder-generator map. Although studied extensively, the issues of whether they have the same generative power of GANs, or learn disentangled representations, have not been fully addressed. We introduce an autoencoder that tackles these issues jointly, which we call Adversarial Latent Autoencoder (ALAE). It is a general architecture that can leverage recent improvements on GAN training procedures. We designed two autoencoders: one based on a MLP encoder, and another based on a StyleGAN generator, which we call StyleALAE. We verify the disentanglement properties of both architectures. We show that StyleALAE can not only generate 1024x1024 face images with comparable quality of StyleGAN, but at the same resolution can also produce face reconstructions and manipulations based on real images. This makes ALAE the first autoencoder able to compare with, and go beyond the capabilities of a generator-only type of architecture. Stanislav Pidhorskyi, Donald A. Adjeroh, Gianfranco Doretto |
CVPR | 2 |
| 2020 | Mapping the Land Development Processes Using Data Transformation and Clustering MethodsabstractUnderstanding the historical trend of land development provides an invaluable source of information for policy-makers, environmental planners, and regional scientists. This information is useful in the study of the impacts of past decisions and natural conditions on the patterns of the trend. In this study, we built a hybrid algorithm that benefits from various techniques of land change detection to map the land development process in a given region. Data transformation and band differencing were used to construct a new feature class from the temporal satellite data. For data of three consecutive decades, a pairwise comparison of images was conducted. A clustering algorithm was applied to the constructed feature class to identify the location of new land developments. The results indicate a user's and producer's accuracy of 90.3%. The findings of this study are useful in the study of historic trends in land development and its patterns. Pariya Pourmohammadi, Donald A. Adjeroh, Michael P. Strager |
IGARSS | 2 |
| 2020 | Centroid of Age Neighborhoods: A New Approach to Estimate Biological AgeabstractEstimation of human biological age is an important and difficult challenge. Different biomarkers and numerous approaches have been studied for biological age prediction, each with its advantages and limitations. In this paper, we propose a new biological age estimation method, and investigate the performance of the new method. We introduce a centroid based approach, using the notion of age neighborhoods. Specifically, we develop a model, based on which we compute biological age using blood biomarkers, by considering the centroid or mediod of specially selected age neighborhoods. Experiments were performed on the National Health and Human Nutrition Examination Survey dataset with biomarkers (21 451 individuals). Compared with current popular methods for biological age prediction, our experiments show that the proposed age neighborhood model results in an improved performance in human biological age estimation. Syed Ashiqur Rahman, Donald A. Adjeroh |
IEEE J. Biomed. Health Informatics | 2 |
| 2019 | FaceSNPs: Identifying Face-Related SNPs from the Human GenomeabstractThe human face is an important part of the human body. Finding SNPs associated with its different regions can open doors to deeper understanding of different genetic diseases affecting the face, improve genotype-to-phenotype analysis for forensics, etc. In this work we propose a process for constructing a panel of SNPs associated with the human face leveraging a database of published work on different face regions. We validate our selection by using them to predict human biogeographical ancestry at both continent and sub-continent levels, as well as by manual confirmation of selected genes. Although our process makes no assumptions on ancestry informativeness of selected SNPs, our SNP panel perform comparatively well at the task, with some subsets outperforming those from published work. We also highlight genes/chromosomes that promise better discriminative power to help those under resource constraint focus on only a few genes/chromosomes for further analyses. Somadina Mbadiwe, Jeremy M. Dawson, Donald A. Adjeroh |
BIBM | 3 |
| 2019 | Proximity Search Method for Mining Biomedical and Genomic InformationabstractThe human face is an important part of the human body. Finding SNPs associated with its different regions can open doors to deeper understanding of different genetic diseases affecting the face, improve genotype-to-phenotype analysis for forensics, etc. In this work we propose a process for constructing a panel of SNPs associated with the human face leveraging a database of published work on different regions of the face and applying proximity search to ensure relevant selections of publications. We validate the process by selecting top SNPs ranked using standard metrics to make predictions on human biogeographical ancestry at the continent level. Although our process makes no assumptions on ancestry informativeness of SNPs, our panel, grouped by chromosomes, performs comparatively well at the task of ancestry classification. Our results highlight genes/chromosomes that promise better discriminative power to help those under resource constraint focus on only a few genes/chromosomes for further analyses. Somadina Mbadiwe, Jeremy M. Dawson, Donald A. Adjeroh |
BIBM | 3 |
| 2019 | Estimating Biological Age from Physical Activity using Deep Learning with 3D CNNabstractWe introduce an approach to predict biological age based on a 3-dimensional deep convolutional neural network (3D-CNN), using human physical activity as recorded by a wearable device. Results on mortality hazard analysis using both the Cox proportional hazard model and Kaplan-Meier curves each show that the proposed method results in an improved performance. This work has significant implications in combining wearable sensors and deep learning techniques for improved health monitoring, for instance, in a mobile health environment. Syed Ashiqur Rahman, Donald A. Adjeroh |
BIBM | 2 |
| 2019 | Predicting Impervious Land Expansion Using Deep Deconvolutional Neural NetworksabstractIn this research, we propose a method for modeling land change using the idea of pixel-wise semantic segmentation through deep deconvolutional neural networks. This analysis is done on a watershed scale with a focus on where developed lands are predicted to expand. The novelty of our approach is in the application of deep learning in the prediction of land transformation, where we integrate high dimensional feature classes encompassing multiple variables. We introduce a method to construct cubes of land patches which include information related to the characteristics of terrain, proximity to other features, population, geo-political-boundaries, and public policy. After modeling the development expansion in a watershed using an encoder-decoder network, the accuracy of the model is computed through the Area Under the Curve of Receiver Operating Characteristics (AUC-ROC). Model performance indicates an accuracy of 80%. Future modeling should consider the use of this technique to better understand and map the spatially explicit landscape changes and aid in land use decision making. Pariya Pourmohammadi, Donald A. Adjeroh, Michael P. Strager |
IGARSS | 2 |
| 2018 | Deep Supervised Hashing with Spherical Embedding
Stanislav Pidhorskyi, Quinn Jones, Saeid Motiian, Donald A. Adjeroh, Gianfranco Doretto |
ACCV (4) | 4 |
| 2018 | Random Subspace Projection for Predicting Biogeographical Ancestry
Tanjin Taher Toma, Tayo Obafemi-Ajayi, Jeremy M. Dawson, Donald A. Adjeroh |
BIBM | 4 |
| 2018 | K2 and K 2 * : efficient alignment-free sequence similarity measurement based on Kendall statisticsabstractMotivation: Alignment-free sequence comparison methods can compute the pairwise similarity between a huge number of sequences much faster than sequence-alignment based methods. Results: We propose a new non-parametric alignment-free sequence comparison method, called K2, based on the Kendall statistics. Comparing to the other state-of-the-art alignment-free comparison methods, K2 demonstrates competitive performance in generating the phylogenetic tree, in evaluating functionally related regulatory sequences, and in computing the edit distance (similarity/dissimilarity) between sequences. Furthermore, the K2 approach is much faster than the other methods. An improved method, K2*, is also proposed, which is able to determine the appropriate algorithmic parameter (length) automatically, without first considering different values. Comparative analysis with the state-of-the-art alignment-free sequence similarity methods demonstrates the superiority of the proposed approaches, especially with increasing sequence length, or increasing dataset sizes. Availability and implementation: The K2 and K2* approaches are implemented in the R language as a package and is freely available for open access (http://community.wvu.edu/daadjeroh/projects/K2/K2_1.0.tar.gz). Contact: [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online. Donald A. Adjeroh, Bing-Hua Jiang, Yue Jiang 0001 |
Bioinform. | 2 |
| 2018 | SSAW: A new sequence similarity analysis method based on the stationary discrete wavelet transformabstractBACKGROUND: Alignment-free sequence similarity analysis methods often lead to significant savings in computational time over alignment-based counterparts. RESULTS: A new alignment-free sequence similarity analysis method, called SSAW is proposed. SSAW stands for Sequence Similarity Analysis using the Stationary Discrete Wavelet Transform (SDWT). It extracts k-mers from a sequence, then maps each k-mer to a complex number field. Then, the series of complex numbers formed are transformed into feature vectors using the stationary discrete wavelet transform. After these steps, the original sequence is turned into a feature vector with numeric values, which can then be used for clustering and/or classification. CONCLUSIONS: Using two different types of applications, namely, clustering and classification, we compared SSAW against the the-state-of-the-art alignment free sequence analysis methods. SSAW demonstrates competitive or superior performance in terms of standard indicators, such as accuracy, F-score, precision, and recall. The running time was significantly better in most cases. These make SSAW a suitable method for sequence analysis, especially, given the rapidly increasing volumes of sequence data required by most modern applications. Donald A. Adjeroh, Bing-Hua Jiang, Yue Jiang 0001 |
BMC Bioinform. | 3 |
| 2017 | Structure-based protein family signature: Efficient comparison of multidomain proteinsabstractThe rapid increase in available protein structure datasets requires new techniques for fast, yet, effective analysis of protein 3D structures. In this work, we propose a structure-based signature for protein families, suitable for rapid analysis of multidomain domain protein structures. Our method is alignment-free, using protein strings as the basic representation. A key novelty is the two-stage approach, whereby an initial list of candidate protein superfamilies are rapidly identified using the protein family signature, and then information retrieval methods are applied only to the members of the candidate superfamilies. This approach is the key to both improved speed, and improved structure retrieval accuracy. Experimental results, including comparative results with state-of-the-art methods, demonstrate the performance of the proposed protein family signature on queries with multidomain protein structures. Donald A. Adjeroh |
BIBM | 2 |
| 2017 | What can one chromosome tell us about human biogeographical ancestry?abstractWe study the problem of predicting human biogeographical ancestry using genomic data. While continental level ancestry is relatively simple using genomic information, distinguishing between individuals from closely associated subpopulations (e.g., from the same continent) is still a difficult challenge. In particular, we focus on the case where the analysis is constrained to using single nucleotide polymorphisms (SNPs) from just one chromosome. We thus propose methods to construct such ancestry informative SNP panels, and assess the performance of such SNP panels from just one chromosome, for both continental-level and sub-population level ancestry prediction. We present results on the performance of the proposed methods, including a comparison with other related methods. Tanjin Taher Toma, Zachary Williams, Jeremy M. Dawson, Donald A. Adjeroh |
BIBM | 4 |
| 2017 | Unified Deep Supervised Domain Adaptation and GeneralizationabstractThis work provides a unified framework for addressing the problem of visual supervised domain adaptation and generalization with deep models. The main idea is to exploit the Siamese architecture to learn an embedding subspace that is discriminative, and where mapped visual domains are semantically aligned and yet maximally separated. The supervised setting becomes attractive especially when only few target data samples need to be labeled. In this scenario, alignment and separation of semantic probability distributions is difficult because of the lack of data. We found that by reverting to point-wise surrogates of distribution distances and similarities provides an effective solution. In addition, the approach has a high “speed” of adaptation, which requires an extremely low number of labeled target training samples, even one per category can be effective. The approach is extended to domain generalization. For both applications the experiments show very promising results. Saeid Motiian, Marco Piccirilli, Donald A. Adjeroh, Gianfranco Doretto |
ICCV | 3 |
| 2017 | Efficient shape classification using region descriptors
Cong Lin 0001, Chi-Man Pun, Chi-Man Vong, Donald A. Adjeroh |
Multim. Tools Appl. | 4 |
| 2016 | Compressing genome resequencing data via the Maximal Longest FactorabstractWe propose a new algorithm for reference-based compression of genome resequencing data. First, we recap a recent reference-based technique for compressing resequencing data via the Longest Previous Factor (LPF). Viewing the problem from a new light, we call this the Sequential Longest Factor (SLF) method, and introduce improvements to the SLF approach. We further leverage the LPF and propose a new compression method: the Maximal Longest Factor (MLF). For the Homo sapiens genome, our proposed MLF achieves a compression ratio of 486, a significant improvement over 399 (newly improved SLF), 360 (original SLF, Beal, et al., BMC Genomics, 2016), 171 (Pinho, et al., NAR, 2011), and 157 (Wang and Zhang, NAR, 2011). Richard Beal, Aliya Farheen, Donald A. Adjeroh |
BIBM | 3 |
| 2016 | K2: Efficient alignment-free sequence similarity measurement using the Kendall statisticabstractAlignment-free sequence comparison methods can compute the similarity between a large number of sequences much faster than methods that depend on sequence alignment. We propose a new alignment-free sequence comparison method, called K2, based on the non-parametric Kendall statistic. Compared with the state-of-the-art alignment-free comparison methods (e.g., D2, D2*, D2sh, and Chisquare(χ2) statistic), K2showed comparative power, demonstrating similar or better performance in computing the edit distance (similarity/dissimilarity) among a huge number of sequences. The K2approach was much faster than each of the other methods, especialy, with long sequence lengths. Donald A. Adjeroh, Bing-Hua Jiang, Yue Jiang 0001 |
BIBM | 2 |
| 2016 | Information Bottleneck Learning Using Privileged Information for Visual RecognitionabstractWe explore the visual recognition problem from a main data view when an auxiliary data view is available during training. This is important because it allows improving the training of visual classifiers when paired additional data is cheaply available, and it improves the recognition from multi-view data when there is a missing view at testing time. The problem is challenging because of the intrinsic asymmetry caused by the missing auxiliary view during testing. We account for such view during training by extending the information bottleneck method, and by combining it with risk minimization. In this way, we establish an information theoretic principle for leaning any type of visual classifier under this particular setting. We use this principle to design a large-margin classifier with an efficient optimization in the primal space. We extensively compare our method with the state-of-the-art on different visual recognition datasets, and with different types of auxiliary data, and show that the proposed framework has a very promising potential. Saeid Motiian, Marco Piccirilli, Donald A. Adjeroh, Gianfranco Doretto |
CVPR | 3 |
| 2016 | Compressed parameterized pattern matching
Richard Beal, Donald A. Adjeroh |
Theor. Comput. Sci. | 2 |
| 2015 | A new algorithm for "the LCS problem" with application in compressing genome resequencing dataabstractThe longest common subsequence (LCS) problem is a classical problem in computer science, and forms the basis of the current best-performing reference-based compression schemes for genome resequencing data. First, we present a new algorithm for the LCS problem. Then, we introduce an LCS-motivated reference-based compression scheme using the components of the LCS, rather than the LCS itself. For the Homo sapiens genome (original size 3,080,436,051 bytes), our proposed scheme compressed the genome to 5,267,656 bytes. This can be compared with the previous best results of 19,666,791 bytes (Wang and Zhang, 2011) and 17,971,030 bytes (Pinho, Pratas, and Garcia, 2011). Thus, our compression ratio is about 3.73 to 3.41 times better than those from the state-of-the-art reference-based compression algorithms. Richard Beal, Tazin Afrin, Aliya Farheen, Donald A. Adjeroh |
BIBM | 4 |
| 2015 | A new encoding scheme for protein structure representationabstractGiven the rapidly increasing quantity of genomic and proteomic data that is now easily available even to a casual observer, the new challenge is in making sense out of the vast quantities of data. Efficient and reliable analysis of protein 3D structures is identified as a major challenge in this post genomic era. Whether the objective of the analysis is for protein classification, protein similarity search, protein structure prediction, discovery of protein structural motifs, or assignment of a functional class to a newly discovered protein, a key aspect in the analysis is the representation used to encode the protein 3D structural information. In this work, we introduce a family of string encodings as an effective descriptor for protein 3D structures. We show how the choice of parameters affects the performance and compare the result with other related research. Donald A. Adjeroh |
BIBM | 2 |
| 2015 | Circular Pattern DiscoveryabstractGiven a text or database T, the circular pattern discovery (CPD) problem is to identify ‘interesting’ circular patterns in T. Here, no specific input pattern is provided, and what is interesting is typically defined in terms of constraints in the search. We propose two algorithms for the CPD problem. The first algorithm uses suffix trees and suffix links to solve the exact CPD problem in time, where m2 is the maximum length of the circular patterns and N is the total length of the sequence database. The second algorithm uses suffix arrays to solve the more challenging approximate CPD (ACPD) problem in worst case, and on average, where k is the maximum allowed error(s). By exploiting the nature of the ACPD problem, the complexity is reduced to time in the worst case, and on average. Yue Jiang 0001, Donald A. Adjeroh |
Comput. J. | 3 |
| 2015 | Efficient pattern matching for RNA secondary structures
Richard Beal, Donald A. Adjeroh |
Theor. Comput. Sci. | 2 |
| 2013 | Compressed Parameterized Pattern MatchingabstractTraditional pattern matching between strings, from the alphabet Σ, is well defined for both uncompressed and compressed sequences. Prior to this work, parameterized pattern matching (p-matching) was defined predominately by the matching between uncompressed parameterized strings (p-strings) from the constant alphabet Σ and the parameter alphabet II. In this work, we define the compressed parameterized pattern matching (compressed p-matching) problem to find all of the p-matches between a pattern P and text T, using only P and the compressed text Tc. Initially, we present parameterized compression (p-compression) as a new way to losslessly compress data to support p-matching. Experimentally, we show that p-compression is competitive with various other standard compression schemes. Subsequently, we provide the compression and decompression algorithms. Using p-compression, we address the compressed p-matching problem. Our general solution is independent of the underlying compression scheme. The results are further examined for the specific case of Tunstall codes. Richard Beal, Donald A. Adjeroh |
DCC | 2 |
| 2012 | Border Array for Structural Strings
Richard Beal, Donald A. Adjeroh |
IWOCA | 2 |
| 2012 | Probabilistic suffix array: efficient modeling and prediction of protein familiesabstractMOTIVATION: Markov models are very popular for analyzing complex sequences such as protein sequences, whose sources are unknown, or whose underlying statistical characteristics are not well understood. A major problem is the computational complexity involved with using Markov models, especially the exponential growth of their size with the order of the model. The probabilistic suffix tree (PST) and its improved variant sparse probabilistic suffix tree (SPST) have been proposed to address some of the key problems with Markov models. The use of the suffix tree, however, implies that the space requirement for the PST/SPST could still be high. RESULTS: We present the probabilistic suffix array (PSA), a data structure for representing information in variable length Markov chains. The PSA essentially encodes information in a Markov model by providing a time and space-efficient alternative to the PST/SPST. Given a sequence of length N, construction and learning in the PSA is done in O(N) time and space, independent of the Markov order. Prediction using the PSA is performed in O(mlog{N is divided by /Σ/}) time, where m is the pattern length, and Σ is the symbol alphabet. In terms of modeling and prediction accuracy, using protein families from Pfam 25.0, SPST and PSA produced similar results (SPST 89.82%, PSA 89.56%), but slightly lower than HMMER3 (92.55%). A modified algorithm for PSA prediction improved the performance to 91.7%, or just 0.79% from HMMER3 results. The average (maximum) practical construction space for the protein families tested was 21.58±6.32N (41.11N) bytes using the PSA, 27.55±13.16N (63.01N) bytes using SPST and 47±24.95N (140.3N) bytes for HMMER3. The PSA was 255 times faster to construct than the SPST, and 11 times faster than HMMER3. Donald A. Adjeroh, Bing-Hua Jiang |
Bioinform. | 2 |
| 2012 | All-Against-All Circular Pattern MatchingabstractGiven a text T=T[1 … n] and a circular pattern P=P[1 … m], the circular pattern matching (CPM) problem is to find all occurrences of P in T. We present the first algorithm that exploits suffix links to solve the exact CPM (ECPM) problem in O(n log |Σ|) time and O(n) space, where Σ is the symbol alphabet. Then, we present a q-gram-based algorithm for the approximate CPM (ACPM) problem using the idea of bidirectional edit distance. Our algorithm finds all k-approximate occurrences of P in T. We then extend each algorithm to solve the all-against-all variant of the CPM problem for both exact and k-approximate matches. Although the CPM problem has been studied since the 1980s, this is the first attempt on the all-against-all variant, without using a trivial application of standard CPM algorithms. Given a database S=S1$1S2$2 … SZ$Z of Z sequences, our algorithms solve the all-against-all ACPM problem in O(kmaN) time on average and O(kmmN2) worst case, where k is the error parameter, N=∑Zi=1(|Si|+1), ma=N/Z and mm=maxi=1,2,… Z{|Si|}. Space complexity is O(N). These can be compared with the O(N2malog ma) average and O(N3log mm) worst case time required by the best available ACPM algorithm. Donald A. Adjeroh |
Comput. J. | 2 |
| 2012 | Parameterized longest previous factor
Richard Beal, Donald A. Adjeroh |
Theor. Comput. Sci. | 2 |
| 2012 | Robust Color Texture Features Under Varying Illumination ConditionsabstractUnder varying illumination, both the statistical and structural contents of color texture are modified, leading to changes in the observed texture surface. We model the effect of illumination as a perturbation on an ideal color texture and show that the spectra of the ambient light have a significant impact on the observed texture patterns in the individual color channels. Motivated by studies in human color constancy, we propose a correlation-based transformation that minimizes the effect of illumination variation in color texture analysis. Experimental results are included, which validate the performance of the proposed minvariance model in the analysis of color texture. Umasankar Kandaswamy, Donald A. Adjeroh, Stephanie Schuckers, Allan Hanbury |
IEEE Trans. Syst. Man Cybern. Part B | 2 |
| 2011 | Can facial metrology predict gender?abstractWe investigate the question of whether facial metrology can be exploited for reliable gender prediction. A new method based solely on metrological information from facial landmarks is developed. Here, metrological features are defined in terms of specially normalized angle and distance measures and computed based on given landmarks on facial images. The performance of the proposed metrology- based method is compared with that of a state-of-the-art appearance-based method for gender classification. Results are reported on two standard face databases, namely, MUCT and XM2VTS containing 276 and 295 images, respectively. The performance of the metrology-based approach was slightly lower than that of the appearance- based method by only about 3.8% for the MUCT database and about 5.7% for the XM2VTS database. Deng Cao, Cunjian Chen, Marco Piccirilli, Donald A. Adjeroh, Thirimachos Bourlai, Arun Ross |
IJCB | 4 |
| 2011 | On lossless image compression using the Burrows-Wheeler TransformabstractThe Burrows-Wheeler Transform (BWT) is known to be very effective in compressing text data. However, there is still a debate on its performance on images. Motivated by a theoretical analysis of the performance of BWT and MTF, we perform a detailed empirical study on the role of MTF in compressing images with the BWT. We propose two schemes for BWT-based image coding, namely BLIC and BLICX, the later being based on the context-ordering property of the BWT. Experimental results using a set of standard test images show that each of the proposed methods outperformed state-of-the-art lossless image coders, such as CALIC, JPEG-LS, and PPAM. Donald A. Adjeroh, Kalyan V. Bhupathiraju |
ICIP | 1 |
| 2011 | Parameterized Longest Previous Factor
Richard Beal, Donald A. Adjeroh |
IWOCA | 2 |
| 2011 | p-Suffix Sorting as Arithmetic Coding
Richard Beal, Donald A. Adjeroh |
IWOCA | 2 |
| 2011 | Random KNN feature selection - a fast and stable alternative to Random ForestsabstractBACKGROUND: Successfully modeling high-dimensional data involving thousands of variables is challenging. This is especially true for gene expression profiling experiments, given the large number of genes involved and the small number of samples available. Random Forests (RF) is a popular and widely used approach to feature selection for such "small n, large p problems." However, Random Forests suffers from instability, especially in the presence of noisy and/or unbalanced inputs. RESULTS: We present RKNN-FS, an innovative feature selection procedure for "small n, large p problems." RKNN-FS is based on Random KNN (RKNN), a novel generalization of traditional nearest-neighbor modeling. RKNN consists of an ensemble of base k-nearest neighbor models, each constructed from a random subset of the input variables. To rank the importance of the variables, we define a criterion on the RKNN framework, using the notion of support. A two-stage backward model selection method is then developed based on this criterion. Empirical results on microarray data sets with thousands of variables and relatively few samples show that RKNN-FS is an effective feature selection approach for high-dimensional data. RKNN is similar to Random Forests in terms of classification accuracy without feature selection. However, RKNN provides much better classification accuracy than RF when each method incorporates a feature-selection step. Our results show that RKNN is significantly more stable and more robust than Random Forests for feature selection when the input data are noisy and/or unbalanced. Further, RKNN-FS is much faster than the Random Forests feature selection method (RF-FS), especially for large scale problems, involving thousands of variables and multiple classes. CONCLUSIONS: Given the superiority of Random KNN in classification performance when compared with Random Forests, RKNN-FS's simplicity and ease of implementation, and its superiority in speed and stability, we propose RKNN-FS as a faster and more stable alternative to Random Forests in classification problems involving feature selection for high-dimensional datasets. Shengqiao Li, E. James Harner, Donald A. Adjeroh |
BMC Bioinform. | 3 |
| 2011 | Comparison of Texture Analysis Schemes Under Nonideal ConditionsabstractSeveral recent advancements in the field of texture analysis prompt some fundamental questions. For instance, what is the true impact of these novel advancements under real-world environments? When do these novel advancements fail to perform? Which methods perform better and under what conditions? In this work, we investigate these and other issues under nonideal image acquisition environments, specifically, environments with changing conditions due to illumination variations and those caused by both affine and nonaffine transformations. We study the performance of nine popular texture analysis algorithms using three different datasets, with varying levels of difficulty. Experiments are performed on nonideal texture datasets under five different setups. We find that most state-of-the-art techniques do not perform well under these conditions. To a large extent, their performance under nonideal conditions depends critically on the nature of the textural surface. Moreover, most techniques fail to perform reliably when the number of classes in the dataset is increased significantly, over the regular-size datasets used in previous work. Multiscale features performed reasonably well against variations caused by illumination and rotation but are prone to fail under changes in scale. Surprisingly, the performance for most of the algorithms is generally stable on structured or periodic textures, even with variations in illumination or affine transformations. Umasankar Kandaswamy, Stephanie Schuckers, Donald A. Adjeroh |
IEEE Trans. Image Process. | 3 |
| 2010 | A statistical approach for shadow detection using spatio-temporal contextsabstractBackground subtraction is an important step used to segment moving regions in surveillance videos. However, cast shadows are often falsely labeled as foreground objects, which may severely degrade the accuracy of object localization and detection. Effective shadow detection is necessary for accurate foreground segmentation, especially for outdoor scenes. Based on the characteristics of shadows, such as luminance reduction, chromaticity invariance and texture invariance, we introduce a nonparametric framework for modeling surface behavior under cast shadows. To each pixel, we assign a potential shadow value with a confidence weight, indicating the probability that the pixel location is an actual shadow point. Given an observed RGB value for a pixel in a new frame, we use its recent spatio-temporal context to compute an expected shadow RGB value. The similarity between the observed and the expected shadow RGB values determines whether a pixel position is a true shadow. Experimental results show the performance of the proposed method on a suite of standard indoor and outdoor video sequences. Donald A. Adjeroh |
ICIP | 2 |
| 2010 | A Novel Congestion Control Protocol for Vital Signs Monitoring in Wireless Biomedical Sensor NetworksabstractRecent developments in biosensor and wireless technology have led to rapid progress in wearable real time health monitoring. Unlike in wired networks, wireless networks are more subject to packet loss and congestion. In this paper, we propose a congestion control and service prioritization protocol for real time monitoring of patients' vital signs using wireless biomedical sensor networks. The proposed system is able to discriminate between different physiological signals and assign them different priorities. Thus, it would be possible to provide better quality of service for transmitting highly important vital signs. Congestion control is performed by considering both the congestion situation in the parent node and the priority of the child nodes in assigning network bandwidth to signals from different patients. Given the dynamic nature of a patient's health condition, the proposed system can detect an anomaly in the received vital signs from a patient and hence assign more priority to patients in need. Simulation results confirm the superior performance of the proposed protocol. To our knowledge, this is the first attempt at a special-purpose congestion control protocol specifically designed for wireless biomedical sensor networks. Mohammad Hossein Yaghmaee Moghaddam, Donald A. Adjeroh |
WCNC | 2 |
| 2009 | Priority-based rate control for service differentiation and congestion control in wireless multimedia sensor networks
Mohammad Hossein Yaghmaee Moghaddam, Donald A. Adjeroh |
Comput. Networks | 2 |
| 2008 | Suffix Sorting via Shannon-Fano-Elias CodesabstractWe propose two algorithms for the direct suffix sorting problem. The first is a simple algorithm that runs in an O(n) average time and space complexity, but with a worst case complexity in O(n log n) time and O(n) space. The second algorithm improves the first algorithm to O(n) time and space in the worst case. The improved algorithm requires only 7n bytes of storage, including the n bytes for the original string, and the 4n bytes for the suffix array. We take a general divide and conquer approach: divide the sequence into two groups; construct the suffix array for the first group; construct the suffix array for the second group, based on the sorted suffix from the first; Merge the suffix arrays from the two groups to form the suffix array for the parent sequence; Perform the above steps recursively to construct the complete suffix array for the entire sequence. Given a string of length n, our algorithm runs in O(n) worst case time and space. Our algorithm differs from previous approaches in the use of a simple partitioning step, and how it exploits this simple partitioning scheme for conflict resolution, using the notion of conflict trees. The basis of our improved algorithm is an extension of Shannon-Fano-Elias codes used in information theory. The space requirement for the proposed algorithm is 7n bytes, including the n bytes required to store the original string. This is a significant improvement when compared with the 13n bytes required by the KS algorithm. The method is also unique in its use of Shannon-Fano-Elias codes in efficient suffix sorting. To our knowledge, this is the first time information- theoretic methods have been used as the basis for solving the suffix sorting problem. Donald A. Adjeroh, Fei Nan |
DCC | 1 |
| 2008 | A Model for Differentiated Service Support in Wireless Multimedia Sensor NetworksabstractDifferent network applications need different Quality of Service (QoS) requirements such as packet delay, packet loss, bandwidth and availability. It is important to develop a network architecture which is able to guaranty quality of service requirements for high priority traffic. In Wireless Multimedia Sensor Networks (WMSNs), a sensor node may have different kinds of sensor which gather different types of data, with differing levels of importance. We argue that the sensor networks should be willing to spend more resources in disseminating packets that carry more important information. Some applications of WMSNs need to send real time traffic toward the sink node. This real time traffic requires low latency and high reliability so that immediate remedial and defensive actions can be taken, where necessary. Similar to wired networks, service differentiation in wireless sensor networks is also very important. In this paper we propose a differentiated service model for WMSNs. The proposed model can provide requested quality of service for high priority real time classes. In the proposed model, we distinguish high priority real time traffic from the low priority non-real time traffic, and input traffic streams are then serviced based on their priorities. Simulation results confirm the efficiency of the proposed model. Mohammad Hossein Yaghmaee Moghaddam, Donald A. Adjeroh |
ICCCN | 2 |
| 2008 | A new priority based congestion control protocol for Wireless Multimedia Sensor NetworksabstractNew applications made possible by the rapid improvements and miniaturization in hardware has motivated recent developments in wireless multimedia sensor networks (WMSNs). As multimedia applications produce high volumes of data which require high transmission rates, multimedia traffic is usually high speed. This may cause congestion in the sensor nodes, leading to impairments in the quality of service (QoS) of multimedia applications. Thus, to meet the QoS requirements of multimedia applications, a reliable and fair transport protocol is mandatory. An important function of the transport layer in WMSNs is congestion control. In this paper, we present a new queue based congestion control protocol with priority support (QCCP-PS), using the queue length as an indication of congestion degree. The rate assignment to each traffic source is based on its priority index as well as its current congestion degree. Simulation results show that the proposed QCCP-PS protocol can detect congestion better than previous mechanisms. Furthermore it has a good achieved priority close to the ideal and near-zero packet loss probability, which make it an efficient congestion control protocol for multimedia traffic in WMSNs. As congestion wastes the scarce energy due to a large number of retransmissions and packet drops, the proposed QCCP-PS protocol can save energy at each node, given the reduced number of retransmissions and packet losses. Mohammad Hossein Yaghmaee Moghaddam, Donald A. Adjeroh |
WOWMOM | 2 |
| 2008 | 3D Axon Structure Extraction and Analysis in Confocal Fluorescence Microscopy ImagesabstractThe morphological properties of axons, such as their branching patterns and oriented structures, are of great interest for biologists in the study of the synaptic connectivity of neurons. In these studies, researchers use triple immunofluorescent confocal microscopy to record morphological changes of neuronal processes. Three-dimensional (3D) microscopy image analysis is then required to extract morphological features of the neuronal structures. In this article, we propose a highly automated 3D centerline extraction tool to assist in this task. For this project, the most difficult part is that some axons are overlapping such that the boundaries distinguishing them are barely visible. Our approach combines a 3D dynamic programming (DP) technique and marker-controlled watershed algorithm to solve this problem. The approach consists of tracking and updating along the navigation directions of multiple axons simultaneously. The experimental results show that the proposed method can rapidly and accurately extract multiple axon centerlines and can handle complicated axon structures such as cross-over sections and overlapping objects. Yong Zhang 0050, Xiaobo Zhou 0001, Ju Lu, Jeff Lichtman, Donald A. Adjeroh, Stephen T. C. Wong |
Neural Comput. | 5 |
| 2008 | Prediction by Partial Approximate Matching for Lossless Image CompressionabstractContext-based modeling is an important step in high-performance lossless data compression. To effectively define and utilize contexts for natural images is, however, a difficult problem. This is primarily due to the huge number of contexts available in natural images, which typically results in higher modeling costs, leading to reduced compression efficiency. Motivated by the prediction by partial matching context model that has been very successful in text compression, we present prediction by partial approximate matching (PPAM), a method for compression and context modeling for images. Unlike the PPM modeling method that uses exact contexts, PPAM introduces the notion of approximate contexts. Thus, PPAM models the probability of the encoding symbol based on its previous contexts, whereby context occurrences are considered in an approximate manner. The proposed method has competitive compression performance when compared with other popular lossless image compression algorithms. It shows a particularly superior performance when compressing images that have common features, such as biomedical images. Yong Zhang 0050, Donald A. Adjeroh |
IEEE Trans. Image Process. | 2 |
| 2007 | Edge-Based Prediction for Lossless Compression of Hyperspectral ImagesabstractWe present two algorithms for error prediction in lossless compression of hyperspectral images. The algorithms are context-based and non-linear, and use a one-band look-ahead, thus requiring a minimal storage buffer. The first algorithm (NPHI) predicts the pixel in the current band based on the information from its context. Prediction contexts are defined based on the neighboring causal pixels in the current band and the corresponding co-located causal pixels in the reference band. EPHI extends NPHI using edge-based analysis. Prediction is performed by classifying the pixels into edge and non-edge pixels. Each pixel is then predicted using information from pixels in the same edge class within the context. Empirical results show that the proposed methods produce competitive results when compared with other state-of-the-art algorithms with comparable complexity. On average, the edge-based technique (EPHI) produced the best overall result, over the images in the test dataset Sushil K. Jain, Donald A. Adjeroh |
DCC | 2 |
| 2006 | Optimal Coding Rate Selection for 3D Video using RCPC CodesabstractSummary form only given. We study the problem of error-resilient transmission of video data at very low bit rates, using unequal error protection (UEP). In particular, we consider a theoretical framework for the selection of parameters for optimal allocation of redundancy when of the protection method is based on rate-compatible punctured convolutional (RCPC) codes, over an AWGN channel. The theoretical results were found to be similar with optimal rate allocation and quantization in traditional lossy data compression, where transmission problems are not considered Donald A. Adjeroh |
DCC | 1 |
| 2006 | On Compressibility of Protein SequencesabstractWe consider the problem of compressibility of protein sequences. Based on an observed genome-scale long-range correlation in concatenated protein sequences from different organisms, we propose a method to exploit this unusual redundancy in compressing the protein sequences. The result is a significant reduction in the number of bits required for representing the sequences. We report results in bits per symbol (bps) of 2.27, 2.55, 3.11 and 3.44 for protein sequences from M. jannaschii, H. influenzae, S. cerevisiae, and H. sapiens respectively, the same protein sequences used by Nevill-Manning and Witten in the "Protein is incompressible" paper. The observed long-range correlations could have significant implications beyond compression and complexity analysis of protein sequences. Donald A. Adjeroh, Fei Nan |
DCC | 1 |
| 2006 | On denoising and compression of DNA microarray images
Donald A. Adjeroh, Yong Zhang 0050, Rahul Parthe |
Pattern Recognit. | 1 |
| 2005 | Prediction by Partial Approximate Matching for Lossless Image CompressionabstractSummary form only given. In this paper, motivated by the characteristics of natural images, we introduce PPAM - prediction by partial approximate matching, a modification of the context model used by the prediction by partial matching (PPM) family of text compression algorithms, for images. Unless there is a strong edge boundary, natural images usually contain locally homogeneous areas. However, the correlation within these local neighborhoods is usually not exact. Repeating contexts often occur in an approximate (rather than exact) form. In PPAM, we exploit these inexact contexts using methods from text pattern matching with errors. In PPM, context searching starts by looking for the maximum-order context and then escapes to shorter contexts until a match is found. PPAM uses a simple context quantization scheme to avoid potentially huge time and space requirements. We group the contexts based on the square of their Euclidean distances (SED) from a reference context. We tested the performance of the proposed PPAM on standard test images and good results were obtained. Yong Zhang 0050, Donald A. Adjeroh |
DCC | 2 |
| 2005 | Effective invariant features for shape-based image retrievalabstractAbstract The success of content‐based image retrieval (CBIR) relies critically on the ability to find effective image features to represent the database images. The shape of an object is a fundamental image feature and belongs to one of the most important image features used in CBIR. In this article we propose a robust and effective shape feature known as the compound image descriptor (CID), which combines the Fourier transform (FT) magnitude and phase coefficients with the global features. The underlying FT coefficients have been shown analytically to be invariant to rotation, translation, and scaling. We also present details of the underlying innovative shape feature extraction method. The global features, besides being incorporated with the FT coefficients to form the CID, are also used to filter out the highly dissimilar images during the image retrieval process. Thus, they serve a dual purpose of improving the accuracy and hence the robustness of the shape descriptor, and of speeding up the retrieval process, leading to a reduced query response time. Experiment results show that the proposed shape descriptor is, in general, robust to changes caused by image shape rotation, translation, and/or scaling. It also outperforms other recently published proposals, such as the generic Fourier descriptor (Zhang & Lu, 2002). Shan Li 0003, Moon-Chuen Lee, Donald A. Adjeroh |
J. Assoc. Inf. Sci. Technol. | 3 |
| 2005 | A comparison of BWT approaches to string pattern matchingabstractRecently a number of algorithms have been developed to search files compressed with the Burrows-Wheeler Transform (BWT) without the need for full decompression first. This allows the storage requirement of data to be reduced through the exceptionally good compression offered by BWT, while allowing fast access to the information for searching by taking advantage of the sorted nature of BWT files. We provide a detailed description of five of these algorithms: BWT-based Boyer-Moore, Binary Search, Suffix Arrays, q-grams and the FM-index, and also present results from a set of extensive experiments that were performed to evaluate and compare the algorithms. Furthermore, we introduce a technique to improve the search times of Binary Search, Suffix Arrays and q-grams by 22% on average, as well as reduce the memory requirement of the latter two by 40% and 31%, respectively. Our results indicate that, while the compressed files of the FM-index are larger than those of the other approaches, it is able to perform searches with considerably less memory. Additionally, when only counting the occurrences of a pattern, or when locating the positions of a small number of matches, it is the fastest algorithm. For larger searches, Binary Search provides the fastest results. Comparative results with non-BWT based search methods are also included. Copyright © 2005 John Wiley & Sons, Ltd. Andrew E. Firth, Timothy C. Bell, Amar Mukherjee, Donald A. Adjeroh |
Softw. Pract. Exp. | 4 |
| 2005 | Efficient Texture Analysis of SAR ImageryabstractWe address the problem of efficiency in texture analysis for synthetic aperture radar (SAR) imagery. Motivated by the statistical occupancy model, we introduce the notion of patch reoccurrences. Using the reoccurrences, we propose the use of approximate textural features in analysis of SAR images. We describe how the proposed approximate features can be extracted for two popular texture analysis methods-the gray-level cooccurrence matrix and Gabor wavelets. Results on image texture classification show that the proposed method can provide an improved efficiency in the analysis of SAR imagery, without introducing any significant degradation in the classification results. Umasankar Kandaswamy, Donald A. Adjeroh, Moon-Chuen Lee |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2004 | Scene-adaptive transform domain video partitioningabstractAn adaptive mechanism for video partitioning in the transform domain is proposed. Different quantitative measures for motion complexity and activity levels in a scene are defined, based on which a video scene can be consistently categorised into identifiable classes. Further, the video quality, as measured by the mean square error (MSE) is related to certain parameters used in video partitioning. Adaptability is realized by tailoring the parameters of the video partitioning algorithm to the specific characteristics of the video scene, as embodied in the video scene class and the video quality. Experimental results are included. Donald A. Adjeroh, Moon-Chuen Lee |
IEEE Trans. Multim. | 1 |
| 2003 | Approximate Pattern Matching Using the Burrows-Wheeler TransformabstractSummary form only given. The approximate pattern matching on the text transformed by the Burrows-Wheeler transform (BWT) was considered. This is an important first step towards developing a compressed pattern matching algorithm for the BWT based compression system. Algorithms are proposed to solve the K-mismatch problem. Tests were performed on different pattern lengths using 133 selected files from the Canterbury, Calgary, and TREC corpus. The results on the K-mismatch pattern matching show that the running time and storage are superior to the fast suffix tree approach. Thus, once the index arrays are created, for repeated pattern search operations and for long patterns, the proposed algorithms perform significantly better than the agrep and ngrep. Using DFA verification, the search time is almost constant. The amortized cost is lower for multiple patterns search operations. Nan Zhang 0005, Amar Mukherjee, Donald A. Adjeroh, Timothy C. Bell |
DCC | 3 |
| 2002 | Pattern Matching in BWT-Transformed TextabstractSummary form only given. The compressed pattern matching problem is to locate the occurrence(s) of a pattern P in a text string T using a compressed representation of T, with minimal (or no) decompression. The BWT performs a permutation of the characters in the text, such that characters in lexically similar contexts will be near to each other. The motivation for our approach is the observation that the BWT provides a lexicographic ordering of the input text as part of its inverse transformation process. Donald A. Adjeroh, Timothy C. Bell, Matt Powell, Nan Zhang 0005, Amar Mukherjee |
DCC | 1 |
| 2002 | Searching BWT Compressed Text with the Boyer-Moore Algorithm and Binary SearchabstractThis paper explores two techniques for on-line exact pattern matching in files that have been compressed using the Burrows-Wheeler transform. We investigate two approaches. The first is an application of the Boyer-Moore algorithm (1977) to a transformed string. The second approach is based on the observation that the transform effectively contains a sorted list of all substrings of the original text, which can be exploited for very rapid searching using a variant of binary search. Both methods are faster than a decompress-and-search approach for small numbers of queries, and binary search is much faster even for large numbers of queries. Timothy C. Bell, Matt Powell, Amar Mukherjee, Donald A. Adjeroh |
DCC | 4 |
| 2001 | On ratio-based color indexingabstractThe color ratio approach to indexing has been found to be robust and effective in indexing image and video databases, in different color spaces, and when using transformed color features, such as those from the Karhunen-Loeve transform (KLT) or the discrete cosine transform (DCT). However, the reason for the superior performance of the color ratio model, especially on different color spaces or with transformed color features has, at best, been speculative. This paper develops a generalized form for the color ratio model, based on which we characterize the general distribution of the color ratios. From the distribution, we present a theory that explains and supports the performance of the color ratio approach in image and video indexing. It is shown that the same theory accounts for its effectiveness in different color spaces and in the transform domain. Some general problems encountered in using the original retinex lightness algorithm, and some other issues specific to ratio-based color indexing are discussed in the light of the theory. Results are presented which show that the proposed theory is supported by empirical evidence. Donald A. Adjeroh, Moon-Chuen Lee |
IEEE Trans. Image Process. | 1 |
| 2000 | An occupancy model for image retrieval and similarity evaluationabstractThe analysis of visual information often involves the manipulation of enormous volumes of data. If some tolerance is allowed in the results, orders of magnitude improvement in efficiency can be achieved in such analysis by appropriate selective processing, without necessarily considering all the data features. To guarantee that the error introduced does not exceed the allowed limit, a certain minimum proportion of the data must be involved in the analysis. This proportion cannot be determined arbitrarily. It should be chosen based on some formal methods, with a consideration of the error inherent in the data. This paper presents some techniques for improving the retrieval efficiency in image-based information systems, with performance guarantees on the reliability of results. Using the statistical theory of occupancy, it develops a model for the formal selection of the minimal subset of image features to be involved in histogram-based similarity evaluation. This guarantees that decisions based on the minimum proportion are always the same as (or close to) the one that would have been reached by considering all the features. Results on real and simulated data show the performance of the model on speedup, robustness, scalability, and performance guarantees. Donald A. Adjeroh, Moon-Chuen Lee |
IEEE Trans. Image Process. | 1 |
| 1999 | A Distance Measure for Video Sequences
Donald A. Adjeroh, Moon-Chuen Lee, Irwin King |
Comput. Vis. Image Underst. | 1 |
| 1998 | Design and Configuration Rationales for Digital Video Storage and Delivery Systems
Cyril U. Orji, Donald A. Adjeroh, Patrick O. Bobbie, Kingsley C. Nwosu |
Multim. Tools Appl. | 2 |
| 1997 | Robust and Efficient Transform Domain Video Sequence Analysis: An Approach from the Generalized Color Ratio Model
Donald A. Adjeroh |
J. Vis. Commun. Image Represent. | 1 |
| 1997 | Quantization of 3D-DCT Coefficients and Scan Order for Video Compression
Raymond K. W. Chan, Donald A. Adjeroh |
J. Vis. Commun. Image Represent. | 3 |
| 1997 | Techniques for Fast Partitioning of Compressed and Uncompressed Video
Donald A. Adjeroh, Moon-Chuen Lee, Cyril U. Orji |
Multim. Tools Appl. | 1 |