VLDB 2026 Research / reviewers in the wild / expert
Mihail Popescu
dblp:55/1120
· DBLP profile ↗
87ranked-venue papers
21as first author
17since 2021 · last 2026
0000-0002-6145-8096ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 45 · 8 first-author · 12 since 2021Artificial intelligence and machine learning · 42 · 13 first-author · 5 since 2021Databases, data management, data science and information retrieval · 3Graphics, computer vision, multimedia, augmented reality and games · 1Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Applications of Large Language Models and Prompt Optimization for Knowledge Extraction From Biological Pathway FiguresabstractRecent developments in Large Language Models (LLMs) have demonstrated remarkable capabilities for image comprehension. This study aims to automate and enhance the extraction of gene interactions from biological pathway images by integrating LLMs and a Genetic Algorithm (GA). A dataset of 200 tumor signaling pathway figures from the recent biological literature was employed to assess the performance of four AI chatbots: GPT-4oV, Claude-3.5V, Gemini-1.5V, and Llama-3.2V, with GA used to optimize prompts for each model. Model performance was evaluated on both directional and non-directional gene relationship extraction. GA-optimized prompts significantly improved extraction accuracies across all LLMs, with GPT 4oV achieving an F1-score of 0.645 (±0.055) and Llama-3.2V achieving an F1-score of 0.616 (±0.068). For non-directional interactions, GPT-4oV outperformed other models, reaching a precision of 0.805, a recall of 0.695, and an F1 score of 0.757, followed by Llama-3.2V and Claude-3.5V with F1-scores of 0.702 and 0.697, respectively, while Gemini-1.5V lagged with 0.612. In directional interaction predictions, all models performed lower, with GPT-4oV leading at 0.687 F1-score, followed by Llama-3.2V at 0.656, Claude-3.5V at 0.641, and Gemini-1.5V at 0.573. While these results demonstrate substantial improvements over traditional OCR-based approaches, further advances in model accuracy and explainability are needed for widespread adoption in critical biomedical applications. Nevertheless, these findings provide a valuable benchmark for the research community and a foundation for future development of specialized, fine-tuned models and scalable multimodal AI frameworks in biomedical data analysis. The source code is publicly available on https://github.com/Muh-aza/LLM_GPV. Hasanain Aldihis, Mihail Popescu, Dong Xu 0002 |
IEEE J. Biomed. Health Informatics | 4 |
| 2025 | Development and evaluation of a 4M taxonomy from nursing home staff text messages using a fine-tuned generative language modelabstractOBJECTIVE: This study aimed to explore the utilization of a fine-tuned language model to extract expressions related to the Age-Friendly Health Systems 4M Framework (What Matters, Medication, Mentation, and Mobility) from nursing home worker text messages, deploy automated mapping of these expressions to a taxonomy, and explore the created expressions and relationships. MATERIALS AND METHODS: The dataset included 21 357 text messages from healthcare workers in 12 Missouri nursing homes. A sample of 860 messages was annotated by clinical experts to form a "Gold Standard" dataset. Model performance was evaluated using classification metrics including Cohen's Kappa (κ), with κ ≥ 0.60 as the performance threshold. The selected model was fine-tuned. Extractions were clustered, labeled, and arranged into a structured taxonomy for exploration. RESULTS: The fine-tuned model demonstrated improved extraction of 4M content (κ = 0.73). Extractions were clustered and labeled, revealing large groups of expressions related to care preferences, medication adjustments, cognitive changes, and mobility issues. DISCUSSION: The preliminary development of the 4M model and 4M taxonomy enables knowledge extraction from clinical text messages and aids future development of a 4M ontology. Results compliment themes and findings in other 4M research. CONCLUSION: This research underscores the need for consensus building in ontology creation and the role of language models in developing ontologies, while acknowledging their limitations in logical reasoning and ontological commitments. Further development and context expansion with expert involvement of a 4M ontology are necessary. Matthew Steven Farmer, Mihail Popescu, Kimberly R. Powell |
J. Am. Medical Informatics Assoc. | 2 |
| 2024 | Evaluation and Integration of Advanced AI Chatbots for Biological Pathway CurationabstractThe rapid expansion of biological literature presents significant challenges in manually curating pathway knowledge from images for biological and medical research. Recent advancements in AI, particularly multimodal AI chatbots like GPT-4V, Claude, and Gemini, offer promising capabilities for image understanding and extraction of biological interactions. This study evaluates the effectiveness of these AI chatbots in curating gene interactions from 50 curated tumor signaling pathway images. Our results highlight the superior performance of GPT-4V, achieving a precision of 0.778, a recall of 0.664, an F1-score of 0.716 in non-directional analysis, and a precision of 0.566, a recall of 0.471, an F1-score of 0.512 in directional analysis. Claude followed with a precision of 0.735, recall of 0.594, and F1-score of 0.6571 in non-directional analysis, and a precision of 0.500, recall of 0.400, and F1-score of 0.445 in directional analysis. Gemini recorded the lowest performance metrics across both analyses. We further developed a model fusion integrating Claude-3 chatbots and our pathway curation pipeline, achieving notable improvements. Krpns Santosh, Dong Xu 0002, Mihail Popescu |
BIBM | 5 |
| 2024 | Predicting Gene Relations with a Graph Transformer Network Integrating DNA, Protein, and Descriptive DataabstractGene relation prediction is crucial for understanding cancer pathways and developing targeted treatments. This study proposes a novel Graph Transformer Network to predict gene relations by integrating DNA sequences, amino acids sequences, and gene descriptions. Utilizing the KEGG pathway database, our model aggregates diverse gene information to enhance gene regulatory pathway predictions. We encoded DNA sequences using DNA-BERT2, protein sequences using ESM2, and generated gene descriptions using ChatGPT, subsequently encoded by Bio-BERT. These embeddings were integrated and processed through a graph transformer with multi-head and single-head attention layers. Our results demonstrate that incorporating protein and description embeddings significantly improves predictive performance, while including DNA sequences shows marginal impact. The proposed model outperforms traditional machine learning and graph neural network-based methods, achieving an F-1 score of 0.893 on identifying whether two genes interact. This study highlights the potential of integrating multi-faceted genomic data to advance gene relation prediction, offering valuable insights for precision medicine and bioinformatics. Dong Xu 0002, Richard D. Hammer, Mihail Popescu |
BIBM | 4 |
| 2024 | pathCLIP: Detection of Genes and Gene Relations From Biological Pathway Figures Through Image-Text Contrastive LearningabstractIn biomedical literature, biological pathways are commonly described through a combination of images and text. These pathways contain valuable information, including genes and their relationships, which provide insight into biological mechanisms and precision medicine. Curating pathway information across the literature enables the integration of this information to build a comprehensive knowledge base. While some studies have extracted pathway information from images and text independently, they often overlook the correspondence between the two modalities. In this paper, we present a pathway figure curation system named pathCLIP for identifying genes and gene relations from pathway figures. Our key innovation is the use of an image-text contrastive learning model to learn coordinated embeddings of image snippets and text descriptions of genes and gene relations, thereby improving curation. Our validation results, using pathway figures from PubMed, showed that our multimodal model outperforms models using only a single modality. Additionally, our system effectively curates genes and gene relations from multiple literature sources. Two case studies on extracting pathway information from literature of non-small cell lung cancer and Alzheimer's disease further demonstrate the usefulness of our curated pathway information in enhancing related pathways in the KEGG database. Fei He 0003, Richard D. Hammer, Dong Xu 0002, Mihail Popescu |
IEEE J. Biomed. Health Informatics | 7 |
| 2022 | Exploration of Disease Pathways Using Simple Graph Comparison Methods
Mikhail Kovalenko, Mihail Popescu |
AMIA | 2 |
| 2022 | Comprehensive Assessment of OCR Tools for Gene Name Recognition in Biological Pathway FiguresabstractOptical Character Recognition (OCR) is becoming more and more effective in text detection in images. However, OCR’s performance in special applications may vary. In particular, OCR in visual representations of complex processes known as pathway figures in the biomedical literature is challenging. The information depicted in a pathway graphic usually represents the article’s most important conclusions. Still, the huge number of pathway figures cannot be automatically processed for large-scale search, data mining, and downstream analysis. Assisted by recent developments in OCR, we have developed a method to extract gene names from pathway images. For usage in the method, we thoroughly evaluated and compared major available OCR tools using 563 genes from 45 pathway images, 1000 images of alphanumeric characters and gene names from HUGO, and KEGG data with 20 random pathway genes curated routes. Our study showed that Google Cloud Vision and MMOCR are best suitable for gene name recognition in pathway figures. Stuart Aldrich, Micheal Olaolu Arowolo, Fei He 0003, Mihail Popescu, Dong Xu 0002 |
BIBM | 4 |
| 2022 | Integration of Gene Regulatory Pathways Found in the Literature in a Graph DatabaseabstractGene regulatory pathways plays a significant role in personalized medicine. Because pathways often branch or converge, graphs are better at describing them. In this work, we introduced Neo4j, a graph database, to represent gene pathways collected from the KEGG Pathway database and from 217 non-small cell lung cancer (NSCLC) articles retrieved from PubMed. We found that the graph representation of disease pathways was able to display regulatory relations between two genes even though they were not mentioned in the same article. Besides, contradictory pathway relations, for example “EGFR activates RAS” and “EGFR inhibits RAS”, self-activation or self-inhibition like “AKT activates AKT”, and pathways with different direction like “EGFR activates RAS” and “RAS activates EGFR” may also be investigated. This contradictory information may come from different conditions of the relations, different scientific opinions or they be mistakes made by our deep learning pipeline that extracts information from the literature. Overall, we concluded that graph representations of disease pathways can give us a panoramic view of the pathway knowledge obtained from various kinds of sources. Although the representation may not manifest real underlying mechanisms, it helps provide research focus and propel the development of personalized medicine. Dong Xu 0002, Mihail Popescu |
BIBM | 3 |
| 2022 | Extraction of Gene Regulatory Relation Using BioBERTabstractRelation Extraction (RE) is a critical task typically carried out after Named Entity recognition for identifying gene-gene association from scientific publication. Current state-of the-art tools have limited capacity as most of them only extract entity relations from abstract texts. The retrieved gene-gene relations typically do not cover gene regulatory relations. There still exists a lot of room for improvement. In this work, we propose GeREx, a transformer based RE tool for identifying gene regulatory relations from full texts by incorporating labeling from pathway figures. GeREx achieved an F1-Score of 83.45% when evaluated on an independent test dataset. Clement Essien, Fei He 0003, Mark Hannink, Mihail Popescu, Dong Xu 0002 |
BIBM | 4 |
| 2022 | Explainable AI for Early Detection of Health Changes Via Streaming ClusteringabstractThe ability to explain the predictions of machine learning models has become increasingly important, especially in healthcare applications. Streaming clustering is an effective tool to recognize normal baseline patterns and to detect early signs of changes in data streams. However, many streaming clustering algorithms are not designed to explain to the users how predictions are made. In this paper, we extend a streaming clustering algorithm, the sequential possibilistic Gaussian mixture model (SPGMM) for early detection of health change to provide algorithm explainability for the results. Four approaches are discussed to explain either the cluster differences or the reason for the algorithm warnings: (i) linguistic summarization for warnings; (ii) annotation distribution of clusters; (iii) SHapley Additive exPlanations (SHAP); (iv) functional health score. The four approaches are validated on one older adult monitored with a collection of motion, bed, and depth sensors over three years. The results obtained on the older adult show that the four approaches aid understanding of how the clusters and warnings are generated, providing strong support for clinicians to take corresponding actions. James Keller 0001, Marjorie Skubic, Mihail Popescu |
FUZZ-IEEE | 4 |
| 2022 | Patient judgments about hypertension control: the role of patient numeracy and graph literacyabstractOBJECTIVE: To assess the impact of patient health literacy, numeracy, and graph literacy on perceptions of hypertension control using different forms of data visualization. MATERIALS AND METHODS: Participants (Internet sample of 1079 patients with hypertension) reviewed 12 brief vignettes describing a fictitious patient; each vignette included a graph of the patient's blood pressure (BP) data. We examined how variations in mean systolic blood pressure, BP standard deviation, and form of visualization (eg, data table, graph with raw values or smoothed values only) affected judgments about hypertension control and need for medication change. We also measured patient's health literacy, subjective and objective numeracy, and graph literacy. RESULTS: Judgments about hypertension data presented as a smoothed graph were significantly more positive (ie, hypertension deemed to be better controlled) then judgments about the same data presented as either a data table or an unsmoothed graph. Hypertension data viewed in tabular form was perceived more positively than graphs of the raw data. Data visualization had the greatest impact on participants with high graph literacy. DISCUSSION: Data visualization can direct patients to attend to more clinically meaningful information, thereby improving their judgments of hypertension control. However, patients with lower graph literacy may still have difficulty accessing important information from data visualizations. CONCLUSION: Addressing uncertainty inherent in the variability between BP measurements is an important consideration in visualization design. Well-designed data visualization could help to alleviate clinical uncertainty, one of the key drivers of clinical inertia and uncontrolled hypertension. Victoria A. Shaffer, Pete Wegier, K. D. Valentine, Sean Duan, Shannon Canfield, Jeffery L. Belden, Linsey M. Steege, Mihail Popescu, Richelle J. Koopman |
J. Am. Medical Informatics Assoc. | 8 |
| 2022 | New Linguistic Description Approach for Time Series and Its Application to Bed Restlessness Monitoring for EldercareabstractTime series analysis has been an active area of research for years, with important applications in forecasting or discovery of hidden information such as patterns or anomalies in observed data. In recent years, the use of time series analysis techniques for the generation of descriptions and summaries in natural language of any variable, such as temperature, heart rate or CO2 emission has received increasing attention. Natural language has been recognized as more effective than traditional graphical representations of numerical data in many cases, in particular in situations where a large amount of data needs to be inspected or when the user lacks the necessary background and skills to interpret it. In this work, we describe a novel mechanism to generate linguistic descriptions of time series using natural language and fuzzy logic techniques. The proposed method generates quality summaries capturing the time series features that are relevant for a user in a particular application, and can be easily customized for different domains. This approach has been successfully applied to the generation of linguistic descriptions of bed restlessness data from residents at TigerPlace (Columbia, Missouri), which is used as a case study to illustrate the modeling process and show the quality of the descriptions obtained. Carmen Martínez-Cruz, Antonio J. Rueda Ruiz, Mihail Popescu, James Keller 0001 |
IEEE Trans. Fuzzy Syst. | 3 |
| 2021 | Identifying Genes and Their Interactions from Pathway Figures and Text in Biomedical ArticlesabstractMany high-quality biological pathways are presented in figures and text in biomedical literature. They are great resources for studies of biological mechanisms and precision medicine practices. These pathways need to be carefully curated, reconciled, and transformed into a computable form. Current manual curation approaches are inadequate in keeping up with the pace of the literature growth. New bio-curation approaches are needed to streamline the identification of gene interactions from pathway figures and text. This paper proposes a pathway curation approach for identifying genes and their interactions using both figures and text of biomedical articles. Our method integrates deep learning-based object detection models with a Google optical character recognition service to extract genes and their interactions from pathway figures. Our pipeline was evaluated on the figures from PubMed publications with manual annotations. The results demonstrated that our model could effectively retrieve genes and their interactions in pathway figures. The proposed pipeline may accelerate various applications of the latest biomedical discoveries. We also developed a web server at http://pathwaydeep.top to provide the gene interaction curation on uploaded pathway figures and corresponding articles. Fei He 0005, Joshua Thompson, Ziting Mao, Yijie Ren, Yulia I. Nussbaum, Olha Kholod, Dmitriy Shin, Mark Hannink, Mihail Popescu, Dong Xu 0002 |
BIBM | 9 |
| 2021 | Early Detection of Health Changes in the Elderly Using In-Home Multi-Sensor Data StreamsabstractThe rapid aging of the population worldwide requires increased attention from healthcare providers and the entire society. For the elderly to live independently, many health issues related to old age, such as frailty and risk of falling, need increased attention and monitoring. When monitoring daily routines for older adults, it is desirable to detect the early signs of health changes before serious health events, such as hospitalizations, happen so that timely and adequate preventive care may be provided. By deploying multi-sensor systems in homes of the elderly, we can track trajectories of daily behaviors in a feature space defined using the sensor data. In this article, we investigate a methodology for tracking the evolution of the behavior trajectories over long periods (years) using high-dimensional streaming clustering and provide very early indicators of changes in health. If we assume that habitual behaviors correspond to clusters in feature space and diseases produce a change in behavior, albeit not highly specific, tracking trajectory deviations can provide hints of early illness. Retrospectively, we visualize the streaming clustering results and track how the behavior clusters evolve in feature space with the help of two dimension-reduction algorithms: Principal Component Analysis and t-distributed Stochastic Neighbor Embedding. Moreover, our tracking algorithm in the original high-dimensional feature space generates early health warning alerts if a negative trend is detected in the behavior trajectory. We validated our algorithm on synthetic data and tested it on a pilot dataset of four TigerPlace residents monitored with a collection of motion, bed, and depth sensors over 10 years. We used the TigerPlace electronic health records to understand the residents’ behavior patterns and to evaluate the health warnings generated by our algorithm. The results obtained on the TigerPlace dataset show that most of the warnings produced by our algorithm can be linked to health events documented in the electronic health records, providing strong support for a prospective deployment of the approach. James Keller 0001, Marjorie Skubic, Mihail Popescu, Kari Lane |
ACM Trans. Comput. Heal. | 4 |
| 2021 | Comments on "TLPCM: Transfer Learning Possibilistic C-Means": Errata and ObservationsabstractThe purpose of this article is twofold. First, and foremost, it fixes an error that somehow made it through all the reviewing, both by the authors and the referees. Second, it provides insights into the meaning and variation of the main PCM parameters in this approach to transfer clustering. Rayan Gargees, James Keller 0001, Mihail Popescu |
IEEE Trans. Fuzzy Syst. | 3 |
| 2021 | TLPCM: Transfer Learning Possibilistic $C$-MeansabstractTraditional machine learning and data mining have made tremendous progress in many knowledge-based areas, such as clustering, classification, and regression. However, the primary assumption in all of these areas is that the training and testing data should be in the same domain and have the same distribution. This assumption is difficult to achieve in real-world applications due to the limited availability of labeled data. Associated data in different domains can be used to expand the availability of prior knowledge about future target data. In recent years, transfer learning has been used to address such cross-domain learning problems by using information from data in a related domain and transferring that data to the target task. In this article, a transfer-learning possibilistic c-means (TLPCM) algorithm is proposed to handle the PCM clustering problem in a domain that has insufficient data. Moreover, TLPCM overcomes the problem of differing numbers of clusters between the source and target domains. The proposed algorithm employs the historical cluster centers of the source data as a reference to guide the clustering of the target data. The experimental studies presented here were thoroughly evaluated, and they demonstrate the advantages of TLPCM in both synthetic and real-world transfer datasets. Rayan Gargees, James Keller 0001, Mihail Popescu |
IEEE Trans. Fuzzy Syst. | 3 |
| 2021 | Extending the Morphological Hit-or-Miss Transform to Deep Neural NetworksabstractWhile most deep learning architectures are built on convolution, alternative foundations such as morphology are being explored for purposes such as interpretability and its connection to the analysis and processing of geometric structures. The morphological hit-or-miss operation has the advantage that it considers both foreground information and background information when evaluating the target shape in an image. In this article, we identify limitations in the existing hit-or-miss neural definitions and formulate an optimization problem to learn the transform relative to deeper architectures. To this end, we model the semantically important condition that the intersection of the hit and miss structuring elements (SEs) should be empty and present a way to express Don't Care (DNC), which is important for denoting regions of an SE that are not relevant to detecting a target pattern. Our analysis shows that convolution, in fact, acts like a hit-to-miss transform through semantic interpretation of its filter differences. On these premises, we introduce an extension that outperforms conventional convolution on benchmark data. Quantitative experiments are provided on synthetic and benchmark data, showing that the direct encoding hit-or-miss transform provides better interpretability on learned shapes consistent with objects, whereas our morphologically inspired generalized convolution yields higher classification accuracy. Finally, qualitative hit and miss filter visualizations are provided relative to single morphological layer. Muhammad Aminul Islam, Bryce Murray, Andrew R. Buck, Derek Anderson, Grant J. Scott, Mihail Popescu, James Keller 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2020 | Experiments with Maximin SamplingabstractTo apply clustering algorithms to big data, or to build clustering ensembles, it is a standard process to sample the original data set in a way that hopefully spans the original distribution. There are at least six ways to initialize the Maximin (MM) sampling algorithm. This paper contains experiments to determine whether samples produced by the six methods differ significantly; and whether they are superior to simple random sampling. Empirical evidence supports two conclusions. First, there is not enough difference in MM samples generated by the six initializations to support using any but the least costly method: viz., using the first sample in the data as the first MM point. Second, unless the input data have subsets (clusters) that are compact and separated in a well-defined sense, random sampling is demonstrably superior to MM sampling for even small data sets. Omar A. Ibrahim, James Keller 0001, James C. Bezdek, Mihail Popescu |
FUZZ-IEEE | 4 |
| 2020 | Corrigendum to "Linguistic summarization of in-home sensor data" [J. Biomed. Inf. 96 (2019) 103240]
Akshay Jain 0004, Mihail Popescu, James Keller 0001, Marilyn Rantz, Brianna Markway |
J. Biomed. Informatics | 2 |
| 2019 | Further Data Visualization Designs to Support Shared Decision Making About Blood Pressure
Jeffery L. Belden, Pete Wegier, Shannon Canfield, Sonal J. Patil, Victoria A. Shaffer, Linsey M. Steege, Mihail Popescu, K. D. Valentine, Akshay Jain 0004, Michael LeFevre, Richelle J. Koopman |
AMIA | 7 |
| 2019 | Informatics Framework to Identify Consistent Diagnostic TechniquesabstractPathologists diagnose diseases according to diagnostic clues (DCs) such as clinical features, morphological features, and expression of antigens on the surface of a cell population. The specific use of these complex DCs constitutes a diagnostic technique. However, interpretation of these DCs may not be the same by all pathologists, subtle differences could lead to misdiagnosis or diagnostic pitfalls. Given that certain DCs express similarly for different disease types, are interrelated, or missing it is challenging to analyze DCs in order to find consistent diagnostic techniques. To address that challenge, we discuss a framework based on Shannon information entropy, which groups cases that represent consistent diagnostic techniques. The analysis of DCs from 35 cases generated three groups. Evaluation studies show that the results are statistically significant. In conclusion, this framework can be useful for the analysis of sparse data and has the potential to discover groups of cases that represent consistent diagnostic heuristics. Pericles S. Giannaris, Zainab Al-Taie, Mikhail Kovalenko, Richard D. Hammer, Mihail Popescu, Dmitriy Shin |
BIBM | 5 |
| 2019 | An Unsupervised Framework for Detecting Early Signs of Illness in EldercareabstractUsing non-wearable sensors in eldercare monitoring is a promising solution for improving care and reducing healthcare costs. Abnormal sensor patterns produced by certain resident behaviors can be linked to early signs of illness. We propose an unsupervised framework for detecting abnormal sensor patterns based on clustering activity sensor sequences. We use a 30-day normal window to build a baseline model of an elderly resident by clustering the activity sequences from these days. Each cluster represents different daily activities that are performed in most (normal) days and correspond to normal routines. If a new day contains fewer routine activities, we flag it as abnormal and label the day as one with a possible sign of early illness. A preliminary analysis of the method was conducted on data collected in TigerPlace, an eldercare facility that promotes aging-in-place, with information from our electronic health records (EHR). On a pilot sensor dataset from three residents, with a total of 1902 days, we achieved an average abnormal events prediction of 0.75. Omar A. Ibrahim, James Keller 0001, Mihail Popescu |
BIBM | 3 |
| 2019 | A New Incremental Cluster Validity Index for Streaming Clustering AnalysisabstractIn this paper, we present an incremental version of the Partition Coefficient and Exponential Separation (PCAES) cluster validity index in the context of streaming data analysis. Incremental PCAES (iPCAES) can be used to monitor evolving structures in streaming data. We investigate the use of the proposed index to understand and analyze the performance of the MU Streaming Clustering (MUSC) algorithm. Synthetic and real-life streaming datasets are used to demonstrate the benefits that can be drawn from such indices such as the appearance of a new structure in the data stream, handling of outlier data samples, and the effect of the streaming sample order on the resultant cluster history. We compare the performance of iPCAES index with the incremental Davies-Boudin index (iDB) because iDB was found to be the most stable among other incremental indices that offer comparable approaches. Omar A. Ibrahim, James Keller 0001, Mihail Popescu |
FUZZ-IEEE | 3 |
| 2019 | Explainable AI For Dataset ComparisonabstractWith the increasing use of intelligent systems to make sense of data, lately, explainable AI systems are gaining a lot of traction. A distance measure that can distinguish data sets in linguistic terms can help AI systems in achieving explainability. We make use of Linguistic Protoform Summaries in tandem with Fuzzy Rules to design a system that can compare datasets numerically, as well as explain the difference in Natural Language. We validate our method with the help of synthetic data and show that it produces high correlation with the well-known Euclidean distance measure. We also employ the proposed method to explain changes in daily pulse rate measurements of an elderly resident living in a sensor equipped smart home. We postulate that the method will help in future endeavors to produce explainable pattern recognition systems. Akshay Jain 0004, James Keller 0001, Mihail Popescu |
FUZZ-IEEE | 3 |
| 2019 | Data Stream Trajectory Analysis Using Sequential Possibilistic Gaussian Mixture ModelabstractData stream processing has gained much attention lately, in the era of big data. Streaming clustering is an effective tool to recognize normal baseline and to detect outliers in sequentially presented data. Perhaps more importantly would be the ability to predict that incoming data indicates movement towards a likely anomaly. In this paper, a Gaussian Mixture Model (GMM) is employed to represent different patterns in the data stream. The Sequential Possibilistic One-Means (SP1M) is used for initialization, and is incorporated into the GMM framework to recognize new mixture components in the data stream. The new proposed algorithm is called Sequential Possibilistic Gaussian Mixture Model (SPGMM). Furthermore, two methods of trajectory analysis, the “maximum typicality decline” and the “trend value measurement,” are used together with SPGMM to detect early signs of pattern changes before unusual pattern data arrive in the stream. The proposed SPGMM is tested on synthetic and real-world datasets, and is shown to have excellent performance on predicting early signs of pattern changes in these sequential streams. James Keller 0001, Marjorie Skubic, Mihail Popescu |
FUZZ-IEEE | 4 |
| 2019 | Non-invasive Classification of Sleep Stages with a Hydraulic Bed Sensor Using Deep LearningabstractThe quality of sleep has a significant impact on health and life. This study adopts the structure of hierarchical classification to develop an automatic sleep stage classification system using ballistocardiogram (BCG) signals. A leave-one-subject-out cross validation (LOSO-CS) procedure is used for testing classification performance. Convolutional Neural Networks (CNNs), Long Short-Term Memory (LSTM), and Deep Neural Networks DNNs are complementary in their modeling capabilities; while CNNs have the advantage of reducing frequency variations, LSTMs are good at temporal modeling. A transfer learning (TL) technique is used to pre-train our CNN model on posture data and then fine-tune it on the sleep stage data. We used a ballistocardiography (BCG) bed sensor to collect both posture and sleep stage data to provide a non-invasive, in-home monitoring system that tracks changes in health of the subjects over time. Polysomnography (PSG) data from a sleep lab was used as the ground truth for sleep stages, with the emphasis on three sleep stages, specifically, awake, rapid eye movement (REM) and non-REM sleep (NREM). Our results show an accuracy of 95.3%, 84% and 93.1% for awake, REM and NREM respectively on a group of patients from the sleep lab. Rayan Gargees, James Keller 0001, Mihail Popescu, Marjorie Skubic |
ICOST | 3 |
| 2019 | Linguistic summarization of in-home sensor data
Akshay Jain 0004, Mihail Popescu, James Keller 0001, Marilyn Rantz, Brianna Markway |
J. Biomed. Informatics | 2 |
| 2018 | Early Sepsis Recognition Based on Ear Localization using Infrared Thermography
Hasanain Al-Sadr, Mihail Popescu, James Keller 0001 |
BIBM | 2 |
| 2017 | Semantic Analysis of a Nursing Home EHR for Sensor Data Annotation
Zhongkai Huang, Mihail Popescu |
AMIA | 2 |
| 2017 | Unsupervised Analysis of Activity Patterns in Eldercare Monitoring
Omar A. Ibrahim, Mihail Popescu, James Keller 0001 |
AMIA | 2 |
| 2017 | Linguistic Summarization of Sensor Data Leading to Health Events
Akshay Jain 0004, Mihail Popescu, James Keller 0001 |
AMIA | 2 |
| 2017 | Assessing Inter-Hospital Clinical Processes Variability with Process Mining
Marius Petruc, Mihail Popescu |
AMIA | 2 |
| 2017 | Linking Resident Behavior to Health Conditions in an Eldercare Monitoring System
Mihail Popescu, Andrew Craver, Lorraine Phillips, Richelle J. Koopman, Laurel Despins, Gregory L. Alexander, Marilyn Rantz |
AMIA | 1 |
| 2017 | Early illness recognition in older adults using transfer learningabstractPredicting early signs of illness in older adults by utilizing a continuous, unobtrusive nursing home monitoring system has been shown to increase the quality of life and decrease the cost of care. Illness prediction is based on sensor data such as motion and bed and uses algorithms such as support vector machine (SVM) or k-nearest neighbor (kNN). One of the greatest challenges in developing prediction algorithms for sensor networks is utilizing knowledge from previous residents for predicting behavior in new ones. In this paper, we employ a transfer learning approach for addressing the cross resident training problem. We validate our method by conducting a retrospective study on three residents from TigerPlace, a retirement community in Columbia, MO, where apartments are fitted with wireless networks of motion and bed sensors. The ground truth, the daily presence or absence of the illness, was manually evaluated using nursing visit reports from an homegrown electronic medical record (EMR) system. In this study, transfer learning SVM approach outperformed other three methods, regular SVM, one class SVM, and one class kNN, resulting in average areas under the curve (AUCs) of 0.74, 0.52, 0.60 and 0.62 for three residents, respectively. Rayan Gargees, James Keller 0001, Mihail Popescu |
BIBM | 3 |
| 2017 | Context preserving representation of daily activities in elder careabstractEldercare monitoring using non-wearable sensors is a candidate solution for improving care and reducing costs. Abnormal sensor patterns produced by certain resident behaviors could be linked to early signs of illness. We propose an unsupervised method for detecting abnormal behavior patterns based on a new context preserving representation of daily activities. A preliminary analysis of the method was conducted on data collected in TigerPlace, an eldercare facility that promotes aging-in-place. Sensors firings of each day are converted into sequences of daily activities. Using the proposed method, a day with hundreds of sequences is converted into a single data point representing that day and preserving the context of the daily routine at the same time. We obtained an average Area Under the Curve (AUC) of 0.9 in detecting days where elder adults need to be assessed. Omar A. Ibrahim, James Keller 0001, Mihail Popescu |
BIBM | 3 |
| 2016 | Messages from the technical program chairsabstractWelcome to The Annual IEEE International Conference on Computational Intelligence in Bioinformatics and Computational Biology (IEEE CIBCB 2016). This 13th edition of IEEE CIBCB is held in Chiang Mai, Thailand, and is only the second time that it is held in Asia. Nipon Theera-Umpon, Mihail Popescu, Patiwet Wuttisarnwattana |
CIBCB | 2 |
| 2016 | A temporal analysis system for early detection of health changesabstractA Gaussian mixture model (GMM), coupled with possibilistic clustering is used to build an adaptive system for analyzing streaming multi-dimensional activity feature vector with the goal of identifying signs of early diseases. The system is based on temporal analysis, including outlier detection, customization and adaption to new changes, together with the creation of new components for GMM in the case of emerging new normal patterns. On the other hand, an alert will be fired when detecting unexpected behavior patterns. When dealing with streaming data from embedded sensors in an eldercare environment, every resident has their unique behavior pattern. Therefore, number of Gaussians for the GMM needs to be individually determined. For this reason, possibilistic C-Means (PCM) and Automatic Merging possibilistic Clustering Method (AMPCM) are combined together to cluster the initial data points, detect anomalies and initialize the GMM. The system achieves our goals when tested on the synthetic datasets simulating an extended period of time. We hope that by applying the proposed system in real datasets, it will help by detecting health changes before real health issue happens. Omar A. Ibrahim, Jingyi Shao, James Keller 0001, Mihail Popescu |
FUZZ-IEEE | 4 |
| 2016 | Random projections fuzzy k-nearest neighbor(RPFKNN) for big data classificationabstractAs the number of features in pattern recognition applications continuously grows, new algorithms are necessary to reduce the dimensionality of the feature space while producing comparable results. For example, a dynamic area of research, activity recognition, produces large quantities of high-velocity, high-dimensionality data that require real time classification. While dimensionality reduction approaches such as principle component analysis (PCA) and feature selection work well for datasets of reasonable size and dimensionality, they fail on big data. A possible approach to classification of high-dimensionality datasets is to combine a typical classifier, fuzzy k-nearest neighbor in our case (FKNN), with feature reduction by random projection (RP). As opposed to PCA where one projection matrix is computed based on least square optimality, in RP, a projection matrix is chosen at random multiple times. As the random projection procedure is repeated many times, the question is how to aggregate the values of the classifier obtained in each projection. In this paper we present a fusion strategy for RP FKNN, denoted as RPFKNN. The fusion strategy is based on the class membership values produced by FKNN and classification accuracy in each projection. We test RPFKNN on several synthetic and activity recognition datasets. Mihail Popescu, James Keller 0001 |
FUZZ-IEEE | 1 |
| 2016 | Random projection below the JL limitabstractThe Johnson-Lindenstrauss (JL) lemma, with known probability, sets a lower bound q0on the dimension for which a random projection of p-dimensional vector data is guaranteed to be within (1±ε) of being an isometry in a randomly projected downspace. We study several ways to identify a “good” rogue random projection when the target downspace has dimensions below the JL limit. The tools used towards this end are Pearson and Spearman correlation coefficients, and a visual imaging method (a cluster heat map) that usually reveals cluster structure in spaces of any dimension. We use four synthetic data sets and the ubiquitous Iris data to study our procedures for tracking the reliability of RRPs. Unsurprisingly, rogue random projection is quite unpredictable. At its best, it is every bit as good as Principal Components Analysis, but at it's worst, it is awful. Pearson and Spearman correlations do signal good and bad projections, but the visual imaging method seems even more effective in determining the quality of RRPs. James C. Bezdek, Xiuyi Ye, Mihail Popescu, James Keller 0001, Alina Zare |
IJCNN | 3 |
| 2016 | A Multidimensional Time-Series Similarity Measure With Applications to Eldercare MonitoringabstractIn the last decade, data mining techniques have been applied to sensor data in a wide range of application domains, such as healthcare monitoring systems, manufacturing processes, intrusion detection, database management, and others. Many data mining techniques are based on computing the similarity between two sensor data patterns. A variety of representations and similarity measures for multiattribute time series have been proposed in the literature. In this paper, we describe a novel method for computing the similarity of two multiattribute time series based on a temporal version of Smith-Waterman (SW), a well-known bioinformatics algorithm. We then apply our method to sensor data from an eldercare application for early illness detection. Our method mitigates difficulties related to data uncertainty and aggregation that often arise when processing sensor data. The experiments take place at an aging-in-place facility, TigerPlace, located in Columbia, MO, USA. To validate our method, we used data from nonwearable sensor networks placed in TigerPlace apartments, combined with information from an electronic health record. We provide a set of experiments that investigate temporal version of SW properties, together with experiments on TigerPlace datasets. On a pilot sensor dataset from nine residents, with a total of 1902 days and around 2.1 million sensor hits of collected data, we obtained an average abnormal events prediction F-measure of 0.75. Zahra Hajihashemi, Mihail Popescu |
IEEE J. Biomed. Health Informatics | 2 |
| 2015 | Random projections fuzzy c-means (RPFCM) for big data clusteringabstractMany contemporary biomedical applications such as physiological monitoring, imaging, and sequencing produce large amounts of data that require new data processing and visualization algorithms. Algorithms such as principal component analysis (PCA), singular value decomposition and random projections (RP) have been proposed for dimensionality reduction. In this paper we propose a new random projection version of the fuzzy c-means (FCM) clustering algorithm denoted as RPFCM that has a different ensemble aggregation strategy than the one previously proposed, denoted as ensemble FCM (EFCM). RPFCM is more suitable than EFCM for big data sets (large number of points, n). We evaluate our method and compare it to EFCM on synthetic and real datasets. Mihail Popescu, James Keller 0001, James C. Bezdek, Alina Zare |
FUZZ-IEEE | 1 |
| 2015 | Recognizing complex instrumental activities of daily living using scene information and fuzzy logic
Tanvi Banerjee, James Keller 0001, Mihail Popescu, Marjorie Skubic |
Comput. Vis. Image Underst. | 3 |
| 2015 | Quantifying care coordination using natural language processing and domain-specific ontologyabstractOBJECTIVE: This research identifies specific care coordination activities used by Aging in Place (AIP) nurse care coordinators and home healthcare (HHC) nurses when coordinating care for older community-dwelling adults and suggests a method to quantify care coordination. METHODS: A care coordination ontology was built based on activities extracted from 11,038 notes labeled with the Omaha Case management category. From the parsed narrative notes of every patient, we mapped the extracted activities to the ontology, from which we computed problem profiles and quantified care coordination for all patients. RESULTS: We compared two groups of patients: AIP who received enhanced care coordination (n=217) and HHC who received traditional care (n=691) using 128,135 narratives notes. Patients were tracked from the time they were admitted to AIP or HHC until they were discharged. We found that patients in AIP received a higher dose of care coordination than HHC in most Omaha problems, with larger doses being given in AIP than in HHC in all four Omaha categories. CONCLUSIONS: 'Communicate' and 'manage' activities are widely used in care coordination. This confirmed the expert hypothesis that nurse care coordinators spent most of their time communicating about their patients and managing problems. Overall, nurses performed care coordination in both AIP and HHC, but the aggregated dose across Omaha problems and categories is larger in AIP. Lori L. Popejoy, Mohammad Khalilia, Mihail Popescu, Colleen Galambos, Vanessa Lyons, Marilyn Rantz, Lanis Hicks, Frank Stetzer |
J. Am. Medical Informatics Assoc. | 3 |
| 2014 | Automated Operative Skill Assessment Using IR Video Motion Analysis
Mihail Popescu, Christopher J. Cooper, Stephen Barnes |
AMIA | 1 |
| 2014 | Relational Fuzzy Self-Organizing Maps for Cluster Visualization and SummarizationabstractThe notion of Best-Matching Unit (BMU) in the proposed Fuzzy Relational Self-Organizing (FRSOM) algorithm is replaced by a membership function where every neuron has a certain degree of matching to an input object. The FRSOM is an extension of the relational self-organizing map. In the proposed FRSOM we incorporate a monotonically increasing fuzzifier and a monotonically decreasing neighborhood kernel. Initially, FRSOM assigns winning neurons. However, as time progresses adjacent neurons begin communicating and sharing information about the stimulus received. The amount of information being shared at a given time is governed by the fuzzifier and the number of neurons sharing information is controlled by the neighborhood kernel. Additionally, in this paper we show that FRSOM is the relational dual of Fuzzy Batch SOM (FBSOM) followed by experimental results comparing both FBSOM and FRSOM on synthetic datasets. Then we will demonstrate the visualization and summarization capabilities of FRSOM on two real relational datasets, Gene Ontology and a patient data consisting of Activity of Daily Living score trajectories. Mohammad Khalilia, Mihail Popescu |
Int. J. Uncertain. Fuzziness Knowl. Based Syst. | 2 |
| 2014 | Uncovering influence links in molecular knowledge networks to streamline personalized medicineabstractOBJECTIVES: We developed Resource Description Framework (RDF)-induced InfluGrams (RIIG) - an informatics formalism to uncover complex relationships among biomarker proteins and biological pathways using the biomedical knowledge bases. We demonstrate an application of RIIG in morphoproteomics, a theranostic technique aimed at comprehensive analysis of protein circuitries to design effective therapeutic strategies in personalized medicine setting. METHODS: RIIG uses an RDF "mashup" knowledge base that integrates publicly available pathway and protein data with ontologies. To mine for RDF-induced Influence Links, RIIG introduces notions of RDF relevancy and RDF collider, which mimic conditional independence and "explaining away" mechanism in probabilistic systems. Using these notions and constraint-based structure learning algorithms, the formalism generates the morphoproteomic diagrams, which we call InfluGrams, for further analysis by experts. RESULTS: RIIG was able to recover up to 90% of predefined influence links in a simulated environment using synthetic data and outperformed a naïve Monte Carlo sampling of random links. In clinical cases of Acute Lymphoblastic Leukemia (ALL) and Mesenchymal Chondrosarcoma, a significant level of concordance between the RIIG-generated and expert-built morphoproteomic diagrams was observed. In a clinical case of Squamous Cell Carcinoma, RIIG allowed selection of alternative therapeutic targets, the validity of which was supported by a systematic literature review. We have also illustrated an ability of RIIG to discover novel influence links in the general case of the ALL. CONCLUSIONS: Applications of the RIIG formalism demonstrated its potential to uncover patient-specific complex relationships among biological entities to find effective drug targets in a personalized medicine setting. We conclude that RIIG provides an effective means not only to streamline morphoproteomic studies, but also to bridge curated biomedical knowledge and causal reasoning with the clinical data in general. Dmitriy Shin, Gerald C. Arthur, Mihail Popescu, Dmitry Korkin, Chi-Ren Shyu |
J. Biomed. Informatics | 3 |
| 2014 | Improvements to the relational fuzzy c-means clustering algorithm
Mohammad Khalilia, James C. Bezdek, Mihail Popescu, James Keller 0001 |
Pattern Recognit. | 3 |
| 2013 | Precise Protein Expression Data Text Mining Curation and Applications
Jia-Fu Chang, Mihail Popescu, Gerald C. Arthur |
AMIA | 2 |
| 2013 | An Early Illness Recognition Framework Using a Temporal Smith Waterman Algorithm and NLP
Zahra Hajihashemi, Mihail Popescu |
AMIA | 2 |
| 2013 | A Cluster Validity Framework Based on Induced Partition DissimilarityabstractWe describe a new cluster validity framework (CVF) that compares structure in the data (in dissimilarity form) to the structure of dissimilarity matrices induced by a matrix transformation of the partition being tested. As part of this framework, we show two possible cluster validation measures: one, visual cluster validity, that that uses visual comparison and another one, correlation cluster validity, based on correlation. Unlike many existing measures, the measures we propose can be applied to crisp or soft partitions obtained by any relational or object data clustering algorithm. We illustrate the new measures and compare them to several well-known existing measures using real and artificial data sets. Mihail Popescu, James C. Bezdek, Timothy C. Havens, James Keller 0001 |
IEEE Trans. Cybern. | 1 |
| 2012 | Quantifying the quality of imaging interpretations
Marius Petruc, Mihail Popescu, John L. Fresen, Dmitriy Shin, Chi-Ren Shyu |
AMIA | 2 |
| 2012 | Developing an Ontology of Nursing Care Coordination using NLP
Lori L. Popejoy, Mihail Popescu, Mohammad Khalilia, Vanessa Lyons, Marilyn Rantz, Colleen Galambos, Lanis Hicks, Frank Stetzer |
AMIA | 2 |
| 2012 | Dirt road segmentation using color and texture features in color imageryabstractThis paper proposes a method for segmenting an unstructured dirt road in color space images using color and texture analysis. A support vector machine (SVM) classifier was trained on samples of on and off road patches from a similar road. Image patches were classified at sparse intervals at a fixed distance from the vehicle. Each patch is described by the Histogram of oriented gradients (HOG), the Local Binary Patters (LBP), a histogram of the color channel, and a histogram of a non linear color transform. The classified patches were transformed to the next frame of the sequence using the scale invariant feature transform (SIFT) to reduce reclassification of image patches. Morphological opening and closing were used to transform the points into a mask, and reduce errors. Experimental results indicated that the algorithm can accurately segment road images given a set of training data from similar road utilizing only color imagery. David B. Lewis, James Keller 0001, Mihail Popescu, Kevin E. Stone |
CISDA | 3 |
| 2012 | Detection of buried objects in FLIR imaging using mathematical morphology and SVMabstractIn this paper we describe a method for detecting buried objects of interest using a forward looking infrared camera (FLIR) installed on a moving vehicle. Infrared (IR) detection of buried targets is based on the thermal gradient between the object and the surrounding soil. The processing of FILR images consists in a spot-finding procedure that includes edge detection, opening and closing. Each spot is then described using texture features such as histogram of gradients (HOG) and local binary patterns (LBP) and assigned a target confidence using a support vector machine (SVM) classifier. Next, each spot together with its confidence is projected and summed in the UTM space. To validate our approach, we present results obtained on 6 one mile long runs recorded with a long wave IR (LWIR) camera installed on a moving vehicle. Mihail Popescu, Alex Paino, Kevin E. Stone, James Keller 0001 |
CISDA | 1 |
| 2012 | Fuzzy relational self-organizing mapsabstractIn this paper we propose a novel fuzzy relational self-organizing map algorithm (FRSOM) that can be used to map a set of n objects described by pairwise dissimilarity values to a two dimensional lattice structure. FRSOM generates a fuzzy membership matrix replacing the crisp best-matching unit matrix in the regular relational SOM (RSOM). We found that FRSOM discovers hard to find substructures in the data that present a challenge to the crisp relational SOM. Furthermore, we observed a triple relationship that seems to exist among the number of data points in the training data, map size and the fuzzifier m. We compare FRSOM and RSOM using several synthetic datasets. Mohammad Khalilia, Mihail Popescu |
FUZZ-IEEE | 2 |
| 2012 | FUMIL-Fuzzy Multiple Instance Learning for early illness recognition in older adultsabstractMany important applications in Health Sciences and Biology have underlying datasets that have ambiguous class membership, that is, individual labels are difficult to establish. In such cases, many times, the training examples are easier to label as a group rather than at the instance level. Multiple Instance Learning (MIL) is a supervised learning strategy that addresses this labeling difficulty by employing training example given as positive and negative bags of instances. In this paper we describe a fuzzy variation of the MIL Diverse Density framework (FUMIL) based on ordered weighted geometric operator (OWG) and fuzzy complement operators. We apply FUMIL for early illness recognition of elderly living alone in their home. The available data consists of wireless non-wearable sensor values aggregated at hour level (instance) and ground truth (medical data) available at day level (bag). In our preliminary experiments FUMIL performed better than the traditional MIL framework. Abhishek Mahnot, Mihail Popescu |
FUZZ-IEEE | 2 |
| 2012 | An extension of a confined space evacuation model to human geographyabstractHuman geography is a phrase that is used to indicate the augmentation of standard geographic layers of information about an area with behavioral variations of the people in the area. In particular, the actions of people can be attributed to both local and regional variations in physical (i.e., terrain) and human (e.g., income, political, cultural) variables. For example, in disaster planning, response, and relief, it is important to understand how individuals and groups of people will move in the environment. Different groups may have different objectives, the time scales are longer, and other factors like food and mobility need to be addressed. Mathematical models with graphics realizations, particularly agent-based models, are an important way to simulate human response to stress. In this paper, we extend a standard mathematical model of agent evacuation to people on an island under the threat of a hurricane. James Keller 0001, Mihail Popescu, Dustin Gibeson |
IGARSS | 2 |
| 2012 | Implementing bounded rationality in disaster agent behavior using OGA operatorsabstractMany agent based models for disaster crowd behavior in closed spaces have been described in the literature. In this paper we propose an extension of the confined space models based on the bounded rationality theory (BRT) that is able to model crowd behavior during large area events with longer time constants, such as natural disasters. We model the BRT factor selection process using an ordered geometric average approach (OGA). Then, we present two examples of crowd behavior related to the evacuation of a hypothetical island. Mihail Popescu, James Keller 0001 |
IGARSS | 1 |
| 2012 | Similarity measure for anomaly detection and comparing human behaviorsabstractHerein, we put forth a new similarity measure for anomaly detection and for comparing human behaviors based on the theories of learning automata, comparison of soft partitions, and temporal probabilistic order relations. In particular, focus is placed on monitoring individuals in a home setting for their own well-being. This work is a high-level investigation focused on the structure of human behavior. Examples demonstrate the utility of this approach for (1) understanding the similarity of pairs of behaviors for an individual (or alternatively between individuals) and (2) detecting significant change between changing behavior and a baseline model. In the context of eldercare, significant change in behavior can be a precursor to cognitive and/or functional health related problems. Simulated resident behavior is used to show different scenarios and the response of the proposed measure. © 2012 Wiley Periodicals, Inc. Derek Anderson, María Ros, James Keller 0001, Manuel P. Cuéllar, Mihail Popescu, Miguel Delgado 0001, Maria-Amparo Vila |
Int. J. Intell. Syst. | 5 |
| 2011 | Comparing soft clusters and partitionsabstractPreviously, we presented a method for comparing soft partitions (i.e. crisp, probabilistic, fuzzy and possibilistic) to a known crisp reference partition. Many of the classical indices that have been used with outputs of crisp clustering algorithms were generalized so that they are applicable for candidate partitions of any type. In particular, focus was placed on generalizations of the Rand index. In this article, we extend our prior work by (1) investigating the behavior of the soft Rand for comparing non-crisp, specifically possibilistic, partitions and (2) we demonstrate how the possibilistic Rand and visual assessment of (cluster) tendency (VAT) algorithm can be used to discover the number of actual clusters and coincident clusters for outputs from the possibilistic c-means (PCM) algorithm. Derek Anderson, James Keller 0001, Ozy Sjahputera, James C. Bezdek, Mihail Popescu |
FUZZ-IEEE | 5 |
| 2011 | Improving disease prediction using ICD-9 ontological featuresabstractDisease prediction has become important in a variety of applications such as health insurance, tailored health communication and public health. Disease prediction is usually performed using publically available datasets such as HCUP, NHANES or MDS that were initially designed for health reporting or health cost evaluation but not for disease prediction. In these datasets, medical diagnoses are traditionally arranged in "diagnose-related groups" (DRGs). In this paper we compare the disease prediction based on crisp DRG features with the results obtained employing a new set of features that consist of the fuzzy membership of patient diagnoses in the DRG groups. The fuzzy membership features were computed using an ICD-9 ontological similarity approach. The prediction results obtained on a subset of 9,000 patients from the 2005 HCUP data representing three diseases (diabetes, atherosclerosis and hypertension) using two classifiers (random forest and SVM trained on 21,000 samples) show significant (about 10%) improvement as measured by the area under the ROC curve (AROC). Mihail Popescu, Mohammad Khalilia |
FUZZ-IEEE | 1 |
| 2011 | Linguistic summarization of long-term trends for understanding change in human behaviorabstractIn this paper, we propose a linguistic summarization procedure for describing long-term trends of change in human behavior. Our objective consists of defining methods that provide information to elders, caregivers, social workers or even family in an understandable language. We adapt a measure that we defined in previous work on soft cluster partition similarity for comparing behaviors that are adapted over time. From that measure, we are able to produce a time series that numerically describes change in behavior over time. In this article, the resulting time series is partitioned and linguistically summarized depending on a user's (caregiver, social worker, etc.) desired time resolution. Simulated resident behavior is used in order to explore a range of different scenarios and the response of the proposed linguistic summarization process is investigated. María Ros, Manuel P. Cuéllar, Miguel Delgado 0001, Maria-Amparo Vila, Derek Anderson, James Keller 0001, Mihail Popescu |
FUZZ-IEEE | 7 |
| 2011 | Correlation cluster validityabstractA common question asked about unlabeled data sets is how many subsets (or clusters) of objects are represented in the data? The answer to this question is usually obtained by first clustering the data, and then employing a cluster validity measure to validate one or more candidate partitions of the objects. In this paper we describe an universal cluster validity measure that, unlike most existing measures, can be applied to partitions obtained by any relational or object data clustering algorithm. We illustrate the new measure, and compare it to several well known existing measures using a variety of artificial data sets. Mihail Popescu, James Keller 0001, James C. Bezdek, Timothy C. Havens |
SMC | 1 |
| 2010 | An ontological fuzzy Smith-Waterman with applications to patient retrieval in Electronic Medical RecordsabstractWith the introduction of Electronic Medical Records (EMR) systems in health care institutions, a huge data repository has been created. By employing computational intelligence (CI) techniques, this data repository can be used to address important health care issues such as improving quality and reducing medical errors. In this paper, we introduce a general word sequence alignment method based on a fuzzy version of the Smith-Waterman (SW) dynamic programming algorithm. The word similarity matrix used in computing the sequence alignment is calculated based on a domain ontology (taxonomy). The fuzzy version of the SW algorithm is designed to accommodate words not present in the initial dictionary used to precompute the similarity matrix, hence avoiding its recalculation. We apply the developed algorithm for patient retrieval in an EMR. Each patient is described by an ordered sequence of ICD9 diagnoses. We analyze various properties of the proposed algorithm on a patient dataset that contains 107 patients described by ICD9 diagnose sequences. Mihail Popescu |
FUZZ-IEEE | 1 |
| 2010 | A Comparison of Five Fuzzy Rand Indices
Derek Anderson, James C. Bezdek, James Keller 0001, Mihail Popescu |
IPMU (1) | 4 |
| 2010 | Comparing Fuzzy, Probabilistic, and Possibilistic PartitionsabstractWhen clustering produces more than one candidate to partition a finite set of objectsO, there are two approaches to validation (i.e., selection of a “best” partition, and implicitly, a best value for c , which is the number of clusters inO). First, we may use an internal index, which evaluates each partition separately. Second, we may compare pairs of candidates with each other, or with a reference partition that purports to represent the “true” cluster structure in the objects. This paper generalizes many of the classical indices that have been used with outputs of crisp clustering algorithms so that they are applicable for candidate partitions of any type (i.e., crisp or soft, with soft comprising the fuzzy, probabilistic, and possibilistic cases). Space prevents inclusion of all of the possible generalizations that can be realized this way. Here, we concentrate on the Rand index and its modifications. We compare our fuzzy-Rand index with those of Campello, Hullermeier and Rifqi, and Brouwer, and show that our extension of the Rand index is O(n), while the other three are all O(n2). Numerical examples are given to illustrate various facets of the new indices. In particular, we show that our indices can be used, even when the partitions are probabilistic or possibilistic, and that our method of generalization is valid for any index that depends only on the entries of the classical (i.e., four-pair types) contingency table for this problem. Derek Anderson, James C. Bezdek, Mihail Popescu, James Keller 0001 |
IEEE Trans. Fuzzy Syst. | 3 |
| 2010 | Computing With Words With the Ontological Self-Organizing MapabstractThis paper addresses thecomputing-with-wordsparadigm by presenting an ontological self-organizing map (OSOM), which produces visualization and summarization information about datasets composed of words, namely, ontological data. The specific data that are used in this paper are the Gene Ontology (GO) annotations of genes and gene products. The OSOM is an extension of the SOM, which was initially developed by Kohonen. We adapt the SOM by integrating ontology-based similarity measures and relational-clustering distance measures. We also develop a novel prototype update. We present results on two datasets composed of GO annotations of genes and gene products. An OSOM-based summarization, which produces the term-based summarizations of the trained OSOM network, is also demonstrated. The results show that the OSOM-based visualization method correctly shows the cluster tendency of the genes and gene products and that the summarization provides useful information about the mapped groups of genes and gene products. Timothy C. Havens, James Keller 0001, Mihail Popescu |
IEEE Trans. Fuzzy Syst. | 3 |
| 2009 | eCCV: A new fuzzy cluster validity measure for large relational bioinformatics datasetsabstractThe existence of BLAST sequence comparison algorithm and microarray technology are among the reasons that make bioinformatics the domain with the most abundant large relational datasets. For example, by BLAST-ing the genes of the human genome (around 30,000 genes) we obtain a 30,000 by 30,000 distance matrix. This matrix can not be currently stored in the memory of a typical desktop PC. In the same time, clustering the resulting matrix using a fuzzy relational clustering algorithm such as Non-Euclidean Fuzzy C-means (NERFCM) requires prior knowledge of the number of clusters existent in the data set. The question is, how can we evaluate the number of clusters if we can't even load the matrix in the memory our PC? To address this problem, we propose to extend the correlation cluster validity (CCV) that we introduced in a previous paper, denoting the new validity measure as eCCV. eCCV consists of two steps: first sampling of the large matrix followed by the estimation of the number of cluster employing CCV of the sampled data. The sampling strategy produces also a significant processing speedup. We illustrate eCCV properties on a large synthetic dataset and on a large subset of human genes obtained from the RefSeq database. Mihail Popescu, James C. Bezdek, James Keller 0001 |
FUZZ-IEEE | 1 |
| 2009 | Clustering in ordered dissimilarity dataabstractThis paper presents a new technique for clustering either object or relational data. First, the data are represented as a matrix D of dissimilarity values. D is reordered to D* using a visual assessment of cluster tendency algorithm. If the data contain clusters, they are suggested by visually apparent dark squares arrayed along the main diagonal of an image I(D*) of D*. The suggested clusters in the object set underlying the reordered relational data are found by defining an objective function that recognizes this blocky structure in the reordered data. The objective function is optimized when the boundaries in I(D*) are matched by those in an aligned partition of the objects. The objective function combines measures of contrast and edginess and is optimized by particle swarm optimization. We prove that the set of aligned partitions is exponentially smaller than the set of partitions that needs to be searched if clusters are sought in D. Six numerical examples are given to illustrate various facets of the algorithm. © 2009 Wiley Periodicals, Inc. Timothy C. Havens, James C. Bezdek, James Keller 0001, Mihail Popescu |
Int. J. Intell. Syst. | 4 |
| 2008 | Ontological self-organizing maps for cluster visualization and functional summarization of gene products using Gene Ontology similarity measuresabstractThis paper presents an ontological self-organizing map (OSOM), which is used to produce visualization and functional summarization information about gene products using gene ontology (GO) similarity measures. The OSOM is an extension of the self-organizing map as initially developed by Kohonen, which trains on data composed of sets of terms. Term-based similarity measures are used as a distance metric as well as in the update of the OSOM training procedure. We present an OSOM-based visualization method that shows the cluster tendency of the gene products. Also demonstrated is an OSOM-based functional summarization which produces the most representative term(s) (MRT) from the GO for each OSOM prototype and, subsequently, each gene product cluster. We validated the results of our method by applying the OSOM to a well-studied set of gene products. Timothy C. Havens, James Keller 0001, Mihail Popescu, James C. Bezdek |
FUZZ-IEEE | 3 |
| 2008 | A new cluster validity measure for bioinformatics relational datasetsabstractMany important applications in biology have underlying datasets that are relational, that is, only the (dis)similarity between biological objects (amino acid sequences, gene expression profiles, etc.) is known and not their feature values in some feature space. Examples of such relational datasets are the gene similarity matrices obtained from BLAST, gene expression data, or gene ontology (GO) similarity measures. Once a relational dataset is obtained, a common question asked is how many groups of objects are represented in the original dataset. The answer to this question is usually obtained by employing a clustering algorithm and a cluster validity measure. In this article we describe a cluster validity measure for non-Euclidean relational fuzzy c-means that is based on the correlation between a relation induced on the data by the cluster memberships and the original relational data. This validity measure can be applied to partitions made by any fuzzy relational clustering algorithm. We illustrate our measure by validating clusters in several dissimilarity matrices for a set of 194 gene products obtained using BLAST and GO similarities. Mihail Popescu, James C. Bezdek, James Keller 0001, Timothy C. Havens, Jacalyn M. Huband |
FUZZ-IEEE | 1 |
| 2008 | Dunn's cluster validity index as a contrast measure of VAT imagesabstractThis paper addresses the relationship between the visual assessment of cluster tendency (VAT) algorithm and Dunnpsilas cluster validity index. We present an analytical comparison in conjunction with numerical examples to demonstrate that the effectiveness of VAT in showing cluster tendency is directly related to Dunnpsilas index. This analysis is important to understanding the underlying theory of VAT and VAT-based algorithms and, more generally, other algorithms that are based on, or similar to, Primpsilas Algorithm. Timothy C. Havens, James C. Bezdek, James Keller 0001, Mihail Popescu |
ICPR | 4 |
| 2007 | GoFuzzKegg: Mapping Genes to KEGG Pathways Using an Ontological Fuzzy Rule SystemabstractIn this paper we present a method for finding the main pathways represented in a set of genes (say obtained from a microarray experiment). The method is based on a fuzzy mapping between genes represented as sets of gene ontology terms and KEGG pathways using a new type of fuzzy rule system called ontological fuzzy rule system (OFRS). As opposed to a crisp mapping, the fuzzy mapping produces a nonzero value even if the gene name is not explicitly listed in a given KEGG pathway. An OFRS is a fuzzy rule system in which the rule memberships are obtained using similarity measures between objects computed based on the gene ontology (GO) annotations. To test our approach, we randomly selected without replacement 10 sets of Arabidopsis thaliana genes from KEGG (each set had 15 genes from 3 different pathways) and tried to predict the pathways they were selected from. Our method was able to find, 90% of the right pathways with a 65% false alarm rate at a p-value of 0.01. The high false alarm rate is due in part to the experimental setting. In a pilot dataset of 526 Arabidopsis thaliana genes we identified 8 clusters which proved to be linked to important pathways such as ATP synthesis and transcription factor Mihail Popescu, Dong Xu 0002, Erik Taylor |
CIBCB | 1 |
| 2007 | Mapping Genes to Pathways Using Ontological Fuzzy Rule SystemsabstractIn this paper we present a novel algorithm for mapping genes to pathways. The approach is based on the concept of ontological fuzzy rule system (OFRS) that, we believe, represents a step closer toward Zadeh's "computing with words" paradigm. An OFRS is a fuzzy rule system that uses ontological mapping between objects according to their representation as sets of terms from an ontology. In our mapping approach the left-hand-side contains genes objects represented using the gene ontology (GO), and the right-hand-side consists in pathways described using a pathway ontology (KEGG). The question of mapping a set of genes to pathways often arises in microarray experiments where one would like to know what are the pathways that can explain the observed gene expression patterns. To compare various mapping approaches, we use a pilot dataset extracted from KEGG that consisted of 10 sets of 15 genes taken from 3 pathways (5 gene/pathway). We conclude that the best matching strategy consists in two steps: in the first step the crisp approach is used to find the pathways involved, and in the second step the fuzzy (ontological) approach is employed to map the genes that could not be found in KEGG. Mihail Popescu, Dong Xu 0002 |
FUZZ-IEEE | 1 |
| 2007 | Relational Analysis of CpG Islands Methylation and Gene Expression in Human Lymphomas Using Possibilistic C-Means Clustering and Modified Cluster Fuzzy DensityabstractHeterogeneous genetic and epigenetic alterations are commonly found in human non-Hodgkin's lymphomas (NHL). One such epigenetic alteration is aberrant methylation of gene promoter-related CpG islands, where hypermethylation frequently results in transcriptional inactivation of target genes, while a decrease or loss of promoter methylation (hypomethylation) is frequently associated with transcriptional activation. Discovering genes with these relationships in NHL or other types of cancers could lead to a better understanding of the pathobiology of these diseases. The simultaneous analysis of promoter methylation using Differential Methylation Hybridization (DMH) and its associated gene expression using Expressed CpG Island Sequence Tag (ECIST) microarrays generates a large volume of methylation-expression relational data. To analyze this data, we propose a set of algorithms based on fuzzy sets theory, in particular Possibilistic c-Means (PCM) and cluster fuzzy density. For each gene, these algorithms calculate measures of confidence of various methylation-expression relationships in each NHL subclass. Thus, these tools can be used as a means of high volume data exploration to better guide biological confirmation using independent molecular biology methods. Ozy Sjahputera, James Keller 0001, J. Wade Davis, Kristen H. Taylor, Farahnaz Rahmatpanah, Huidong Shi, Derek Anderson, Samuel Blisard, Robert H. Luke III, Mihail Popescu, Gerald C. Arthur, Charles William Caldwell |
IEEE ACM Trans. Comput. Biol. Bioinform. | 10 |
| 2006 | OntoQuest: A Physician Decision Support System based on Ontological Queries of the Hospital Database
Mihail Popescu, Gerald C. Arthur |
AMIA | 1 |
| 2006 | Summarization of Patient Groups Using the Fuzzy C-Means and Ontology Similarity MeasuresabstractThis paper addresses the problem of constructing a summarization of groups of patients that are found by clustering a hospital database where diagnoses are encoded in a controlled medical vocabulary, called ICD-9. Our method finds the "most representative terms" (MRTs) for a patient cluster by using weights from a fuzzy partition matrix generated by fuzzy clustering the patient similarity matrix. We present a novel approach to computing patient similarity by using OWA operators. Finally, we apply our method to a set of 2077 cardiology patients. Mihail Popescu, James Keller 0001 |
FUZZ-IEEE | 1 |
| 2006 | Bioinformatics and Fuzzy LogicabstractMany biological systems and objects are intrinsically fuzzy. Fuzzy set theory and fuzzy logic are ideal frameworks for describing some biological systems/objects and providing suitable computational methods for a widely range of bioinformatics problems. In this paper, we present two examples of using fuzzy set theory in bioinformatics, one in fuzzy measurement of ontological similarity and its application in bioinformatics, and the other in the application of the fuzzy k-nearest neighbor algorithm in protein secondary structure prediction. We also review other "fuzzy" methods for bioinformatics applications. Dong Xu 0002, Rajkumar Bondugula, Mihail Popescu, James Keller 0001 |
FUZZ-IEEE | 3 |
| 2006 | Gene Ontology Similarity Measures Based on Linear Order StatisticsabstractThe standard method for comparing gene products (proteins or RNA) is to compare their DNA or amino acid sequences. Additional information about some gene products may come from multiple sources, including the set of Gene Ontology (GO) annotations and the set of journal abstracts related to each gene product. Gene product similarity measures can be based on evaluating sets of descriptor terms found in the GO taxonomy, and/or the index term sets of the related documents (MeSH annotations). While our techniques can be applied to term sets from any taxonomy, we restrict our examples in this article to GO annotations. We investigate the use of linear order statistics (LOS) to build similarity relations on pairs of terms that are used in the GO as linguistic descriptors of genes and gene products. One of our objectives is to investigate the construction and utility of visual assessments of relational data (in this case, dissimilarity matrices) for discovering tendencies of groups of gene products to "cluster together". We use gene product data derived from a group of 194 gene products representing three protein families extracted from ENSEMBL. Our examples suggest that LOS similarity measures are more effective than traditional sequence-based similarity measures at capturing relationships between pairs of gene products in ENSEMBL families when annotation information is available. We show examples of how these similarity measures can assist in knowledge discovery and gene product family validation. James Keller 0001, James C. Bezdek, Mihail Popescu, Nikhil R. Pal, Joyce A. Mitchell, Jacalyn M. Huband |
Int. J. Uncertain. Fuzziness Knowl. Based Syst. | 3 |
| 2006 | Fuzzy Measures on the Gene Ontology for Gene Product SimilarityabstractOne of the most important objects in bioinformatics is a gene product (protein or RNA). For many gene products, functional information is summarized in a set of Gene Ontology (GO) annotations. For these genes, it is reasonable to include similarity measures based on the terms found in the GO or other taxonomy. In this paper, we introduce several novel measures for computing the similarity of two gene products annotated with GO terms. The fuzzy measure similarity (FMS) has the advantage that it takes into consideration the context of both complete sets of annotation terms when computing the similarity between two gene products. When the two gene products are not annotated by common taxonomy terms, we propose a method that avoids a zero similarity result. To account for the variations in the annotation reliability, we propose a similarity measure based on the Choquet integral. These similarity measures provide extra tools for the biologist in search of functional information for gene products. The initial testing on a group of 194 sequences representing three proteins families shows a higher correlation of the FMS and Choquet similarities to the BLAST sequence similarities than the traditional similarity measures such as pairwise average or pairwise maximum. Mihail Popescu, James Keller 0001, Joyce A. Mitchell |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2006 | Fuzzy spatial pattern processing using linguistic hidden Markov modelsabstractIn this work, we propose a hidden Markov model (HMM), called the linguistic HMM, suitable for processing sequences of fuzzy vectors. A fuzzy vector B is an n-tuple of fuzzy numbers. Since fuzzy numbers are often associated with linguistic terms, such as "small," "medium," etc., a fuzzy vector can also be called a linguistic vector. Similarly, an HMM that processes linguistic vectors can be called a "linguistic HMM." The derivation of the linguistic HMM (LHMM) from the continuous HMM is performed using the extension principle and the decomposition theorem. We prove that a LHMM behaves in the same fashion as the CHMM in the degenerate linguistic case when the fuzzy numbers are singletons (real numbers). We also provide an example where an LHMM was used for the recognition of a play (pick-and-shoot) during a basketball game. The positions of the players were described using spatial fuzzy relations. For the recognition experiment, we generated two sets of 100 sequences containing pick-and-shoot and non pick-and-shoot sequences, respectively. The LHMM results obtained for the fuzzy sequences were compared to the CHMM results obtained on a crisp version of the same sequences. The results obtained showed that the fuzzy spatial relations together with the LHMM provide a better description of the movement than the CHMM. Mihail Popescu, Paul D. Gader, James Keller 0001 |
IEEE Trans. Fuzzy Syst. | 1 |
| 2005 | Soft Computing in Bioinformatics
James Keller 0001, Mihail Popescu |
FUZZ-IEEE | 2 |
| 2005 | Gene Ontology Automatic Annotation Using a Domain Based Gene Product Similarity MeasureabstractRecent years have seen an explosive growth in the amount of biological data available for analysis. The large volume of data collected makes it necessary to automatically classify and sort such data on a very large scale. Typically, investigators use computational sequence analysis tools to assign functions to newly found gene products. The problem is to find the functions of a (unknown) gene product given its amino acid sequence. In this work we search for functional similarity between gene products by matching the functional domains that they contain. The domain-based approach addresses the main problem of sequence-based similarity, i.e., when the region of a gene product that is matched by a query sequence is not related to the function of that gene product. We use the hidden Markov representation of a gene product domain as described in the PFAM database, and then infer annotations that come from the Gene Ontology. To compute domain similarity between two gene products we introduce a fuzzy Jaccard similarity measure. We tested our domain-based similarity for the functional annotation of a set of 194 gene products extracted from the ENSEMBL Web site. We compared the domain similarity approach to the traditional way of performing functional annotation using a sequence-based similarity (BLAST and Smith-Waterman). The annotation was performed in all cases using a fuzzy K-nearest neighbor algorithm. We found that our domain-based annotation was better than the most common BLAST approach, but not as good as complex Smith-Waterman technique. The domain-based annotation has about 70% correct annotation rate at 17% false annotation rate Mihail Popescu, James Keller 0001, Joyce A. Mitchell |
FUZZ-IEEE | 1 |
| 2004 | Taxonomy-based soft similarity measures in bioinformaticsabstractOne of the most important objects in bioinformatics is a gene product (a protein or an RNA). Besides the gene sequence and expression values found following a microarray experiment, for many gene products, additional functional information comes from the set of gene ontology (GO) annotations and the set of journal abstracts related to the gene product. For these genes, it is reasonable to include similarity measures based on the terms found in the GO and/or the index term sets of the related documents (MeSH annotations). We propose a fuzzy measure-based similarity (FMS) for computing the similarity of two gene products annotated with terms from ontology. The advantage of FMS is that it takes into consideration the context of the whole set when computing the similarity. For the case when the two gene products are not annotated by common ontology terms, we propose a method that avoids a zero similarity result. In dealing with large groups of documents describing the objects under consideration, not only do we determine the similarity between the document pairs, but, by introducing the Choquet integral to the scenario, we can fuse this partial agreement function on pairs of documents into a single value relating the gene products. We present examples of FMS calculation for specific situations where two genes are described by a set of terms from the gene ontology, comparing our measures to others from the literature. James Keller 0001, Mihail Popescu, Joyce A. Mitchell |
FUZZ-IEEE | 2 |
| 2003 | Linguistic hidden Markov modelsabstractIn this paper we develop a hidden Markov model (HMM), called the linguistic HMM (LHMM), suitable for processing sequences of fuzzy vectors. A fuzzy vector B is an n-tuple of fuzzy numbers. Since fuzzy numbers are often associated with linguistic terms, such as "small," "medium," etc., a fuzzy vector can also be called a linguistic vector. The derivation of the linguistic HMM (LHMM) from the numeric HMM is done using the extension principle and the decomposition theorem. We show that the LHMM behaves in the same way as the HMM in the degenerate linguistic case when the fuzzy numbers are singletons (real numbers). We also derive the related algorithms for LHMM training (linguistic Baum-Welch) and for LHMM recognition (linguistic Viterbi). Several examples of LHMM training and recognition are given. Mihail Popescu, James Keller 0001, Paul D. Gader |
FUZZ-IEEE | 1 |
| 1999 | An Agent-Based Approach for Interpreting Medical ImagesabstractWe present an agent-based approach for medical image interpretation. The system is based on the concept of active fusion for image recognition and is composed of two major types of intelligent agents: radiologist agents and patient representative agents. A patient representative agent takes images from the patient through a Web-based interface, asks for multiple opinions from radiologist agents in interpreting them, and then integrates the opinions for the user. A radiologist agent decomposes the image interpretation task into smaller subtasks, uses multiple agents to solve the subtasks, and combines the solutions to the subtasks intelligently to solve the image interpretation problem. Mihail Popescu, Yi Shang |
ICTAI | 1 |
| 1998 | Using Co-Occurrence Data to Determine a Thesaurus Structure
James E. Andrews, Timothy B. Patrick, David E. Moxley, Colleen M. Meyer, Mihail Popescu, MaryEllen C. Sievert |
AMIA | 5 |