EDBT 2026 Demo / reviewers in the wild / expert
Chi-Ren Shyu
dblp:44/6873
· DBLP profile ↗
89ranked-venue papers
7as first author
8since 2021 · last 2025
0000-0001-9197-9522ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 77 · 4 first-author · 7 since 2021Artificial intelligence and machine learning · 6 · 3 first-authorGraphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-authorDatabases, data management, data science and information retrieval · 3 · 1 since 2021Human-computer interaction and ubiquitous computing · 3 · 2 since 2021Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Enhancing Student Self-Efficacy and Interest in Microelectronics Through Immersive Virtual Reality in an Informal Learning EnvironmentabstractThis study investigates iVRLab, an immersive virtual reality (VR) microfabrication training system, and its influence on college students' self-efficacy and interest in microelectronics during a one-hour workshop held within a 1.5-day training camp. iVRLab simulates photolithography cleanroom operations, providing hands-on virtual experiences that merge theory with practice. A pre-post analysis revealed a significant rise in self-efficacy (M = 4.11 to 4.37, p = 0.0066), demonstrating the system's effectiveness in building confidence. Although interest only showed slight gains, high baseline scores (3.6-3.7 on a 4-point scale) indicate a ceiling effect. Correlations among engagement, embodiment (r = 0.91, p < 0.001), and immersion (r = 0.60, p < 0.01) underscore the value of active, embodied VR learning. Qualitative feedback further highlighted iVRLab's immersive realism and its capacity to spark curiosity about microfabrication careers. These findings suggest that VR training can effectively boost self-efficacy in microelectronics, reinforcing students' enthusiasm in an informal learning setting and indicating the potential broader adoption across STEM fields. Yupei Duan, Fang Wang 0024, Chi-Ren Shyu, Syed K. Islam, Sazia A. Eliza, Jim Flink, Hao He 0009, Shangman Li, Scottie Murrell, Amith Nalmas |
ICALT | 4 |
| 2025 | Students' Motivation in STEM Education through a Microelectronics Training Program: Through the Lens of the ARCS Model of Motivational DesignabstractIn this study, we examined the effects of a 1.5-day microelectronics training program on students' self-efficacy. Using a quasi-experimental design (N = 27), we analyzed the data through a paired t-test complemented by qualitative insights. The findings revealed that: (1) a short-term microelectronics training program, designed based on the Attention, Relevance, Confidence, and Satisfaction (ARCS) motivational model, can positively influence learners' self-efficacy toward STEM education; and (2) incorporating diverse instructional strategies-such as guest talks from industry and academic experts, along with hands-on, real-world experiences-is essential for making abstract textbook concepts more concrete and inspiring students to pursue STEM-related careers. Chi-Ren Shyu, Syed K. Islam, Sazia A. Eliza, Jim Flink |
ICALT | 3 |
| 2025 | Identifying homogenous patient subgroups using transformer based hierarchical clustering of heterogeneous Mixed-Modality medical data
William Baskett, Benjamin Black, Adnan I. Qureshi, Chi-Ren Shyu |
J. Biomed. Informatics | 4 |
| 2024 | Identifying gene expression programs in single-cell RNA-seq data using linear correlation explanation
Yulia I. Nussbaum, K. S. M. Tozammel Hossain, Jussuf T. Kaifi, Wesley C. Warren, Chi-Ren Shyu, Jonathan B. Mitchem |
J. Biomed. Informatics | 5 |
| 2023 | Understanding common key indicators of successful and unsuccessful cancer drug trials using a contrast mining framework on ClinicalTrials.govabstractClinical trials are essential to the process of new drug development. As clinical trials involve significant investments of time and money, it is crucial for trial designers to carefully investigate trial settings prior to designing a trial. Utilizing trial documents from ClinicalTrials.gov, we aim to understand the common characteristics of successful and unsuccessful cancer drug trials to provide insights about what to learn and what to avoid. In this research, we first computationally classified cancer drug trials into successful and unsuccessful cases and then utilized natural language processing to extract eligibility criteria information from the trial documents. To provide explainable and potentially modifiable recommendations for new trial design, contrast mining was applied to discover highly contrasted patterns with a significant difference in prevalence between successful (completion with advancement to the next phase) and unsuccessful (suspended, withdrawn, or terminated) groups. Our method identified contrast patterns consisting of combinations of drug categories, eligibility criteria, study organization, and study design for nine major cancers. In addition to a literature review for the qualitative validation of mined contrast patterns, we found that contrast-pattern-based classifiers using the top 200 contrast patterns as feature representations can achieve approximately 80% F1 score for eight out of ten cancer types in our experiments. In summary, aligning with the modernization efforts of ClinicalTrials.gov, our study demonstrates that understanding the contrast characteristics of successful and unsuccessful cancer trials may provide insights into the decision-making process for trial investigators and therefore facilitate improved cancer drug trial design. Shu-Kai Chang, Danlu Liu, Jonathan B. Mitchem, Christos Papageorgiou, Jussuf T. Kaifi, Chi-Ren Shyu |
J. Biomed. Informatics | 6 |
| 2022 | Found in Translation: Transforming Rules-Based Diabetes Phenotyping Algorithms into Reproducible Diabetes Cohorts from Real-World Data
Erin M. Tallon, Mark A. Clements, Chi-Ren Shyu |
AMIA | 3 |
| 2022 | RHPTree - Risk Hierarchical Pattern Tree for Scalable Long Pattern MiningabstractRisk patterns are crucial in biomedical research and have served as an important factor in precision health and disease prevention. Despite recent development in parallel and high-performance computing, existing risk pattern mining methods still struggle with problems caused by large-scale datasets, such as redundant candidate generation, inability to discover long significant patterns, and prolonged post pattern filtering. In this article, we propose a novel dynamic tree structure, Risk Hierarchical Pattern Tree (RHPTree), and a top-down search method, RHPSearch, which are capable of efficiently analyzing a large volume of data and overcoming the limitations of previous works. The dynamic nature of the RHPTree avoids costly tree reconstruction for the iterative search process and dataset updates. We also introduce two specialized search methods, the extended target search (RHPSearch-TS) and the parallel search approach (RHPSearch-SD), to further speed up the retrieval of certain items of interest. Experiments on both UCI machine learning datasets and sampled datasets of the Simons Foundation Autism Research Initiative (SFARI)—Simon’s Simplex Collection (SSC) datasets demonstrate that our method is not only faster but also more effective in identifying comprehensive long risk patterns than existing works. Moreover, the proposed new tree structure is generic and applicable to other pattern mining problems. Danlu Liu, Yu Li 0052, William Baskett, Dan Lin 0001, Chi-Ren Shyu |
ACM Trans. Knowl. Discov. Data | 5 |
| 2021 | Explainable artificial intelligence in high-throughput drug repositioning for subgroup stratifications with interventionable potentialabstractEnabling precision medicine requires developing robust patient stratification methods as well as drugs tailored to homogeneous subgroups of patients from a heterogeneous population. Developing de novo drugs is expensive and time consuming with an ultimately low FDA approval rate. These limitations make developing new drugs for a small portion of a disease population unfeasible. Therefore, drug repositioning is an essential alternative for developing new drugs for a disease subpopulation. This shows the importance of developing data-driven approaches that find druggable homogeneous subgroups within the disease population and reposition the drugs for these subgroups. In this study, we developed an explainable AI approach for patient stratification and drug repositioning. Contrast pattern mining and network analysis were used to discover homogeneous subgroups within a disease population. For each subgroup, a biomedical network analysis was done to find the drugs that are most relevant to a given subgroup of patients. The set of candidate drugs for each subgroup was ranked using an aggregated drug score assigned to each drug. The proposed method represents a human-in-the-loop framework, where medical experts use the data-driven results to generate hypotheses and obtain insights into potential therapeutic candidates for patients who belong to a subgroup. Colorectal cancer (CRC) was used as a case study. Patients' phenotypic and genotypic data was utilized with a heterogeneous knowledge base because it gives a multi-view perspective for finding new indications for drugs outside of their original use. Our analysis of the top candidate drugs for the subgroups identified by medical experts showed that most of these drugs are cancer-related, and most of them have the potential to be a CRC regimen based on studies in the literature. Zainab Al-Taie, Danlu Liu, Jonathan B. Mitchem, Christos Papageorgiou, Jussuf T. Kaifi, Wesley C. Warren, Chi-Ren Shyu |
J. Biomed. Informatics | 7 |
| 2020 | Drug Repositioning for Colorectal Cancer Patient Subgroups Using Exploratory Mining and Network Analysis
Zainab Al-Taie, Danlu Liu, Christos Papageorgiou, Jussuf T. Kaifi, Jonathan B. Mitchem, Chi-Ren Shyu |
AMIA | 6 |
| 2020 | Differentiating Genetic Risk Factors Between Triple Negative and Luminal Breast Cancer Subtypes Using Long Contrast Patterns Mining
Danlu Liu, Zainab Al-Taie, Christos Papageorgiou, Chi-Ren Shyu |
AMIA | 5 |
| 2020 | Using novel exploratory subgroup mining to determine complication patterns of post colectomy based on hospital characteristics
Danlu Liu, Chi-Ren Shyu, Jonathan B. Mitchem |
AMIA | 3 |
| 2020 | A Deep Exploratory Mining Approach for Discovering Associations between Family History of Type 1 Diabetes and Autoimmune and Related Conditions
Erin M. Tallon, Mark A. Clements, Danlu Liu, Katrina Boles, Chi-Ren Shyu |
AMIA | 5 |
| 2020 | Development of A Blockchain Framework for Virtual Clinical Trials
Yan Zhuang 0011, Lincoln Sheets, Xiyuan Gao, Zon-Yin Shae, Jeffrey J. P. Tsai, Chi-Ren Shyu |
AMIA | 7 |
| 2020 | Exploratory Data Mining for Subgroup Cohort Discoveries and PrioritizationabstractFinding small homogeneous subgroup cohorts in large heterogeneous populations is a critical process for hypothesis development in biomedical research. Concurrent computational approaches are still lacking in robust answers to the question "what hypotheses are likely to be novel and to produce clinically relevant results with well thought-out study designs?" We have developed a novel subgroup discovery method which employs a deep exploratory mining process to slice and dice thousands of potential subpopulations and prioritize potential cohorts based on their explainable contrast patterns and which may provide interventionable insights. We conducted computational experiments on both synthesized data and a clinical autism data set to assess performance quantitatively for coverage of pre-defined cohorts and qualitatively for novel knowledge discovery, respectively. We also conducted a scaling analysis using a distributed computing environment to suggest computational resource needs for when the subpopulation number increases. This work will provide a robust data-driven framework to automatically tailor potential interventions for precision health. Danlu Liu, William Baskett, David Q. Beversdorf, Chi-Ren Shyu |
IEEE J. Biomed. Health Informatics | 4 |
| 2020 | A Patient-Centric Health Information Exchange Framework Using Blockchain TechnologyabstractHealth Information Exchange (HIE) exhibits remarkable benefits for patient care such as improving healthcare quality and expediting coordinated care. The Office of the National Coordinator (ONC) for Health Information Technology is seeking patient-centric HIE designs that shift data ownership from providers to patients. There are multiple barriers to patient-centric HIE in the current system, such as security and privacy concerns, data inconsistency, timely access to the right records across multiple healthcare facilities. After investigating the current workflow of HIE, this paper provides a feasible solution to these challenges by utilizing the unique features of blockchain, a distributed ledger technology which is considered "unhackable". Utilizing the smart contract feature, which is a programmable self-executing protocol running on a blockchain, we developed a blockchain model to protect data security and patients' privacy, ensure data provenance, and provide patients full control of their health records. By personalizing data segmentation and an "allowed list" for clinicians to access their data, this design achieves patient-centric HIE. We conducted a large-scale simulation of this patient-centric HIE process and quantitatively evaluated the model's feasibility, stability, security, and robustness. Yan Zhuang 0011, Lincoln Sheets, Yin-Wu Chen, Zon-Yin Shae, Jeffrey J. P. Tsai, Chi-Ren Shyu |
IEEE J. Biomed. Health Informatics | 6 |
| 2019 | Applying Blockchain Technology to Enhance Clinical Trial Recruitment
Yan Zhuang 0011, Lincoln Sheets, Zon-Yin Shae, Yin-Wu Chen, Jeffrey J. P. Tsai, Chi-Ren Shyu |
AMIA | 6 |
| 2019 | Quantifying and understanding the differences in visual activities with contrast subsequencesabstractUnderstanding differences and similarities between scanpaths has been one of the primary goals for eye tracking research. Sequences of areas of interest mapped from fixations are a major focus for many analytic techniques since these sequences directly relate to the semantic meaning of the visual input. Many studies analyze complete sequences while overlooking the micro-transitions in subsequences. In this paper, we propose a method which extracts subsequences as features and finds contrasting patterns between different viewer groups. The contrast patterns help domain experts to quantify variations between visual activities and understand reasoning processes for complex visual tasks. Experiments were conducted with 39 expert and novice radiographers using nine radiology images corresponding to nine levels of task complexity. Identified contrast patterns, validated by an expert, prove that the method effectively reveals visual reasoning processes that are otherwise hidden. Yu Li 0052, Carla Allen, Chi-Ren Shyu |
ETRA | 3 |
| 2019 | Advancing health equity and access using telemedicine: a geospatial assessmentabstractINTRODUCTION: Health disparity affects both urban and rural residents, with evidence showing that rural residents have significantly lower health status than urban residents. Health equity is the commitment to reducing disparities in health and in its determinants, including social determinants. OBJECTIVE: This article evaluates the reach and context of a virtual urgent care (VUC) program on health equity and accessibility with a focus on the rural underserved population. MATERIALS AND METHODS: We studied a total of 5343 patient activation records and 2195 unique encounters collected from a VUC during the first 4 quarters of operation. Zip codes served as the analysis unit and geospatial analysis and informatics quantified the results. RESULTS: The reach and context were assessed using a mean accumulated score based on 11 health equity and accessibility determinants calculated for each zip code. Results were compared among VUC users, North Carolina (NC), rural NC, and urban NC averages. CONCLUSIONS: The study concluded that patients facing inequities from rural areas were enabled better healthcare access by utilizing the VUC. Through geospatial analysis, recommendations are outlined to help improve healthcare access to rural underserved populations. Saif S. Khairat, Timothy L. Haithcoat, Songzi Liu, Tanzila Zaman, Barbara Edson, Robert Gianforcaro, Chi-Ren Shyu |
J. Am. Medical Informatics Assoc. | 7 |
| 2018 | Leveraging Location to Integrate and Access Information for Analyzing Health Context
Timothy L. Haithcoat, Chi-Ren Shyu |
AMIA | 2 |
| 2018 | Applying Blockchain Technology for Health Information Exchange and Persistent Monitoring for Clinical Trials
Yan Zhuang 0011, Lincoln Sheets, Zon-Yin Shae, Jeffrey J. P. Tsai, Chi-Ren Shyu |
AMIA | 5 |
| 2018 | Heritable genotype contrast mining reveals novel gene associations specific to autism subgroupsabstractThough the genetic etiology of autism is complex, our understanding can be improved by identifying genes and gene-gene interactions that contribute to the development of specific autism subtypes. Identifying such gene groupings will allow individuals to be diagnosed and treated according to their precise characteristics. To this end, we developed a method to associate gene combinations with groups with shared autism traits, targeting genetic elements that distinguish patient populations with opposing phenotypes. Our computational method prioritizes genetic variants for genome-wide association, then utilizes Frequent Pattern Mining to highlight potential interactions between variants. We introduce a novel genotype assessment metric, the Unique Inherited Combination support, which accounts for inheritance patterns observed in the nuclear family while estimating the impact of genetic variation on phenotype manifestation at the individual level. High-contrast variant combinations are tested for significant subgroup associations. We apply this method by contrasting autism subgroups defined by severe or mild manifestations of a phenotype. Significant associations connected 286 genes to the subgroups, including 193 novel autism candidates. 71 pairs of genes have joint associations with subgroups, presenting opportunities to investigate interacting functions. This study analyzed 12 autism subgroups, but our informatics method can explore other meaningful divisions of autism patients, and can further be applied to reveal precise genetic associations within other phenotypically heterogeneous disorders, such as Alzheimer's disease. Matthew Spencer, T. Nicole Takahashi, Sounak Chakraborty, Judith H. Miles, Chi-Ren Shyu |
J. Biomed. Informatics | 5 |
| 2017 | Efficient GPU-accelerated extraction of imperfect inverted repeats from DNA sequencesabstractInverted Repeats in DNA sequences have long been known to have both major beneficial and detrimental effects in regards to how DNA is transcribed and duplicated. Palindromic sequences are frequently translated into proteins and may also facilitate DNA repair in some instances. However, they are also associated with significantly increased risk of mutation. Current methods are either slow or limited in the ways they can process imperfections due to tradeoffs between computational complexity and completeness in results. Our method allows for the efficient extraction of imperfect inverted repeats, featuring the ability to define the level of imperfection by the proportion of mismatching bases. By using GPU acceleration, we achieve order of magnitude speedups compared to the current leading method for imperfect inverted repeat extraction while allowing for more flexible results. We conducted a study on protein-coding exons contained entirely within inverted repeats. We found that these exons were significantly more likely to be included in multiple gene transcripts and were less likely to be spliced out. William Baskett, Matthew Spencer, Chi-Ren Shyu |
BIBM | 3 |
| 2017 | Geospatial health context tableabstractThis project develops a Big Data table that allows researchers to query across and among multiple data sources integrated by location. The big table created in this way uses location as the fundamental linkage between data sets. This is the power of geospatial analysis and forms the foundation for the development and interaction with the Health Context Table. The approach utilizes a dense point file populated with attribution derived or obtained directly from public data sources and associated geospatial analysis. The database created extends across the entire continental United States comprising over 300 million points. The data table has at its core, functional socio-demographic data that is pre-processed, cleaned, integrated and represented in its spatial context. To this core, is being added environmental, infrastructure, cultural, physical, as well as geo-analytically derived layers (i.e. remoteness, isolation). These data span multiple spatial scales (Census Block Group, Zip Code Tabulation Areas, County, etc.). The interface to this Big Data table will allow a user to visualize, data mine, analyze uncertainty, and perform data analytics on these data. The Geospatial Health Context Table's goal is to address the gap in health research and application for an underpinned spatial framework to address real-world issues and research in the context of place. Timothy L. Haithcoat, Chi-Ren Shyu |
BIBM | 2 |
| 2017 | Discovering multifactorial associations with the development of age-related cataract using contrast miningabstractCataract is a cloudiness of eye lens and studies have reported many risk factors for the development of cataract. However, the cumulative effect of multiple factors along with clinical and systemic disease conditions have not been adequately tested due to a limitation in methodology. The collection of a large volume of Electronic Health Records (EHR) offers an opportunity to apply computational tools for knowledge discovery in databases (KDD) process which enable to discover and extract hidden patterns and relationships among a large number of variables. This approach is possible because of the computational friendly EHR database such as the Cerner Health Facts Database. The main goal of this paper is to investigate the factors which are associated with the development of age-related cataracts using EHR data. This study demonstrates the potential of applying data mining tools for risk assessments using large-scale EHR data. Murugesan Raju, Danlu Liu, Frederick W. Fraunfelder, Chi-Ren Shyu |
BIBM | 4 |
| 2017 | Quasi-palindrome effects on DNA sequence evolutionabstractQuasi-palindromes can be harmful or helpful, but most of this functionality is attributed to the formation of cruciforms. Unfortunately, the general properties a sequence must have to facilitate cruciform formation are poorly understood, as most research has focused on case studies of the unusual secondary structures of a few specific sequences. In this work, we investigate the general sequence properties leading to the formation of cruciforms by exploiting the fact that cruciforms lead to an increased mutation rate. We use evolutionary divergences in protein-coding sequences to calculate the mutation rates in quasi-palindromes with different degrees of self-complementarity. We perform this calculation to test different quasi-palindrome properties, then compare the observed patterns with the expected results. The most promising trials indicate that quasi-palindromes must be 10-30bp long with a self-complementarity at least 94% to form cruciforms, and that nucleotides near the center of the sequence are protected from mutation. Matthew Spencer, Jacob Gotberg, Chi-Ren Shyu |
BIBM | 3 |
| 2017 | In-Memory Distributed Indexing for Large-Scale Media Data RetrievalabstractData retrieval serves a critical role in the development of multimedia applications. However, due to the exponential growth of multimedia data, high-speed and efficient indexing is becoming more and more difficult than ever. In this paper, we propose a novel approach to speed up the retrieval process by adopting a distributed computing paradigm through the Apache Spark framework. Utilizing search trees in a Big Data ecosystem leads to fast and cost-effective media database retrievals by caching indexing structures into memory and aggregating ranked results with flexibilities for users to specify the importance of search cues. We conducted computational experiments on large-scaled vector files for remote sensing image database and synthesized pollen image database to demonstrate the effectiveness and scalability of our system with reasonably high accuracy. Yinmiao Ma, Danlu Liu, Grant J. Scott, Jeffrey Uhlmann, Chi-Ren Shyu |
ISM | 5 |
| 2016 | RDF-Based Method to Uncover Implicit Health Communication Episodes from Unstructured Healthcare Data
Pericles S. Giannaris, Zainab Al-Taie, Nattaphon Thanintorn, Ilker Ersoy, Chi-Ren Shyu, Richard D. Hammer, Dmitriy Shin |
AMIA | 5 |
| 2016 | Automatic Workflow Extraction Using Sequential Pattern Mining and Electronic Medical Record Usage Data
Tim A. Green, Michael Phinney, Chi-Ren Shyu |
AMIA | 3 |
| 2016 | Contrasting Autism Subgroup Genotypes Using Frequent Pattern Mining
Matthew Spencer, Chi-Ren Shyu |
AMIA | 2 |
| 2016 | Large scale extraction of perfect and imperfect DNA palindromes using in-memory computingabstractDNA palindromes are known to have many beneficial and detrimental functions in cell biology. Most of these functions are shared between perfect and imperfect palindromic sequences, but imperfect palindromes can be of particular interest due to their additional ability to dramatically increase the local mutation rate. Almost all available tools capable of extracting genetic palindromes were designed to accommodate inverted repeats, which are palindromic sequences with centralized non-matching nucleotides. We developed computational methods that focus specifically on perfect and imperfect palindromes, allowing us to take advantage of palindrome-specific properties which led to an 8× to 40× increase in speed and an approximately 10× decrease in memory requirements compared to the leading inverted repeats detection method. William Baskett, Matthew Spencer, Chi-Ren Shyu |
BIBM | 3 |
| 2016 | Semi-hypothesis guided exploratory analysis for biomedical applicationsabstractMedical research and clinical trials are often based on hypotheses that were observed from clinical practice with noticeable evidence. Forming clinically significant hypotheses will greatly benefit the success of clinical research and ensure both external and internal validity of the trial. In this talk, I will introduce a knowledge discovery approach to automatically identify populations of subjects with commonly occurred comorbidities, genotypes, and phenotypes that present statistically high contract between populations. To focus on a confined set of medical problems as most of medical researchers would like to target (hypertension and diabetes versus all chronic diseases), this approach is able to take a set of selected attributes of interest and expand knowledge discoveries from the initial set. The computational approach consists of a forward floating search method for population selection, a hierarchical frequent pattern mining tree to efficiently handle dense associations, contrast mining for identifying actionable plans, and accumulated contrast (ac-)index for ranking mining results for biomedical researchers. I will present exploratory analysis process and results from the Simon's Simplex Collection (SSC) by the Simons Foundation Autism Research Initiative (SFARI) which comprises data representing 11,560 individuals from 2,591 families. Putative autism subtypes were explored by partitioning families based on demographics and autism phenotypes. An extended contrast mining procedure identified genetic combinations showing preferential association for one of the contrasted subgroups, emphasizing combinations novel to the autistic proband within each family tree. Potentials for other biomedical applications will also be discussed. Chi-Ren Shyu |
BIBM | 1 |
| 2016 | Detection of Simulated Vocal Dysfunctions Using Complex sEMG PatternsabstractSymptoms of voice disorder may range from slight hoarseness to complete loss of voice; from modest vocal effort to uncomfortable neck pain. But even minor symptoms may still impact personal and especially professional lives. While early detection and diagnosis can ameliorate that effect, to date, we are still largely missing reliable and valid data to help us better screen for voice disorders. In our previous study, we started to address this gap in research by introducing an ambulatory voice monitoring system using surface electromyography (sEMG) and a robust algorithm (HiGUSSS) for pattern recognition of vocal gestures. Here, we expand on that work by further analyzing a larger set of simulated vocal dysfunctions. Our goal is to demonstrate that such a system has the potential to recognize and detect real vocal dysfunctions from multiple individuals with high accuracy under both intra and intersubject conditions. The proposed system relies on four sEMG channels to simultaneously process various patterns of sEMG activation in the search for maladaptive laryngeal activity that may lead to voice disorders. In the results presented here, our pattern recognition algorithm detected from two to ten different classes of sEMG patterns of muscle activation with an accuracy as high as 99%, depending on the subject and the testing conditions. Nicholas R. Smith, Luis A. Rivera, Maria Dietrich, Chi-Ren Shyu, Matthew P. Page, Guilherme N. DeSouza |
IEEE J. Biomed. Health Informatics | 4 |
| 2015 | Integrated Clinical Decision Support Systems: Systematic Review and Classification of Online Medical Calculators
Tim A. Green, Chi-Ren Shyu |
AMIA | 2 |
| 2015 | Data Mining to Predict Healthcare Utilization in Managed Care Patients
Lincoln Sheets, Michael Phinney, Sean Lander, Jerry C. Parker, Chi-Ren Shyu |
AMIA | 5 |
| 2015 | Large scale multi-species palindrome study using distributed in-memory computingabstractPalindromic DNA has many interesting and functional properties, including the ability to form non-canonical DNA structures such as hairpins, cruciforms, and slipped strand structures. Palindromes also serve important roles in binding sites and enzyme activity, and have a strong effect on mutation rates. Palindromes are abundant in most genomes, often occurring within coding sequences, though in many instances it is still not clear how their presence affects genomic functions. The identification and study of palindromic DNA is essential to the progression of our understanding of the genome. To address this need, we present a novel method using an in-memory computing environment for identifying, extracting, and indexing palindromes in a searchable database for all mammals in Ensembl release 80. We discuss the preliminary results of a multi-species study on palindromic DNA, focusing on the size, frequency, and distribution of palindromes. Utilizing a Big Data ecosystem enables us to generate the largest palindrome database to date, comprising 42 genomes. Our study offers new insight into the dynamics of palindromes and facilitates future investigation. Devin Petersohn, Matthew Spencer, Alex Fratila, Chi-Ren Shyu |
BIBM | 4 |
| 2015 | An Integrated Approach to Sequence-Independent Local Alignment of Protein Binding SitesabstractAccurate alignment of protein-protein binding sites can aid in protein docking studies and constructing templates for predicting structure of protein complexes, along with in-depth understanding of evolutionary and functional relationships. However, over the past three decades, structural alignment algorithms have focused predominantly on global alignments with little effort on the alignment of local interfaces. In this paper, we introduce the PBSalign (Protein-protein Binding Site alignment) method, which integrates techniques in graph theory, 3D localized shape analysis, geometric scoring, and utilization of physicochemical and geometrical properties. Computational results demonstrate that PBSalign is capable of identifying similar homologous and analogous binding sites accurately and performing alignments with better geometric match measures than existing protein-protein interface comparison tools. The proportion of better alignment quality generated by PBSalign is 46, 56, and 70 percent more than iAlign as judged by the average match index (MI), similarity index (SI), and structural alignment score (SAS), respectively. PBSalign provides the life science community an efficient and accurate solution to binding-site alignment while striking the balance between topological details and computational complexity. Bin Pang 0001, David Schlessman, Xingyan Kuang, Nan Zhao 0002, Daniel Shyu, Dmitry Korkin, Chi-Ren Shyu |
IEEE ACM Trans. Comput. Biol. Bioinform. | 7 |
| 2014 | MRSMRS: Mining repetitive sequences in a MapReduce settingabstractRecent research suggests DNA repeats play critical roles in cellular regulatory functions and disease development. Also, repeat variability among different species, or the same species, is an important indicator for the development of specific phenotypes. Similarities in repetitive sequences among different species have been shown to indicate deeply conserved functions. Patterns such as ultra conserved elements (UCEs), tandem repeats, and palindromes have been of interest. Researchers utilize various computational approaches to aid in the identification of each of these types of patterns. The challenge associated with identifying repeats across a collection of genomes arises from the amount of data stored within DNA. The human genome alone consists of more than 3.1 billion base pairs, and intermediate data generated by alignment- and hash-based approaches are substantial. This sort of all-against-all analysis on a large collection of genomic sequence data often requires data to be reprocessed when new genomes are collected. To handle data of this scale, we utilize the Hadoop Distributed File System running on a cluster of 11 relatively inexpensive nodes, each containing a quad-core commodity processor. Furthermore, to alleviate redundant computation, intermediate data are organized in HBase, allowing us to incrementally process new genomic data without having to reprocess existing genomes. Our approach lends a cost-effective, flexible, robust, and scalable solution to the challenge of identifying various types of repetitive sequences across a collection of genomes. In this study, we benchmark our method using a collection of 6 genomes, summing to an approximate total of 14.2 billion base pairs. Three case studies are presented, demonstrating a 10.4 times speedup over previous state-of-the-art approaches and linear scalability. Hongfei Cao, Michael Phinney, Devin Petersohn, Benjamin Merideth, Chi-Ren Shyu |
BIBM | 5 |
| 2014 | Non-invasive ambulatory monitoring of complex sEMG patterns and its potential application in the detection of vocal dysfunctionsabstractVoice disorders are non-trivial when it comes to their early detection. Symptoms range from slight hoarseness to complete loss of voice, and may seriously impact personal and professional life. To date, we are still largely missing reliable data to help us better understand and screen voice pathologies. In this paper, we present an ambulatory voice monitoring system using surface electromyography (sEMG) and a robust algorithm for pattern recognition of vocal gestures. The system, which can process up to four sEMG channels simultaneously, also can store large amounts of data (up to 13 hours of continuous use) and in the future will be used to analyze on-the-fly various patterns of sEMG activation in the search for maladaptive laryngeal activity that may lead to voice disorders. In the preliminary results presented here, our pattern recognition algorithm (Hierarchical GUSSS) detected six different sEMG patterns of activation, and it achieved 90% accuracy. Nicholas R. Smith, Teekayu Klongtruagrok, Guilherme N. DeSouza, Chi-Ren Shyu, Maria Dietrich, Matthew P. Page |
Healthcom | 4 |
| 2014 | Uncovering influence links in molecular knowledge networks to streamline personalized medicineabstractOBJECTIVES: We developed Resource Description Framework (RDF)-induced InfluGrams (RIIG) - an informatics formalism to uncover complex relationships among biomarker proteins and biological pathways using the biomedical knowledge bases. We demonstrate an application of RIIG in morphoproteomics, a theranostic technique aimed at comprehensive analysis of protein circuitries to design effective therapeutic strategies in personalized medicine setting. METHODS: RIIG uses an RDF "mashup" knowledge base that integrates publicly available pathway and protein data with ontologies. To mine for RDF-induced Influence Links, RIIG introduces notions of RDF relevancy and RDF collider, which mimic conditional independence and "explaining away" mechanism in probabilistic systems. Using these notions and constraint-based structure learning algorithms, the formalism generates the morphoproteomic diagrams, which we call InfluGrams, for further analysis by experts. RESULTS: RIIG was able to recover up to 90% of predefined influence links in a simulated environment using synthetic data and outperformed a naïve Monte Carlo sampling of random links. In clinical cases of Acute Lymphoblastic Leukemia (ALL) and Mesenchymal Chondrosarcoma, a significant level of concordance between the RIIG-generated and expert-built morphoproteomic diagrams was observed. In a clinical case of Squamous Cell Carcinoma, RIIG allowed selection of alternative therapeutic targets, the validity of which was supported by a systematic literature review. We have also illustrated an ability of RIIG to discover novel influence links in the general case of the ALL. CONCLUSIONS: Applications of the RIIG formalism demonstrated its potential to uncover patient-specific complex relationships among biological entities to find effective drug targets in a personalized medicine setting. We conclude that RIIG provides an effective means not only to streamline morphoproteomic studies, but also to bridge curated biomedical knowledge and causal reasoning with the clinical data in general. Dmitriy Shin, Gerald C. Arthur, Mihail Popescu, Dmitry Korkin, Chi-Ren Shyu |
J. Biomed. Informatics | 5 |
| 2014 | Determining Effects of Non-synonymous SNPs on Protein-Protein Interactions using Supervised and Semi-supervised LearningabstractSingle nucleotide polymorphisms (SNPs) are among the most common types of genetic variation in complex genetic disorders. A growing number of studies link the functional role of SNPs with the networks and pathways mediated by the disease-associated genes. For example, many non-synonymous missense SNPs (nsSNPs) have been found near or inside the protein-protein interaction (PPI) interfaces. Determining whether such nsSNP will disrupt or preserve a PPI is a challenging task to address, both experimentally and computationally. Here, we present this task as three related classification problems, and develop a new computational method, called the SNP-IN tool (non-synonymous SNP INteraction effect predictor). Our method predicts the effects of nsSNPs on PPIs, given the interaction's structure. It leverages supervised and semi-supervised feature-based classifiers, including our new Random Forest self-learning protocol. The classifiers are trained based on a dataset of comprehensive mutagenesis studies for 151 PPI complexes, with experimentally determined binding affinities of the mutant and wild-type interactions. Three classification problems were considered: (1) a 2-class problem (strengthening/weakening PPI mutations), (2) another 2-class problem (mutations that disrupt/preserve a PPI), and (3) a 3-class classification (detrimental/neutral/beneficial mutation effects). In total, 11 different supervised and semi-supervised classifiers were trained and assessed resulting in a promising performance, with the weighted f-measure ranging from 0.87 for Problem 1 to 0.70 for the most challenging Problem 3. By integrating prediction results of the 2-class classifiers into the 3-class classifier, we further improved its performance for Problem 3. To demonstrate the utility of SNP-IN tool, it was applied to study the nsSNP-induced rewiring of two disease-centered networks. The accurate and balanced performance of SNP-IN tool makes it readily available to study the rewiring of large-scale protein-protein interaction networks, and can be useful for functional annotation of disease-associated SNPs. SNIP-IN tool is freely accessible as a web-server at http://korkinlab.org/snpintool/. Nan Zhao 0002, Jing Ginger Han, Chi-Ren Shyu, Dmitry Korkin |
PLoS Comput. Biol. | 3 |
| 2013 | Object behavior mining from mitochondiral shape databases for degenerated nerve disease studiesabstractTo make a biological information system valuable for the life sciences community, data analytics have to provide explainable results that can demonstrate the linkages between measurable features and biological meaning of interest. Neither traditional classification nor information retrieval methods are sufficient for biomedical applications if such linkages cannot be established. An example area is the understanding of mitochondrial dynamics in Drosophila segmental nerves. Simple shape analysis or image classification of mitochondrial objects is a far cry from in-depth study of the temporal patterns and transitions of mitochondria which are essential organelles of eukaryotic cells where mitochondria undergo frequent shape changes as well as active movement. We present a computational approach to perform mitochondrial shape analysis, pattern mining, and dynamic behavior retrievals. The capability of the computational methods may provide new insights into scientific discoveries for degenerated nerve diseases. Hongfei Cao, Jing Ginger Han, Chi-Ren Shyu |
BIBM | 3 |
| 2013 | PBSalign: Sequence-independent local alignment of protein binding sitesabstractAccurate alignment of protein-protein binding sites can aid in protein docking studies and constructing templates for predicting structure of protein complexes, along with in-depth understanding of evolutionary and functional relationships. However, over the past three decades, structural alignment algorithms have focused predominantly on global alignments with little effort on the alignment of local interfaces. In this paper, we introduce the PBSalign (Protein-protein Binding Site alignment) method, which integrates techniques in graph theory, 3D localized shape analysis, geometric scoring, and utilization of physiochemical and geometrical properties. Computational results demonstrate that PBSalign is capable of identifying similar homologous and analogous binding sites accurately and performing alignments with better geometric match measures than existing protein-protein interface comparison tools. The proportion of better alignment quality generated by PBSalign is 46%, 56%, and 70% more than Align as judged by the average match index (MI), similarity index (SI), and structural alignment score (SAS), respectively. PBSalign provides the life science community a fast and accurate solution to binding-site alignment while striking the balance between topological details and computational complexity. Bin Pang 0001, David Schlessman, Xingyan Kuang, Nan Zhao 0002, Daniel Shyu, Dmitry Korkin, Chi-Ren Shyu |
BIBM | 7 |
| 2013 | Mining repetitive sequences using a big data ecosystemabstractIdentifying repetitive gene sequences occurring within DNA sequences that span a collection of species is a challenge that is conceptually simple yet computationally challenging. Biological research suggests that certain regions within genomic sequences may be unchanged for hundreds of millions of years; understanding and identifying these highly preserved regions is a major challenge faced by bioinformaticians. Taking an evolutionary perspective on DNA, pinpointing these repetitive sequences is the first step to understanding functional similarities and diversities. The difficulty of this problem arises from the volume of the data required for analysis; it grows with every genome that is sequenced. Traditional approaches used to identify repetitive sequences often require the pair-wise comparison of chromosomes, which takes a significant amount of time to gather results. When comparing n chromosomes, n(n-l) individual comparisons must be made. To avoid exhaustive pair-wise comparisons, we designed an algorithm that partitions genomic sequences into search key values representing potential repetitive sequences, which are hashed into bins. With the introduction of new genomes, we only process the new sequences and aggregate new results with those that were previously processed. Michael Phinney, Hongfei Cao, Andi Dhroso, Chi-Ren Shyu |
BIBM | 4 |
| 2013 | TeleMDID: Mobile technology applications for interactive diagnoses in teledermatology clinicsabstractOver the past two decades, teledermatology has typically applied a hybrid model with both live-interactive and store-and-forward imaging technology during the diagnosis process. The primary challenges for teledermatology are integrations of diagnostic information and images captured by heterogeneous digital cameras used by the far sites, image sharing tools by the dermatologists and patients, and means for longitudinal diagnosis of skin diseases. These challenges not only cause disruptive clinic flow, but also limit the capability for the efficacy of diagnoses. The objectives of this study are to adapt mobile applications from in-person clinical to telehealth settings, and evaluate and compare its usability, effectiveness, and changes in clinic flow. Mirna Becevic, Blake Anderson, Jing Ginger Han, E. Rachel Mutrux, Lanis Hicks, Karen Edison, Chi-Ren Shyu |
Healthcom | 7 |
| 2013 | Comparing limb-volume measurement techniques: 3D models from an infrared depth sensor versus water displacementabstractIn our previous work, a new method for measuring limb volume based on infrared depth sensors was presented. The system, which can be operated in the comfort of our homes, allows for the early detection of swelling associated with lymphedema - a chronic disease caused by failure in the lymphatic system. Early detection and management can significantly reduce the potential for symptoms and complications; however, many patients fail to seek medical assistance at the first sign of the disease. So, the proposed system can potentially affect the lives of nearly 500,000 people in the U.S. who suffer from lymphedema with over 2.6 million breast cancer survivors and over 230,000 new cases every year1who are at-risk for developing this disease at some point in their life. In this paper, a series of improvements made to the system is presented. The changes led to the complete automation of the process of 3D imaging the arms. The proposed technique for limb-volume measurement was compared with the water displacement and the perometry. Being an ongoing research, the results presented here are limited to 14 arms of healthy volunteers. In the future, test will include a larger number of limbs of healthy as well as cancer patients. Guannan Lu, Guilherme N. DeSouza, Jane Armer, Chi-Ren Shyu |
Healthcom | 4 |
| 2012 | A Study of Dermatological Image Search, Archive, and Use Behaviors
Blake Anderson, Jonathan Dyer, Chi-Ren Shyu |
AMIA | 5 |
| 2012 | Quantifying the quality of imaging interpretations
Marius Petruc, Mihail Popescu, John L. Fresen, Dmitriy Shin, Chi-Ren Shyu |
AMIA | 5 |
| 2012 | Improving disease management through a mobile application for lymphedema patientsabstractBreast cancer-related lymphedema (LE) is one type of incurable progressive chronic disease caused by cancer treatment or surgery that damages a patient's lymphatic system. Many patients are unaware of available treatment and places to seek for help with LE. This paper introduces a framework for a mobile application package for LE patients or patients at-risk to 1) locate available trained therapists using embedded Google Map technology; 2) monitor disease progression using profile match and rules mining; and 3) acquire up-to-date LE research findings based on patient characteristics. It aims to increase patients' accessibility to available resources and improve patient's quality of life by a mHealth-driven disease management approach. Christine Shuyu Xu, Blake Anderson, Jane Armer, Chi-Ren Shyu |
Healthcom | 4 |
| 2012 | Associative semantic ranking of satellite images using PathFinder Network Scaling ensemble methodsabstractThis article proposes a methodology to reduce overfitting when ranking high-resolution satellite images by domain semantics. Our approach uses PathFinder Network Scaling ensemble methods. We generate cross-fold co-occurrence matrices for relevance of feature subspaces to each semantic. Each matrix is then reduced using the PathFinder network scaling algorithm. Irrelevant nodes are removed using node strength metrics resulting in an optimized model for ranking by semantic that generalizes better to new images. The experiments show that, when using this approach, the quality of ranking by semantic can be significantly improved. Results show that Mean Average Precision (MAP) of ranking over cross-fold experiments increased by a 13.2% while standard deviation of MAP was reduced by 16.8% relatively to experiments without PathFinder network scaling. Adrian Barb, Chi-Ren Shyu |
IGARSS | 2 |
| 2012 | Fast protein binding site comparisons using visual words representationabstractMOTIVATION: Finding geometrically similar protein binding sites is crucial for understanding protein functions and can provide valuable information for protein-protein docking and drug discovery. As the number of known protein-protein interaction structures has dramatically increased, a high-throughput and accurate protein binding site comparison method is essential. Traditional alignment-based methods can provide accurate correspondence between the binding sites but are computationally expensive. RESULTS: In this article, we present a novel method for the comparisons of protein binding sites using a 'visual words' representation (PBSword). We first extract geometric features of binding site surfaces and build a vocabulary of visual words by clustering a large set of feature descriptors. We then describe a binding site surface with a high-dimensional vector that encodes the frequency of visual words, enhanced by the spatial relationships among them. Finally, we measure the similarity of binding sites by utilizing metric space operations, which provide speedy comparisons between protein binding sites. Our experimental results show that PBSword achieves a comparable classification accuracy to an alignment-based method and improves accuracy of a feature-based method by 36% on a non-redundant dataset. PBSword also exhibits a significant efficiency improvement over an alignment-based method. Bin Pang 0001, Nan Zhao 0002, Dmitry Korkin, Chi-Ren Shyu |
Bioinform. | 4 |
| 2012 | Role of domain knowledge in developing user-centered medical-image indexingabstractAbstract An efficient and robust medical‐image indexing procedure should be user‐oriented. It is essential to index the images at the right level of description and ensure that the indexed levels match the user's interest level. This study examines 240 medical‐image descriptions produced by three different groups of medical‐image users (novices, intermediates, and experts) in the area of radiography. This article reports several important findings: First, the effect of domain knowledge has a significant relationship with the use of semantic image attributes in image‐users' descriptions. We found that experts employ more high‐level image attributes which require high‐reasoning or diagnostic knowledge to search for a medical image (Abstract Objects and Scenes) than do novices; novices are more likely to describe some basic objects which do not require much radiological knowledge to search for an image they need (Generic Objects) than are experts. Second, all image users in this study prefer to use image attributes of the semantic levels to represent the image that they desired to find, especially using those specific‐level and scene‐related attributes. Third, image attributes generated by medical‐image users can be mapped to all levels of the pyramid model that was developed to structure visual information. Therefore, the pyramid model could be considered a robust instrument for indexing medical imagery. Sanda Erdelez, Carla Allen, Blake Anderson, Hongfei Cao, Chi-Ren Shyu |
J. Assoc. Inf. Sci. Technol. | 6 |
| 2012 | Multi-Index Multi-Object Content-Based RetrievalabstractIn many large-scale content-based retrieval (CBR) applications, the input to the search process is a complex query that may be composed of several constituent parts. The proposed approach performs CBR queries by breaking down a complex query into several smaller heterogeneous queries. Object-based queries in an imagery search application can be performed by executing a search over several distinct feature space indexes. For example, CBR indexes may exist for spectral, texture, and shape feature vectors extracted from objects. A query for similar objects can be completed by aggregating the results from these multiple indexes. Complementing this concept, a multi-object search can be used to identify relevant groups of objects which match a given set of query objects. For example, a set of objects identified in satellite imagery could be used as a CBR query in order to identify similar groups of objects. Thus, a query can be performed for each object, and these results can be aggregated into multi-object search results by determining the optimal match of the query objects to those in each resulting group. We introduce the absence penalty method and obligatory object query algorithms for performing multi-index and multi-object CBR searches and provide experimental results that show that the proposed approaches efficiently provide search results with a high degree of precision with minimal error. The experimental results shown demonstrate the efficiency and accuracy of the proposed methods; moreover, through the fusion of multi-index and multi-object search techniques, we are able to construct new sophisticated query mechanisms. Matthew N. Klaric, Grant J. Scott, Chi-Ren Shyu |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2011 | Computable visually observed phenotype ontological framework for plantsabstractBACKGROUND: The ability to search for and precisely compare similar phenotypic appearances within and across species has vast potential in plant science and genetic research. The difficulty in doing so lies in the fact that many visual phenotypic data, especially visually observed phenotypes that often times cannot be directly measured quantitatively, are in the form of text annotations, and these descriptions are plagued by semantic ambiguity, heterogeneity, and low granularity. Though several bio-ontologies have been developed to standardize phenotypic (and genotypic) information and permit comparisons across species, these semantic issues persist and prevent precise analysis and retrieval of information. A framework suitable for the modeling and analysis of precise computable representations of such phenotypic appearances is needed. RESULTS: We have developed a new framework called the Computable Visually Observed Phenotype Ontological Framework for plants. This work provides a novel quantitative view of descriptions of plant phenotypes that leverages existing bio-ontologies and utilizes a computational approach to capture and represent domain knowledge in a machine-interpretable form. This is accomplished by means of a robust and accurate semantic mapping module that automatically maps high-level semantics to low-level measurements computed from phenotype imagery. The framework was applied to two different plant species with semantic rules mined and an ontology constructed. Rule quality was evaluated and showed high quality rules for most semantics. This framework also facilitates automatic annotation of phenotype images and can be adopted by different plant communities to aid in their research. CONCLUSIONS: The Computable Visually Observed Phenotype Ontological Framework for plants has been developed for more efficient and accurate management of visually observed phenotypes, which play a significant role in plant genomics research. The uniqueness of this framework is its ability to bridge the knowledge of informaticians and plant science researchers by translating descriptions of visually observed phenotypes into standardized, machine-understandable representations, thus enabling the development of advanced information retrieval and phenotype annotation analysis tools for the plant science community. Jaturon Harnsomburana, Jason M. Green, Adrian Barb, Mary L. Schaeffer, Leszek Vincent, Chi-Ren Shyu |
BMC Bioinform. | 6 |
| 2011 | A SNOMED supported ontological vector model for subclinical disorder detection using EHR similarity
Lawrence Wing-Chi Chan, Ying Liu 0004, Chi-Ren Shyu, Iris F. F. Benzie |
Eng. Appl. Artif. Intell. | 3 |
| 2011 | Entropy-Balanced Bitmap Tree for Shape-Based Object Retrieval From Large-Scale Satellite Imagery DatabasesabstractIn this paper, we present a novel indexing structure that was developed to efficiently and accurately perform content-based shape retrieval of objects from a large-scale satellite imagery database. Our geospatial information retrieval and indexing system, GeoIRIS, contains 45 GB of high-resolution satellite imagery. Objects of multiple scales are automatically extracted from satellite imagery and then encoded into a bitmap shape representation. This shape encoding compresses the total size of the shape descriptors to approximately 0.34% of the imagery database size. We have developed the entropy-balanced bitmap (EBB) tree, which exploits the probabilistic nature of bit values in automatically derived shape classes. The efficiency of the shape representation coupled with the EBB tree allows us to index approximately 1.3 million objects for fast content-based retrieval of objects by shape. Grant J. Scott, Matthew N. Klaric, Curt H. Davis, Chi-Ren Shyu |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2010 | An accurate classification of native and non-native protein-protein interactions using supervised and semi-supervised learning approachesabstractThe progress in experimental and computational structural biology has led to a rapid growth of experimentally resolved structures and computational models of protein-protein interactions. However, distinguishing between the physiological and non-physiological interactions remains a challenging problem. In this work, two related problems of interface classification have been addressed. The first problem is concerned with classification of the physiological and crystal-packing interactions. The second problem deals with the classification of the physiological interactions, or their accurate models, and decoys obtained from the inaccurate docking models. We have defined a universal set of interface features and employed supervised and semi-supervised learning approaches to accurately classify the interactions in both problems. Furthermore, we formulated the second problem as a semi-supervised learning problem and employed a transductive SVM to improve the accuracy of classification. Finally, we showed that using the scoring functions from the obtained classifiers, one can improve the accuracy of the docking methods. Nan Zhao 0002, Bin Pang 0001, Chi-Ren Shyu, Dmitry Korkin |
BIBM | 3 |
| 2010 | Visual information mining and ranking using graded relevance assessments in satellite image databasesabstractWith recent technological advances, the geospatial industry produces digital image data at an astonishing rate. Such large amounts of data need to be analyzed for visual content in a timely fashion. For in-depth analysis of the geospatial there is a need to find efficient methods to process the visual information into actionable knowledge. One of the most promising methods is to evaluate the relevance of geospatial images to domain-specific visual semantics. Most of existing methods for annotating semantic meaning to geospatial images are trained using binary feedback from users. Such approaches may lead to suboptimal models especially due to the fact that semantic relevance of images is rarely a binary problem. In this paper, we report an algorithm to link low-level image features with high-level visual semantics using graded relevance feedback from image analysts. This linkage is done using flexible possibility functions that mathematically model the existence of visual semantics in new images added to the database. Our experimental results show that our technique improves the knowledge discovery process as evidenced by increased mean average precision of semantic queries. Adrian Barb, Chi-Ren Shyu |
IGARSS | 2 |
| 2010 | Progressive spatial clustering of content-based satellite imagery retrieval resultsabstractThe ProgressiveDBSCAN algorithm allows for the progressive clustering of results from a geospatial information retrieval system. Results can be clustered by a combination of both their spatial and non-spatial attributes. The benefit of this clustering is that users are able to sort through the results returned from a geospatial information retrieval system in a spatial context. No longer are results from disparate locations presented to the user, but instead compact spatial clusters are displayed. There is a 98% reduction in the spatial distance between consecutive CBIR results and the spatial distance between the compact clusters; this leads to more efficient analysis of results by reducing the amount of time users spend context switching while on average only adding a few seconds to the query time. Matthew N. Klaric, Grant J. Scott, Chi-Ren Shyu |
IGARSS | 3 |
| 2010 | Visual-Semantic Modeling in Content-Based Geospatial Information Retrieval Using Associative Mining TechniquesabstractAutomatic learning of geospatial intelligence is challenging due to the complexity of articulating knowledge from visual patterns and to the ever-increasing quantities of image data generated on a daily basis. In this setting, human inspection and annotation is subjective and, more importantly, impractical. In this letter, we propose a knowledge-discovery algorithm that uses content-based methods to link low-level image features with high-level visual semantics in an effort to automate the process of retrieving semantically similar images. Our algorithm represents geospatial images by using a high-dimensional feature vector and generates a set of association rules that correlate semantic terms with visual patterns represented by discrete feature intervals. We also provide a mathematical model to customize the relevance of feature measurements to semantic assignments as well as methods of querying by semantics and by example. Adrian Barb, Chi-Ren Shyu |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2009 | Efficient SCOP-fold classification and retrieval using index-based protein substructure alignmentsabstractMOTIVATION: To investigate structure-function relationships, life sciences researchers usually retrieve and classify proteins with similar substructures into the same fold. A manually constructed database, SCOP, is believed to be highly accurate; however, it is labor intensive. Another known method, DALI, is also precise but computationally expensive. We have developed an efficient algorithm, namely, index-based protein substructure alignment (IPSA), for protein-fold classification. IPSA constructs a two-layer indexing tree to quickly retrieve similar substructures in proteins and suggests possible folds by aligning these substructures. RESULTS: Compared with known algorithms, such as DALI, CE, MultiProt and MAMMOTH, on a sample dataset of non-redundant proteins from SCOP v1.73, IPSA exhibits an efficiency improvement of 53.10, 16.87, 3.60 and 1.64 times speedup, respectively. Evaluated on three different datasets of non-redundant proteins from SCOP, average accuracy of IPSA is approximately equal to DALI and better than CE, MAMMOTH, MultiProt and SSM. With reliable accuracy and efficiency, this work will benefit the study of high-throughput protein structure-function relationships. AVAILABILITY: IPSA is publicly accessible at http://ProteinDBS.rnet.missouri.edu/IPSA.php Pin-Hao Chi, Bin Pang 0001, Dmitry Korkin, Chi-Ren Shyu |
Bioinform. | 4 |
| 2009 | Countering imbalanced datasets to improve adverse drug event predictive models in labor and delivery
Laritza M. Taft, R. Scott Evans, Chi-Ren Shyu, Marlene J. Egger, Nitesh V. Chawla, Joyce A. Mitchell, Sidney N. Thornton, Bruce E. Bray, Michael W. Varner |
J. Biomed. Informatics | 3 |
| 2008 | Ontology Driven Content Mining and Semantic Queries for Satellite Image DatabasesabstractExtracting domain-specific knowledge from image databases is challenging and requires a deep understanding of the domain. For example, in the geospatial domain knowledge discovery computationally expensive due to the huge amount of generated imagery. Existing content-based image retrieval systems utilize models that are trained and optimized to experts' knowledge using expert-in-the-loop approaches. However, such approaches may lead to suboptimal models especially when the number of training images is small. In this paper, we propose incorporating existing domain knowledge resources into knowledge discovery. More specifically, we have developed methods for using ontological relationships between geospatial semantics to oversample under-represented semantics. Our experimental results show that our technique improves the knowledge discovery process, as evidenced by increased precision of semantic queries. Adrian Barb, Chi-Ren Shyu |
IGARSS (3) | 2 |
| 2008 | GeoCDX: An Automated Change Detection & Exploitation System for High Resolution Satelite ImageryabstractThe demand for high-resolution commercial satellite imagery (HR-CSI) has increased significantly over the last 5 years for a wide variety of applications. This demand has driven an increase in volume, frequency of acquisition, and spatial resolution of HR-CSI. In turn, this has spurred the need for more accurate and time-efficient processing tools for analyzing geospatial information to support user-specific applications. One such application is change detection between multi-temporal HR-CSI data. The significant increase in quantity and quality of multi-temporal HR-CSI data makes traditional manual analysis impractical. Thus, there is a need for a fully automated change detection system that not only identifies areas of change, but also allows users to filter, sort through, and analyze areas of change quickly and efficiently. Here we describe a tool - GeoCDX (Geospatial Change Detection and Exploitation) - to meet this need. GeoCDX is an integrated system that performs image ingestion and registration, feature extraction, and change analysis. The change detection results from GeoCDX are web accessible with additional interfaces to Google Earth™ (GE) and Google Maps™ (GM). In the near future GeoCDX will be integrated with its sister system GeoIRIS (Geospatial Information Retrieval and Indexing System). This integration will produce a very powerful HR-CSI analysis package with the change detection capabilities of GeoCDX and the content-based image retrieval system of GeoIRIS. Ozy Sjahputera, Curt H. Davis, Brian C. Claywell, Nicholas J. Hudson 0002, James Keller 0001, Michael G. Vincent, Matthew N. Klaric, Chi-Ren Shyu |
IGARSS (5) | 9 |
| 2008 | The effects of conceptual description and search practice on users' mental models and information seeking in a case-based reasoning retrieval system
Wu He, Sanda Erdelez, Feng-Kwei Wang, Chi-Ren Shyu |
Inf. Process. Manag. | 4 |
| 2007 | User-specific semantics for modeling content-based geospatial knowledgeabstractModern technology enables organizations to build huge geospatial data repositories. But collecting and storing information is not sufficient if it is not backed-up by accurate and flexible methods of extracting knowledge encapsulated in data. Image analysts use individualized models to represent visual patterns found in images. These models may not coincide with the models created by computer algorithms. To be successful, computer systems need to adapt to the subjective views of image analysts. In this article we introduce a novel method for fast query customization that provides users with individualized computer models to assign semantics to visual patterns in images. These models are evolved according to user input by adjusting possibility functions that mathematically map the assignment of semantics into low-level features. Our approach provides a flexible method for querying image databases using semantics, and potentially provides a knowledge exchange method for the geospatial community. Adrian Barb, Chi-Ren Shyu |
IGARSS | 2 |
| 2007 | GeoIRIS: Geospatial Information Retrieval and Indexing System - Content Mining, Semantics Modeling, and Complex QueriesabstractSearching for relevant knowledge across heterogeneous geospatial databases requires an extensive knowledge of the semantic meaning of images, a keen eye for visual patterns, and efficient strategies for collecting and analyzing data with minimal human intervention. In this paper, we present our recently developed content-based multimodal Geospatial Information Retrieval and Indexing System (GeoIRIS) which includes automatic feature extraction, visual content mining from large-scale image databases, and high-dimensional database indexing for fast retrieval. Using these underpinnings, we have developed techniques for complex queries that merge information from heterogeneous geospatial databases, retrievals of objects based on shape and visual characteristics, analysis of multiobject relationships for the retrieval of objects in specific spatial configurations, and semantic models to link low-level image features with high-level visual descriptors. GeoIRIS brings this diverse set of technologies together into a coherent system with an aim of allowing image analysts to more rapidly identify relevant imagery. GeoIRIS is able to answer analysts' questions in seconds, such as "given a query image, show me database satellite images that have similar objects and spatial relationship that are within a certain radius of a landmark." Chi-Ren Shyu, Matthew N. Klaric, Grant J. Scott, Adrian Barb, Curt H. Davis, Kannappan Palaniappan |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2007 | Knowledge-Driven Multidimensional Indexing Structure for Biomedical Media Database RetrievalabstractToday, biomedical media data are being generated at rates unimaginable only years ago. Content-based retrieval of biomedical media from large databases is becoming increasingly important to clinical, research, and educational communities. In this paper, we present the recently developed entropy balanced statistical (EBS) k-d tree and its applications to biomedical media, including a high-resolution computed tomography (HRCT) lung image database and the first real-time protein tertiary structure search engine. Our index utilizes statistical properties inherent in large-scale biomedical media databases for efficient and accurate searches. By applying concepts from pattern recognition and information theory, the EBS k-d tree is built through top-down decision tree induction. Experimentation shows similarity searches against a protein structure database of 53 363 structures consistently execute in less than 8.14 ms for the top 100 most similar structures. Additionally, we have shown improved retrieval precision over adaptive and statistical k-d trees. Retrieval precision of the EBS k-d tree is 81.6% for content-based retrieval of HRCT lung images and 94.9% at 10% recall for protein structure similarity search. The EBS k-d tree has enormous potential for use in biomedical applications embedded with ground-truth knowledge and multidimensional signatures. Grant J. Scott, Chi-Ren Shyu |
IEEE Trans. Inf. Technol. Biomed. | 2 |
| 2006 | Representing Clinical Questions by Semantic Type for Better Classification
Tetsuya Kobayashi, Chi-Ren Shyu |
AMIA | 2 |
| 2006 | Knowledge Discovery using Domain-Concept Mining Approach for the Behavioral Risk Factor Surveillance System (BRFSS) Data
Wannapa Kay Mahamaneerat, Chi-Ren Shyu |
AMIA | 2 |
| 2006 | Applying Sequential Forward Floating Selection to Protein Structure Prediction with a Study of HIV-1 RP
Jeff Reneker, Chi-Ren Shyu |
AMIA | 2 |
| 2006 | Predicting Cancer Interaction Networks Using Text-Mining and Structure Understanding
Christopher M. Topinka, Chi-Ren Shyu |
AMIA | 2 |
| 2006 | Predicting Ranked SCOP Domains by Mining Associations of Visual Contents in Distance Matrices
Pin-Hao Chi, Chi-Ren Shyu |
APBC | 2 |
| 2006 | Fusion of Spectral and Spatial Information for Automated Change Detection in High Resolution Satellite ImageryabstractHere we propose a pixel-based change detector utilizing both spectral and spatial features. Traditional pixel- based change detection methods that only utilize spectral data are inherently sensitive to spectral variation. This presents problems in mitigating the impact of spectral changes due to uninteresting types of change without severely limiting the sensitivity of the detector. Here we introduce a change detector that utilizes both spectral and spatial information, including linear features and texture measures, as a method to decrease sensitivity to spectral variation and increase detection rates. For each pixel, a fuzzy value representing the similarity of each pixel feature (both spectral and spatial) is computed. All features are then fused into a single overall similarity score by weighted averaging. This is followed by thresholding and morphological extraction of the detected regions of change. The algorithm was evaluated using panchromatic and pan-sharpened multi-spectral imagery of Springfield, Missouri acquired during different seasons and covering approximately 20 square kilometers of urban, suburban, and rural terrain. Preliminary results show a 70% change detection probability for types of change unrelated to seasonal variation with rates of only 2.4 uninteresting detections per km 2 and 0.09 false alarms per km 2 . Brian C. Claywell, Curt H. Davis, Chi-Ren Shyu |
IGARSS | 3 |
| 2006 | Mining Visual Associations from User Feedback for Weighting Multiple Indexes in Geospatial Image RetrievalabstractGeospatial content-based image retrieval (CBIR) systems can be used to query for visually similar images by identifying similar patterns between a query image and those in the database. When several different classes of features are used, some queries require that each class should be given a different degree of weight; to this end, CBIR indexes are built for each class of features. This paper proposes an approach for weighting multiple indexes in a geospatial CBIR system by mining information from user feedback. After a small number of iterations of relevance feedback and data mining, index weights can be determined dynamically per query. Using this technique geospatial retrieval system precision of results increased from 70% to 79% after 5 iterations of feedback. Matthew N. Klaric, Grant J. Scott, Chi-Ren Shyu |
IGARSS | 3 |
| 2006 | A Framework for Geospatial Satellite Imagery Retrieval SystemsabstractThis paper presents a framework for the efficient retrieval of satellite imagery from large-scale databases. With the ever-expanding volume of imagery being acquired from satellite platforms it has become increasingly important to locate specific areas of interest within a large database of images. Identifying relevant areas within image databases can be thought of as finding the needle in the haystack problem; too often for a particular task or application there exist a small number of useful images hidden among millions of images. The motivation behind the work presented here is that through the use of geospatial image retrieval systems, the number of scenes that image analysts must manually examine may be decreased dramatically. By using a geospatial image retrieval system as a tool, analysts no longer must manually examine the entire database of imagery, but instead can limit their search to a subset identified by our retrieval system. The techniques that are introduced in this paper have been developed in our image retrieval system named GeoIRIS: Geospatial Information Retrieval and Indexing System. Matthew N. Klaric, Grant J. Scott, Chi-Ren Shyu, Curt H. Davis, Kannappan Palaniappan |
IGARSS | 3 |
| 2006 | Knowledge Discovery by Mining Association Rules and Temporal-Spatial Information from Large-Scale Geospatial Image DatabasesabstractDiscovering relevant knowledge from large-scale geospatial image databases is challenging because of the complexity of describing visual semantics, the computational cost of processing petabytes of data, and the difficulty in summarizing and presenting knowledge. In this paper, we revisit a selective set of core data mining algorithms, namely association rules mining, spatial mining, and temporal mining. We then customize these algorithms using visual content and potential objects extracted from geospatial image databases with other relevant information, such as text-based annotations. Queries utilizing the mining results are also discussed in this paper. These mining and query processing algorithms play an important role in GeoIRIS- Geospatial Information Retrieval and Indexing System. Chi-Ren Shyu, Matthew N. Klaric, Grant J. Scott, Wannapa Kay Mahamaneerat |
IGARSS | 1 |
| 2006 | A fast SCOP fold classification system using content-based E-Predict algorithmabstractBACKGROUND: Domain experts manually construct the Structural Classification of Protein (SCOP) database to categorize and compare protein structures. Even though using the SCOP database is believed to be more reliable than classification results from other methods, it is labor intensive. To mimic human classification processes, we develop an automatic SCOP fold classification system to assign possible known SCOP folds and recognize novel folds for newly-discovered proteins. RESULTS: With a sufficient amount of ground truth data, our system is able to assign the known folds for newly-discovered proteins in the latest SCOP v1.69 release with 92.17% accuracy. Our system also recognizes the novel folds with 89.27% accuracy using 10 fold cross validation. The average response time for proteins with 500 and 1409 amino acids to complete the classification process is 4.1 and 17.4 seconds, respectively. By comparison with several structural alignment algorithms, our approach outperforms previous methods on both the classification accuracy and efficiency. CONCLUSION: In this paper, we build an advanced, non-parametric classifier to accelerate the manual classification processes of SCOP. With satisfactory ground truth data from the SCOP database, our approach identifies relevant domain knowledge and yields reasonably accurate classifications. Our system is publicly accessible at http://ProteinDBS.rnet.missouri.edu/E-Predict.php. Pin-Hao Chi, Chi-Ren Shyu, Dong Xu 0002 |
BMC Bioinform. | 2 |
| 2005 | Automated object extraction through simplification of the differential morphological profile for high-resolution satellite imagery
Matthew N. Klaric, Grant J. Scott, Chi-Ren Shyu, Curt H. Davis |
IGARSS | 3 |
| 2005 | Mining image content associations for visual semantic modeling in geospatial information indexing and retrieval
Chi-Ren Shyu, Adrian Barb, Curt H. Davis |
IGARSS | 1 |
| 2005 | Refined repetitive sequence searches utilizing a fast hash function and cross species information retrievalsabstractBACKGROUND: Searching for small tandem/disperse repetitive DNA sequences streamlines many biomedical research processes. For instance, whole genomic array analysis in yeast has revealed 22 PHO-regulated genes. The promoter regions of all but one of them contain at least one of the two core Pho4p binding sites, CACGTG and CACGTT. In humans, microsatellites play a role in a number of rare neurodegenerative diseases such as spinocerebellar ataxia type 1 (SCA1). SCA1 is a hereditary neurodegenerative disease caused by an expanded CAG repeat in the coding sequence of the gene. In bacterial pathogens, microsatellites are proposed to regulate expression of some virulence factors. For example, bacteria commonly generate intra-strain diversity through phase variation which is strongly associated with virulence determinants. A recent analysis of the complete sequences of the Helicobacter pylori strains 26695 and J99 has identified 46 putative phase-variable genes among the two genomes through their association with homopolymeric tracts and dinucleotide repeats. Life scientists are increasingly interested in studying the function of small sequences of DNA. However, current search algorithms often generate thousands of matches -- most of which are irrelevant to the researcher. RESULTS: We present our hash function as well as our search algorithm to locate small sequences of DNA within multiple genomes. Our system applies information retrieval algorithms to discover knowledge of cross-species conservation of repeat sequences. We discuss our incorporation of the Gene Ontology (GO) database into these algorithms. We conduct an exhaustive time analysis of our system for various repetitive sequence lengths. For instance, a search for eight bases of sequence within 3.224 GBases on 49 different chromosomes takes 1.147 seconds on average. To illustrate the relevance of the search results, we conduct a search with and without added annotation terms for the yeast Pho4p binding sites, CACGTG and CACGTT. Also, a cross-species search is presented to illustrate how potential hidden correlations in genomic data can be quickly discerned. The findings in one species are used as a catalyst to discover something new in another species. These experiments also demonstrate that our system performs well while searching multiple genomes -- without the main memory constraints present in other systems. CONCLUSION: We present a time-efficient algorithm to locate small segments of DNA and concurrently to search the annotation data accompanying the sequence. Genome-wide searches for short sequences often return hundreds of hits. Our experiments show that subsequently searching the annotation data can refine and focus the results for the user. Our algorithms are also space-efficient in terms of main memory requirements. Source code is available upon request. Jeff Reneker, Chi-Ren Shyu |
BMC Bioinform. | 2 |
| 2005 | A Fast Protein Structure Retrieval System Using Image-Based Distance Matrices and Multidimensional IndexabstractIndexing protein tertiary structures has been shown to provide a scalable solution for structure-to-structure comparisons in large protein structure retrieval systems. To conduct similarity searches against 53,356 polypeptide chains in a database with real-time responses, two critical issues must be addressed, information extraction and suitable indexing. In this paper, we apply computer vision techniques to extract the predominant information encoded in each 2D distance matrix, generated from 3D coordinates of protein chains. Distance matrices are capable of representing specific protein structural topologies, and similar proteins will generate similar matrices. Once meaningful features are extracted from distance images, an advanced indexing structure, Entropy Balanced Statistical (EBS) k-d tree, can be utilized to index the multidimensional data. With a limited amount of training data from domain experts, namely structural classification of a subset of available protein chains, we apply various techniques in the pattern recognition field to determine clusters of proteins in the multi-dimensional feature space. Our system is able to recall search results in a ranked order from the protein database in seconds, exhibiting a reasonably high degree of precision. Pin-Hao Chi, Grant J. Scott, Chi-Ren Shyu |
Int. J. Softw. Eng. Knowl. Eng. | 3 |
| 2005 | Knowledge representation and sharing using visual semantic modeling for diagnostic medical image databasesabstractInformation technology offers great opportunities for supporting radiologists' expertise in decision support and training. However, this task is challenging due to difficulties in articulating and modeling visual patterns of abnormalities in a computational way. To address these issues, well established approaches to content management and image retrieval have been studied and applied to assist physicians in diagnoses. Unfortunately, most of the studies lack the flexibility of sharing both explicit and tacit knowledge involved in the decision making process, while adapting to each individual's opinion. In this paper, we propose a knowledge repository and exchange framework for diagnostic image databases called "evolutionary system for semantic exchange of information in collaborative environments" (Essence). This framework uses semantic methods to describe visual abnormalities, and offers a solution for tacit knowledge elicitation and exchange in the medical domain. Also, our approach provides a computational and visual mechanism for associating synonymous semantics of visual abnormalities. We conducted several experiments to demonstrate the system's capability of matching synonym terms, and the benefit of using tacit knowledge in improving the meaningfulness of semantic queries. Adrian Barb, Chi-Ren Shyu, Yash Sethi |
IEEE Trans. Inf. Technol. Biomed. | 2 |
| 2004 | Semantic Integration and Knowledge Exchange for Diagnostic Medical Image DatabasesabstractInformation technology offers great opportunities to radiologists to utilize their expertise in decision support and training. Collaborative approaches in these areas enable physicians to access relevant cases diagnosed by experts from other health care groups. Unfortunately, there is little agreement on a single model of semantic representation and information exchange. In this context, semantic interoperability among heterogeneous groups plays an important role in a collaborative setting. In this paper, we propose a model for semantics integration and knowledge exchange in collaborative environments that feature heterogeneous semantics integration. It provides a computational and visual mechanism to associate synonymous semantics of visual abnormalities related to lung pathologies. We also offer a solution for system level communication that improves the retrieval precision using peer domain expertise. From our experiments we obtained a high degree of matching between different semantics that describe the same visual pattern of lung pathology. Also, our experiments show that, using the knowledge exchange mechanism, the default system setting adjusts well over time to increase the retrieval precision for new users. Adrian Barb, Chi-Ren Shyu, Yash Sethi |
BIBE | 2 |
| 2004 | A Fast Protein Structure Retrieval System Using Image-Based Distance Matrices and Multidimensional IndexabstractIndexing protein structures has been shown to provide a scalable solution for structure-to-structure comparisons in large protein structure retrieval systems. To conduct similarity searches against 46,075 polypeptide chains in a database with real-time responses, two critical issues must be addressed, information extraction and suitable indexing. In this paper, we apply computer vision techniques to extract the predominant information encoded in each 2D distance matrix, generated from 3D coordinates of protein chains. Distance matrices are capable of representing specific protein structural topologies, and similar proteins will generate similar matrices. Once meaningful features are extracted from distance images, an advanced indexing structure, entropy balanced statistical (EBS) k-d tree, can be utilized to index the multidimensional data. With a limited amount of training data from domain experts, namely structural classification of a subset of available protein chains, we apply various techniques in the pattern recognition field to determine clusters of proteins in the multi-dimensional feature space. Our system is able to recall search results in a ranked order from the protein database in seconds, exhibiting a reasonably high degree of precision. Pin-Hao Chi, Grant J. Scott, Chi-Ren Shyu |
BIBE | 3 |
| 2003 | Semantics modeling in diagnostic medical image databases using customized fuzzy membership functionsabstractIt is widely recognized that fuzzy methods play an important role in image database retrieval, especially in the context of semantic queries. Known approaches that use crisp hierarchical semantic networks have been studied and applied to content-based image retrieval (CBIR) to narrow the gap between semantics and image features. Unfortunately, most of the studies lack the flexibility to adapt to an individual's preferences and/or to establish a general-purpose semantic network for sharing the perceptual understanding. In this paper, we propose a semantic query system for diagnostic image database retrieval that uses physician-defined linguistic variables. Users can obtain more desirable retrieval results by creating new, customized semantic terms, and by modeling a suite of membership functions to reflect their preferences. The system brings an increased versatility for image retrieval, and a great amount of possibilities for customizing the semantic terms using customized fuzzy mappings. Our unique approach provides various query methods that use the semantic terms within the domain of HRCT images of the lung and allows individual users to bring the contribution to the common knowledge base. Adrian Barb, Chi-Ren Shyu |
FUZZ-IEEE | 2 |
| 2002 | Using Human Perceptual Categories for Content-Based Retrieval from a Medical Image Database
Chi-Ren Shyu, Christina Pavlopoulou, Avinash C. Kak, Carla E. Brodley, Lynn S. Broderick |
Comput. Vis. Image Underst. | 1 |
| 2001 | Spatial Lesion Indexing for Medical Image Databases Using Force HistogramsabstractIt is often difficult to find a well-principled approach for the selection of a spatial indexing mechanism for medical image databases. Spatial information concerning lesions in medical images is critically important in disease diagnosis and plays an important role in image retrieval. Unfortunately, images are rarely indexed properly for clinically useful retrieval. One example is the well-known R-tree and its variants which index image objects based on their physical locations in an "absolute" way. However, such information is not meaningful in medical content-based image retrieval systems, and the approaches suffer from problems caused by variations in object size and shape, imprecise image centering, etc. A more appropriate approach, which does not require object registration, is to model the spatial relationships between lesions and anatomical landmarks. To convey diagnostic information, lesions must exist in certain locations with regard to landmarks. In this paper, we show that the histogram of forces (which represents the relative position between two objects) provides an efficient spatial indexing mechanism in the medical domain. Chi-Ren Shyu, Pascal Matsakis |
CVPR (2) | 1 |
| 1999 | The Customized-Queries Approach to CBIR Using EMabstractThis paper makes two contributions. The first contribution is an approach called the "customized-queries" approach (CQA) to content-based image retrieval. The second is an algorithm called FSSEM that performs feature selection and clustering simultaneously. The customized queries approach first classifies a query using the features that best differentiate the major classes and then customizes the query to that class by using the features that best distinguish the images within the chosen major class. This approach is motivated by the observation that the features that are most effective in discriminating among images from different classes may not be the most effective for retrieval of visually similar images within a class. This occurs for domains in which not all pairs of images within one class have equivalent visual similarity, i.e., subclasses exists. Because we are not given subclass labels, we must simultaneously find the features that best discriminate the subclasses and at the same time find these subclasses. We use FSSEM to find these features. We apply this approach to content-based retrieval of high-resolution tomographic images of patients with lung disease and show that this approach radically improves the retrieval precision over the traditional approach that performs retrieval using a single feature vector. Jennifer G. Dy, Carla E. Brodley, Avinash C. Kak, Chi-Ren Shyu, Lynn S. Broderick |
CVPR | 4 |
| 1999 | ASSERT: A Physician-in-the-Loop Content-Based Retrieval System for HRCT Image Databases
Chi-Ren Shyu, Carla E. Brodley, Avinash C. Kak, Akio Kosaka, Alex M. Aisen, Lynn S. Broderick |
Comput. Vis. Image Underst. | 1 |