EDBT 2026 Demo / reviewers in the wild / expert
Guadalupe Canahuate
dblp:40/4280
· DBLP profile ↗
29ranked-venue papers
3as first author
10since 2021 · last 2025
0000-0001-5873-5454ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 20 · 3 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 since 2021Artificial intelligence and machine learning · 4 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | PRO-Based Stratification Improves Model Prediction for Toxicity and Survival of Head and Neck Cancer PatientsabstractPatient-Reported Outcomes (PRO) consist of information provided directly by the patients about their health status including symptom ratings. PROs are commonly used in clinical practice to support clinical decision-making and have recently been incorporated into machine learning models to improve risk prediction. In this work, we aim to evaluate whether the inclusion of a patient stratification based on 12-month post-treatment predicted Patient Reported Outcomes improves risk prediction of radiation-induced toxicity and overall survival for head and neck cancer patients. A bidirectional long-short term memory (Bi-LSTM) recurrent neural network was used to model the longitudinal PRO data and to predict symptom ratings 12 months post-treatment. Patients were stratified using hierarchical clustering over the LSTM-predicted data. A logistic regression model was trained to predict Xerostomia at 12 months and a Cox regression model to predict overall survival. Results show that the inclusion of symptom burden clusters derived from the predicted Patient Reported Outcomes improves radiation-induced toxicity and overall survival prediction for head and neck cancer patients. Eric Ababio Anyimadu, Carla Floricel, Serageldin Kamel, Clifton D. Fuller, G. Elisabeta Marai, Guadalupe Canahuate |
IEEE J. Biomed. Health Informatics | 8 |
| 2025 | DITTO: A Visual Digital Twin for Interventions and Temporal Treatment Outcomes in Head and Neck CancerabstractDigital twin models are of high interest to Head and Neck Cancer (HNC) oncologists, who have to navigate a series of complex treatment decisions that weigh the efficacy of tumor control against toxicity and mortality risks. Evaluating individual risk profiles necessitates a deeper understanding of the interplay between different factors such as patient health, spatial tumor location and spread, and risk of subsequent toxicities that can not be adequately captured through simple heuristics. To support clinicians in better understanding tradeoffs when deciding on treatment courses, we developed DITTO, a digital-twin and visual computing system that allows clinicians to analyze detailed risk profiles for each patient, and decide on a treatment plan. DITTO relies on a sequential Deep Reinforcement Learning digital twin (DT) to deliver personalized risk of both long-term and short-term disease outcome and toxicity risk for HNC patients. Based on a participatory collaborative design alongside oncologists, we also implement several visual explainability methods to promote clinical trust and encourage healthy skepticism when using our system. We evaluate the efficacy of DITTO through quantitative evaluation of performance and case studies with qualitative feedback. Finally, we discuss design lessons for developing clinical visual XAI applications for clinical end users. Andrew Wentzel, Serageldin Kamel, Guadalupe Canahuate, Clifton D. Fuller, G. Elisabeta Marai |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2024 | Collaborative Filtering for the Imputation of Patient Reported Outcomes
Eric Ababio Anyimadu, Clifton D. Fuller, G. Elisabeta Marai, Guadalupe Canahuate |
DEXA (1) | 5 |
| 2024 | Roses Have Thorns: Understanding the Downside of Oncological Care Delivery Through Visual Analytics and Sequential Rule MiningabstractPersonalized head and neck cancer therapeutics have greatly improved survival rates for patients, but are often leading to understudied long-lasting symptoms which affect quality of life. Sequential rule mining (SRM) is a promising unsupervised machine learning method for predicting longitudinal patterns in temporal data which, however, can output many repetitive patterns that are difficult to interpret without the assistance of visual analytics. We present a data-driven, human-machine analysis visual system developed in collaboration with SRM model builders in cancer symptom research, which facilitates mechanistic knowledge discovery in large scale, multivariate cohort symptom data. Our system supports multivariate predictive modeling of post-treatment symptoms based on during-treatment symptoms. It supports this goal through an SRM, clustering, and aggregation back end, and a custom front end to help develop and tune the predictive models. The system also explains the resulting predictions in the context of therapeutic decisions typical in personalized care delivery. We evaluate the resulting models and system with an interdisciplinary group of modelers and head and neck oncology researchers. The results demonstrate that our system effectively supports clinical and symptom research. Carla Floricel, Andrew Wentzel, Abdallah Sherif Radwan Mohamed, Clifton D. Fuller, Guadalupe Canahuate, G. Elisabeta Marai |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2023 | Evaluating Autoencoders for Dimensionality Reduction of MRI-derived Radiomics and Classification of Malignant Brain TumorsabstractMalignant brain tumors including parenchymal metastatic (MET) lesions, glioblastomas (GBM), and lymphomas (LYM) account for 29.7% of brain cancers. However, the characterization of these tumors from MRI imaging is difficult due to the similarity of their radiologically observed image features. Radiomics is the extraction of quantitative imaging features to characterize tumor intensity, shape, and texture. Applying machine learning over radiomic features could aid diagnostics by improving the classification of these common brain tumors. However, since the number of radiomic features is typically larger than the number of patients in the study, dimensionality reduction is needed to balance feature dimensionality and model complexity. Autoencoders are a form of unsupervised representation learning that can be used for dimensionality reduction. It is similar to PCA but uses a more complex and non-linear model to learn a compact latent space. In this work, we examine the effectiveness of autoencoders for dimensionality reduction on the radiomic feature space of multiparametric MRI images and the classification of malignant brain tumors: GBM, LYM, and MET. We further aim to address the class imbalances imposed by the rarity of lymphomas by examining different approaches to increase overall predictive performance through multiclass decomposition strategies. Mikayla Biggs, Neetu Soni, Sarv Priya, Girish Bathla, Guadalupe Canahuate |
SSDBM | 6 |
| 2023 | DASS Good: Explainable Data Mining of Spatial Cohort DataabstractDeveloping applicable clinical machine learning models is a difficult task when the data includes spatial information, for example, radiation dose distributions across adjacent organs at risk. We describe the co-design of a modeling system, DASS, to support the hybrid human-machine development and validation of predictive models for estimating long-term toxicities related to radiotherapy doses in head and neck cancer patients. Developed in collaboration with domain experts in oncology and data mining, DASS incorporates human-in-the-loop visual steering, spatial data, and explainable AI to augment domain knowledge with automatic data mining. We demonstrate DASS with the development of two practical clinical stratification models and report feedback from domain experts. Finally, we describe the design lessons learned from this collaborative experience. Andrew Wentzel, Carla Floricel, Guadalupe Canahuate, Mohamed A. Naser, Abdallah S. Mohamed, Clifton D. Fuller, Lisanne van Dijk, G. Elisabeta Marai |
Comput. Graph. Forum | 3 |
| 2022 | A Tale of Two Centers: Visual Exploration of Health Disparities in Cancer CareabstractThe annual incidence of head and neck cancers (HNC) worldwide is more than 550,000 cases, with around 300,000 deaths each year. However, the incidence rates and disease-characteristics of HNC differ between treatment centers and different populations, due to undetermined reasons, which may or not include socioeconomic factors. The multi-faceted and multi-variate nature of the data in the context of the emerging field of health disparities research makes automated analysis impractical. Hence, we present a visual analysis approach to explore the health disparities in the data of HNC patients from two different cohorts at two cancer care centers. Our approach integrates data from multiple sources, including census data and city data, with custom visual encodings and with a nearest neighbor approach. Our design, created in collaboration with oncology experts, makes it possible to analyze the patients' demographic, disease characteristics, treatments and outcomes, and to make significant comparisons of these two cohorts and of individual patients. We evaluate this approach through two case studies performed with domain experts. The results demonstrate that this visual analysis approach successfully accomplishes the goal of comparing two cohorts in terms of different significant factors, and can provide insights into the main source of health disparities between the two centers. Sanjana Srabanti, Michael Tran, Virginie Achim, Clifton D. Fuller, Guadalupe Canahuate, Fabio Miranda 0001, G. Elisabeta Marai |
PacificVis | 5 |
| 2022 | THALIS: Human-Machine Analysis of Longitudinal Symptoms in Cancer TherapyabstractAlthough cancer patients survive years after oncologic therapy, they are plagued with long-lasting or permanent residual symptoms, whose severity, rate of development, and resolution after treatment vary largely between survivors. The analysis and interpretation of symptoms is complicated by their partial co-occurrence, variability across populations and across time, and, in the case of cancers that use radiotherapy, by further symptom dependency on the tumor location and prescribed treatment. We describe THALIS, an environment for visual analysis and knowledge discovery from cancer therapy symptom data, developed in close collaboration with oncology experts. Our approach leverages unsupervised machine learning methodology over cohorts of patients, and, in conjunction with custom visual encodings and interactions, provides context for new patients based on patients with similar diagnostic features and symptom evolution. We evaluate this approach on data collected from a cohort of head and neck cancer patients. Feedback from our clinician collaborators indicates that THALIS supports knowledge discovery beyond the limits of machines or humans alone, and that it serves as a valuable tool in both the clinic and symptom research. Carla Floricel, Nafiul Nipu, Mikayla Biggs, Andrew Wentzel, Guadalupe Canahuate, Lisanne van Dijk, Abdallah Sherif Radwan Mohamed, Clifton D. Fuller, G. Elisabeta Marai |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2021 | Identifying Symptom Clusters Through Association Rule Mining
Mikayla Biggs, Carla Floricel, Lisanne van Dijk, Abdallah Sherif Radwan Mohamed, Clifton D. Fuller, G. Elisabeta Marai, Guadalupe Canahuate |
AIME | 8 |
| 2021 | Predicting late symptoms of head and neck cancer treatment using LSTM and patient reported outcomesabstractPatient-Reported Outcome (PRO) surveys are used to monitor patients' symptoms during and after cancer treatment. Acute symptoms refer to those experienced during treatment and late symptoms refer to those experienced after treatment. While most patients experience severe symptoms during treatment, these usually subside in the late stage. However, for some patients, late toxicities persist negatively affecting the patient's quality of life (QoL). In the case of head and neck cancer patients, PRO surveys are recorded every week during the patient's visit to the clinic and at different follow-up times after the treatment has concluded. In this paper, we model the PRO data as a time-series and apply Long-Short Term Memory (LSTM) neural networks for predicting symptom severity in the late stage. The PRO data used in this project corresponds to MD Anderson Symptom Inventory (MDASI) questionnaires collected from head and neck cancer patients treated at the MD Anderson Cancer Center. We show that the LSTM model is effective in predicting symptom ratings under the RMSE and NRMSE metrics. Our experiments show that the LSTM model also outperforms other machine learning models and time-series prediction models for these data. Guadalupe Canahuate, Lisanne van Dijk, Abdallah Sherif Radwan Mohamed, Clifton D. Fuller, G. Elisabeta Marai |
IDEAS | 2 |
| 2020 | High-dimensional similarity searches using query driven dynamic quantization and distributed indexing
Gheorghi Guzun, Guadalupe Canahuate |
Distributed Parallel Databases | 2 |
| 2020 | Cohort-based T-SSIM Visual Computing for Radiation Therapy Prediction and ExplorationabstractWe describe a visual computing approach to radiation therapy (RT) planning, based on spatial similarity within a patient cohort. In radiotherapy for head and neck cancer treatment, dosage to organs at risk surrounding a tumor is a large cause of treatment toxicity. Along with the availability of patient repositories, this situation has lead to clinician interest in understanding and predicting RT outcomes based on previously treated similar patients. To enable this type of analysis, we introduce a novel topology-based spatial similarity measure, T-SSIM, and a predictive algorithm based on this similarity measure. We couple the algorithm with a visual steering interface that intertwines visual encodings for the spatial data and statistical results, including a novel parallel-marker encoding that is spatially aware. We report quantitative results on a cohort of 165 patients, as well as a qualitative evaluation with domain experts in radiation oncology, data management, biostatistics, and medical imaging, who are collaborating remotely. Andrew Wentzel, Peter Hanula, Timothy Luciani, Baher Elgohari, Hesham Elhalawani, Guadalupe Canahuate, David M. Vock, Clifton D. Fuller, G. Elisabeta Marai |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2019 | Precision Risk Analysis of Cancer Therapy with Interactive Nomograms and Survival PlotsabstractWe present the design and evaluation of an integrated problem solving environment for cancer therapy analysis. The environment intertwines a statistical martingale model and a K Nearest Neighbor approach with visual encodings, including novel interactive nomograms, in order to compute and explain a patient's probability of survival as a function of similar patient results. A coordinated views paradigm enables exploration of the multivariate, heterogeneous and few-valued data from a large head and neck cancer repository. A visual scaffolding approach further enables users to build from familiar representations to unfamiliar ones. Evaluation with domain experts show how this visualization approach and set of streamlined workflows enable the systematic and precise analysis of a patient prognosis in the context of cohorts of similar patients. We describe the design lessons learned from this successful, multi-site remote collaboration. G. Elisabeta Marai, Chihua Ma, Andrew Thomas Burks, Filippo Pellolio, Guadalupe Canahuate, David M. Vock, Abdallah Sherif Radwan Mohamed, Clifton D. Fuller |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2018 | Distributed query-aware quantization for high-dimensional similarity searchesabstractfraction of the points while a constant penalty is applied for the rest of the points. QED not only improves the quality of the distance metric, but also improves query time performance by filtering out non relevant data. We propose a distributed indexing and query algorithm to efficiently compute QED. Our experimental results show improvements in classification accuracy as well as query performance up to one order of magnitude faster than Manhattan-based sequential scan NN queries over datasets with hundreds of dimensions. Gheorghi Guzun, Guadalupe Canahuate |
EDBT | 2 |
| 2016 | Power efficient big data analytics algorithms through low-level operationsabstractWe present an empirical performance evaluation of algorithms that replace arithmetic operations with low-level bit operations for power-aware Big Data processing. Specifically, we compare two different data structures in terms of both execution time and power efficiency: (a) a baseline design using arrays, and (b) a design using bit-slice indexing (BSI) and distributed BSI arithmetic. We evaluate two types of queries popular in Big Data analytics: aggregations and top-k. These queries were implemented using each of the two data structure designs on Apache Spark running on a server cluster that was instrumented with specialized hardware for synchronized real-time power measurement for each server in the cluster. We performed a series of experiments running the above queries on several different datasets. These experiments show that the bit-slicing algorithm consistently outperforms the array algorithm in both power efficiency and execution time. An interesting observation is that the power efficiency improvement of the bit-slicing algorithm over the array method is comparable to or greater than the improvement in execution time for both queries evaluated. Gheorghi Guzun, Josiah McClurg, Guadalupe Canahuate, Raghuraman Mudumbai |
IEEE BigData | 3 |
| 2016 | On-demand aggregation of gridded data over user-specified spatio-temporal domainsabstractThe advent of satellite imagery, remote sensing products, and global scale numerical climate models over the last two decades has created an explosion of available gridded environmental data. These space-time explicit datasets are produced and distributed using different spatial and temporal resolutions. Current approaches for comparing two different products generally involve offline pre-computation of aggregations to a common spatio-temporal resolution. This limits the user's ability to interactively compare different data products or transform data products into the required input resolution for modeling. Joel E. Tosado, Gheorghi Guzun, Guadalupe Canahuate, Ricardo Mantilla |
SIGSPATIAL/GIS | 3 |
| 2016 | A Two-Phase MapReduce Algorithm for Scalable Preference Queries over High-Dimensional DataabstractPreference (top-k) queries play a key role in modern data analytics tasks. Top-k techniques rely on ranking functions in order to determine an overall score for each of the objects across all the relevant attributes being examined. This ranking function is provided by the user at query time, or generated for a particular user by a personalized search engine which prevents the pre-computation of the global scores. Executing this type of queries is particularly challenging for high-dimensional data. Recently, bit-sliced indices (BSI) were proposed to answer these high-dimensional preference queries efficiently in a centralized environment. Gheorghi Guzun, Guadalupe Canahuate, David Chiu 0001 |
IDEAS | 2 |
| 2016 | Performance evaluation of word-aligned compression methods for bitmap indices
Gheorghi Guzun, Guadalupe Canahuate |
Knowl. Inf. Syst. | 2 |
| 2016 | Hybrid query optimization for hard-to-compress bit-vectors
Gheorghi Guzun, Guadalupe Canahuate |
VLDB J. | 2 |
| 2015 | Scalable preference queries for high-dimensional data using map-reduceabstractPreference (top-k) queries play a key role in modern data analytics tasks. Top-k techniques rely on ranking functions in order to determine an overall score for each of the objects across all the relevant attributes being examined. This ranking function is provided by the user at query time, or generated for a particular user by a personalized search engine which prevents the pre-computation of the global scores. Executing this type of queries is particularly challenging for high-dimensional data. Recently, bit-sliced indices (BSI) were proposed to answer these preference queries efficiently in a non-distributed environment for data with hundreds of dimensions. As MapReduce and key-value stores proliferate as the preferred methods for analyzing big data, we set up to evaluate the performance of BSI in a distributed environment, in terms of index size, network traffic, and execution time of preference (top-k) queries, over data with thousands of dimensions. Indexing is implemented on top of Apache Spark for both column and row stores and shown to outperform Hive when running on Map-reduce, and Tez for top-k (preference) queries. Gheorghi Guzun, Joel E. Tosado, Guadalupe Canahuate |
IEEE BigData | 3 |
| 2014 | A tunable compression framework for bitmap indicesabstractBitmap indices are widely used for large read-only repositories in data warehouses and scientific databases. Their binary representation allows for the use of bitwise operations and specialized run-length compression techniques. Due to a trade-off between compression and query efficiency, bitmap compression schemes are aligned using a fixed encoding length size (typically the word length) to avoid explicit decompression during query time. In general, smaller encoding lengths provide better compression, but require more decoding during query execution. However, when the difference in size is considerable, it is possible for smaller encodings to also provide better execution time. We posit that a tailored encoding length for each bit vector will provide better performance than a one-size-fits-all approach. We present a framework that optimizes compression and query efficiency by allowing bitmaps to be compressed using variable encoding lengths while still maintaining alignment to avoid explicit decompression. Efficient algorithms are introduced to process queries over bitmaps compressed using different encoding lengths. An input parameter controls the aggressiveness of the compression providing the user with the ability to tune the tradeoff between space and query time. Our empirical study shows this approach achieves significant improvements in terms of both query time and compression ratio for synthetic and real data sets. Compared to 32-bit WAH, VAL-WAH produces up to 1.8× smaller bitmaps and achieves query times that are 30% faster. Gheorghi Guzun, Guadalupe Canahuate, David Chiu 0001, Jason Sawin |
ICDE | 2 |
| 2014 | Optimizing query execution for variable-aligned length compression of bitmap indicesabstractIndexing is a fundamental mechanism for efficient data access. Recently, we proposed the Variable-Aligned Length (VAL) bitmap index encoding framework, which generalizes the commonly used word-aligned compression techniques. VAL presented a variable-aligned compression framework, which allows columns of a bitmap to be compressed using different encoding lengths. This flexibility creates a tunable compression that balances the trade-off between space and query processing time. The variable format of VAL presents several unique opportunities for query optimization. Ryan Slechta, Jason Sawin, Ben McCamish, David Chiu 0001, Guadalupe Canahuate |
IDEAS | 5 |
| 2013 | Dynamic bitmap index recompression through workload-based optimizationsabstractMany large-scale read-only databases and data warehouses use bitmap indices in an effort to speed up data analysis. These indices have the dual properties of compressibility and being able to leverage fast bit-wise operations for query processing. Numerous hybrid run-length encoding compression schemes have been proposed that greatly compress the index and enable querying without the need to decompress. Typically, these schemes align their compression with the computer architecture's word size to further accelerate queries. Fredton Doan, David Chiu 0001, Brasil Perez Lukes, Jason Sawin, Gheorghi Guzun, Guadalupe Canahuate |
IDEAS | 6 |
| 2009 | Secondary bitmap indexes with vertical and horizontal partitioningabstractTraditional bitmap indexes are utilized as a special type of primary or clustered indexes where the queries are answered by performing fast logical operations supported by hardware. Answers are mapped to the physical data by using the row id of each tuple. Bitmaps represent the i-th tuple in the original table with the i-th bit position of the index. Run-length compression is used to reduce the size of the bitmaps and it has been shown that ordered data is significantly better compressed. However, for large-scale and dynamic datasets it is infeasible to keep the data always sorted. Partitioning can be used to keep the data in smaller and manageable chunks, where a different bitmap index is built for each chunk. We propose a novel bitmap index design with partitioning which serves as basis for non-clustered bitmap indexes. Individual bitmaps are not stored, only an Existence Bitmap (EB) for the existing ranks of the full table is maintained. This approach improves update performance of sorted bitmaps and does not require maintaining a heap as the underlying table, nor the same ordering for all the partitions. A one dimensional index is used over the ranks to map the bits in the EB to the physical order of the data, which allows queries to run even faster. The proposed approach, called ranked Non-Clustered Bitmaps (rNCB), is compared against traditional bitmaps using FastBit and shows significant performance gains. 1. Guadalupe Canahuate, Tan Apaydin, Ahmet Sacan, Hakan Ferhatosmanoglu |
EDBT | 1 |
| 2008 | Online Index Recommendations for High-Dimensional Databases Using Query WorkloadsabstractHigh-dimensional databases pose a challenge with respect to efficient access. High-dimensional indexes do not work because of the often-cited "curse of dimensionality." However, users are usually interested in querying data over a relatively small subset of the entire attribute set at a time. A potential solution is to use lower dimensional indexes that accurately represent the user access patterns. A query response using the physical database design that is developed based on a static snapshot of the query workload may significantly degrade if the query patterns change. To address these issues, we introduce a parameterizable technique to recommend indexes based on index types that are frequently used for high-dimensional data sets and to dynamically adjust indexes as the underlying query workload changes. We incorporate a query pattern change detection mechanism to determine when the access patterns have changed enough to warrant change in the physical database design. By adjusting analysis parameters, we trade off analysis speed against analysis resolution. We perform experiments with a number of data sets, query sets, and parameters to show the effect that varying these characteristics has on analysis results. Michael Gibas, Guadalupe Canahuate, Hakan Ferhatosmanoglu |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2007 | Update Conscious Bitmap IndicesabstractBitmap indices have been widely used in several domains such as data warehousing and scientific applications due to their efficiency in answering certain query types over large data sets. However, their utilization has been largely limited to read-only data sets or to static snapshots of data due to the cost associated with the update and append of new data. Typically, several bitmaps are associated with each indexed attribute in a table, i.e. one for each attribute value, bin, or range. Each one of these bitmaps needs to be updated to reflect a new, appended row. Since a given table could be represented by hundreds or even thousands of bitmaps, the insertion of a single record can be prohibitively costly. In order to transfer the fast query response times offered by bitmap indices to dynamic database domains, we propose an update conscious bitmap index that provides a mechanism to quickly update bitmaps to reflect dynamic database changes. For an insert operation only the bitmaps that represent the values being inserted need to be updated. We formalize the insert and delete operations of the proposed technique and provide a cost model for bitmap updates. We compare the update conscious bitmaps to traditional bitmaps in terms of storage space, update performance, and query execution time. Guadalupe Canahuate, Michael Gibas, Hakan Ferhatosmanoglu |
SSDBM | 1 |
| 2006 | Indexing Incomplete Databases
Guadalupe Canahuate, Michael Gibas, Hakan Ferhatosmanoglu |
EDBT | 1 |
| 2006 | Approximate Encoding for Direct Access and Query Processing over Compressed Bitmaps
Tan Apaydin, Guadalupe Canahuate, Hakan Ferhatosmanoglu, Ali Saman Tosun |
VLDB | 2 |
| 2006 | Efficient parallel processing of range queries through replicated declustering
Hakan Ferhatosmanoglu, Ali Saman Tosun, Guadalupe Canahuate, Aravind Ramachandran |
Distributed Parallel Databases | 3 |