VLDB 2026 Research / reviewers in the wild / expert
Girmaw Abebe
dblp:182/8656 · also Girmaw Abebe Tadesse
· DBLP profile ↗
18ranked-venue papers
8as first author
13since 2021 · last 2026
0000-0002-2648-9102ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 5 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 4 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Machine Learning for Sustainable Rice Production: Region-Scale Monitoring of Water-Saving Practices in Punjab, IndiaabstractRice cultivation supplies half the world's population with staple food, while also being a major driver of freshwater depletion--consuming roughly a quarter of global freshwater--and accounting for ~48% of greenhouse gas emissions from croplands. In regions like Punjab, India, where groundwater levels are plummeting at 41.6 cm/year, adopting water-saving rice farming practices is critical. Direct-Seeded Rice (DSR) and Alternate Wetting and Drying (AWD) can cut irrigation water use by 20–40% without hurting yields, yet lack of spatial data on adoption impedes effective adaptation policy and climate action. We present a machine learning framework to bridge this data gap by monitoring sustainable rice farming at scale. In collaboration with agronomy experts and a large-scale farmer training program, we obtained ground-truth data from ~1,400 fields across Punjab. Leveraging this partnership, we developed a novel dimensional classification approach that decouples sowing and irrigation practices, achieving F1 scores of 0.8 and 0.74 respectively, solely employing Sentinel-1 satellite imagery. Explainability analysis reveals that DSR classification is robust while AWD classification depends primarily on planting schedule differences, as Sentinel-1's 12-day revisit frequency cannot capture the higher frequency irrigation cycles characteristic of AWD practices. Applying this model across 3 million fields reveals spatial heterogeneity in adoption at the state level, highlighting gaps and opportunities for policy targeting. Our district-level adoption rates correlate well with government estimates (Spearman's=0.69 and Rank Biased Overlap=0.77). This study provides policymakers and sustainability programs a powerful tool to track practice adoption, inform targeted interventions, and drive data-driven policies for water conservation and climate mitigation at regional scale. Ando Shah, Rajveer Singh, Akram Zaytar, Girmaw Abebe, Caleb Robinson, Negar Tafti, Stephen A. Wood, Rahul Dodhia, Juan M. Lavista Ferres |
AAAI | 4 |
| 2024 | Weak Labeling for Cropland Mapping in AfricaabstractCropland mapping can play a vital role in addressing environmental, agricultural, and food security challenges. However, in the context of Africa, practical applications are often hindered by the limited availability of high-resolution cropland maps. Such maps typically require extensive human labeling, thereby creating a scalability bottleneck. To address this, we propose an approach that utilizes unsupervised object clustering to refine existing weak labels, such as those obtained from global cropland maps. The refined labels, in conjunction with sparse human annotations, serve as training data for a semantic segmentation network designed to identify cropland areas. We conduct experiments to demonstrate the benefits of the improved weak labels generated by our method. In a scenario where we train our model with only 33 human-annotated labels, the F1score for the cropland category increases from 0.53 to 0.84 when we add the mined negative labels. Gilles Quentin Hacheme, Akram Zaytar, Girmaw Abebe, Caleb Robinson, Rahul Dodhia, Juan M. Lavista Ferres, Stephen A. Wood |
IGARSS | 3 |
| 2023 | Spatially Constrained Adversarial Attack Detection and Localization in the Representation Space of Optical Flow NetworksabstractOptical flow estimation have shown significant improvements with advances in deep neural networks. However, these flow networks have recently been shown to be vulnerable to patch-based adversarial attacks, which poses security risks in real-world applications, such as self-driving cars and robotics. We propose SADL, a Spatially constrained adversarial Attack Detection and Localization framework, to detect and localize these patch-based attack without requiring a dedicated training. The detection of an attacked input sequence is performed via iterative optimization on the features from the inner layers of flow networks, without any prior knowledge of the attacks. The novel spatially constrained optimization ensures that the detected anomalous subset of features comes from a local region. To this end, SADL provides a subset of nodes within a spatial neighborhood that contribute more to the detection, which will be utilized to localize the attack in the input sequence. The proposed SADL is validated across multiple datasets and flow networks. With patch attacks 4.8% of the size of the input image resolution on RAFT, our method successfully detects and localizes them with an average precision of 0.946 and 0.951 for KITTI-2015 and MPI-Sintel datasets, respectively. The results show that SADL consistently achieves higher detection rates than existing methods and provides new localization capabilities. Hannah Kim 0002, Celia Cintas, Girmaw Abebe, Skyler Speakman |
IJCAI | 3 |
| 2023 | Quantifying the impact of COVID-19 on essential health services: a comparison of interrupted time series analysis using Prophet and Poisson regression modelsabstractBACKGROUND: Coronavirus disease 2019 (COVID-19) altered healthcare utilization patterns. However, there is a dearth of literature comparing methods for quantifying the extent to which the pandemic disrupted healthcare service provision in sub-Saharan African countries. OBJECTIVE: To compare interrupted time series analysis using Prophet and Poisson regression models in evaluating the impact of COVID-19 on essential health services. METHODS: We used reported data from Uganda's Health Management Information System from February 2018 to December 2020. We compared Prophet and Poisson models in evaluating the impact of COVID-19 on new clinic visits, diabetes clinic visits, and in-hospital deliveries between March 2020 to December 2020 and across the Central, Eastern, Northern, and Western regions of Uganda. RESULTS: The models generated similar estimates of the impact of COVID-19 in 10 of the 12 outcome-region pairs evaluated. Both models estimated declines in new clinic visits in the Central, Northern, and Western regions, and an increase in the Eastern Region. Both models estimated declines in diabetes clinic visits in the Central and Western regions, with no significant changes in the Eastern and Northern regions. For in-hospital deliveries, the models estimated a decline in the Western Region, no changes in the Central Region, and had different estimates in the Eastern and Northern regions. CONCLUSIONS: The Prophet and Poisson models are useful in quantifying the impact of interruptions on essential health services during pandemics but may result in different measures of effect. Rigor and multimethod triangulation are necessary to study the true effect of pandemics on essential health services. William Ogallo, Irene Wanyana, Girmaw Abebe, Catherine Wanjiru, Victor Akinwande, Steven Kabwama, Sekou L. Remy, Charles Wachira, Sharon Okwako, Susan Kizito, Rhoda Wanyenze, Suzanne Kiwanuka, Aisha Walcott-Bryant |
J. Am. Medical Informatics Assoc. | 3 |
| 2022 | Principled Subpopulation Analysis of the BetterBirth Study and the Impact of WHO's Safe Childbirth Checklist Intervention
Girmaw Abebe, Megan Marx Delaney, Victor Akinwande, William Ogallo, Claire-Helene Mershon, Katherine Semrau, Skyler Speakman |
AMIA | 1 |
| 2022 | Model-free feature selection to facilitate automatic discovery of divergent subgroups in tabular dataabstractData-centric AI encourages the need for cleaning, evaluating, and understanding data in order to achieve trustworthy AI. Existing technologies, such as AutoML, make it easier to design and train models automatically, but there is a lack of a similar level of capability to extract data-centric insights. Manual stratification of tabular data per a given feature of interest (e.g., gender) is limited to scaling up for higher feature dimension, which could be addressed using automatic discovery of divergent/anomalous subgroups. Nonetheless, these automatic discovery techniques often search across potentially exponential combinations of features which could be simplified using a preceding feature selection step. Existing feature selection techniques for tabular data often involve fitting a particular model (e.g., XGBoost) in order to select important features. However, such model-based selection is prone to model-bias and spurious correlations in addition to requiring extra resources to design, fine-tune and train a model. In this paper, we propose a model-free and sparsity-based automatic feature selection (SAFS) framework to facilitate automatic discovery of divergent subgroups. Different to filter-based selection techniques, we exploit the sparsity of objective measures among feature values to rank and select features. We validated SAFS across two publicly available datasets (MIMIC-III and Allstate Claims) and compared it with six existing feature selection methods. SAFS achieves a reduction of the feature selection time by a factor of 81× and 104×, averaged cross the existing methods in the MIMIC-III and Claims datasets, respectively. SAFS-selected features are also shown to achieve competitive detection performance, e.g., 18.3% of features selected by SAFS detected similar divergent group compared to using the whole features, in the Claims dataset, with a Jaccard similarity of 0.95 but with a 16× reduction in detection time. Girmaw Abebe, William Ogallo, Celia Cintas, Skyler Speakman |
IEEE Big Data | 1 |
| 2022 | Systematic Discovery of Bias in DataabstractDetecting bias in data is an integral component of trustworthy and responsible ML. For researchers and data scientists, investigating, detecting, and becoming aware of biases present in data is an important step to correcting and making better ML decisions. Bias exists in the form of subsets that deviate from global expectations. Typically, researchers begin with a set of pre-defined protected/sensitive attributes and use them as the basis upon which deviation from expectation is examined. For instance, a researcher may examine under- or over-representation of a particular gender or race and adjust ML models accordingly. While this works for most settings, it is suboptimal, because it does not cover the true scale of all possible enumerations of subsets in the data. In this paper, we argue for a different approach to bias discovery. Instead of performing stratification across a pre-defined set of features, we ask the more open-ended question — which subset has the highest deviation between observed and expected outcomes? To answer this question, we leverage subset scanning, which efficiently maximizes measures of divergence over exponentially many combinations of feature values. We demonstrate the capabilities and advantages of subset scanning over pre-defined stratification by analyzing scanning results on the Stanford Open Policing dataset. In so doing, we uncover anomalous subsets within the data which, to the best of our knowledge, have not been discovered before and show that it is impossible to uncover such anomalies by stratifying across a set of pre-defined features. John Wamburu, Girmaw Abebe, Celia Cintas, Adebayo Oshingbesan, Tanya Akumu, Skyler Speakman |
IEEE Big Data | 2 |
| 2022 | Towards Creativity Characterization of Generative Models via Group-Based Subset ScanningabstractDeep generative models, such as Variational Autoencoders (VAEs) and Generative Adversarial Networks (GANs), have been employed widely in computational creativity research. However, such models discourage out-of-distribution generation to avoid spurious sample generation, thereby limiting their creativity. Thus, incorporating research on human creativity into generative deep learning techniques presents an opportunity to make their outputs more compelling and human-like. As we see the emergence of generative models directed toward creativity research, a need for machine learning-based surrogate metrics to characterize creative output from these models is imperative. We propose group-based subset scanning to identify, quantify, and characterize creative processes by detecting a subset of anomalous node-activations in the hidden layers of the generative models. Our experiments on the standard image benchmarks and their ``creatively generated'' variants reveal that the proposed subset scores distribution is more useful for detecting novelty in creative processes in the activation space rather than the pixel space. Further, we found that creative samples generate larger subsets of anomalies than normal or non-creative samples across datasets. The node activations highlighted during the creative decoding process are different from those responsible for the normal sample generation. Lastly, we assess if the images from the subsets selected by our method were also found creative by human evaluators, presenting a link between creativity perception in humans and node activations within deep neural nets. Celia Cintas, Brian Quanz, Girmaw Abebe, Skyler Speakman |
IJCAI | 4 |
| 2022 | Pattern detection in the activation space for identifying synthesized content
Celia Cintas, Skyler Speakman, Girmaw Abebe, Victor Akinwande, Edward McFowland, Komminist Weldemariam |
Pattern Recognit. Lett. | 3 |
| 2021 | Data-Driven Sequential Uptake Pattern Discovery for Family Planning Studies
Celia Cintas, Victor Akinwande, Ramya Raghavendra, Girmaw Abebe, Aisha Walcott-Bryant, Charity Wayua, Fredrick Makumbi, Rhoda Wanyenze, Komminist Weldemariam |
AMIA | 4 |
| 2021 | Automatic Stratification of Tabular Health Data
Skyler Speakman, Girmaw Abebe, Victor Akinwande, William Ogallo, Claire-Helene Mershon, Nosa Orobaton, Daniel B. Neill |
AMIA | 2 |
| 2021 | Racial Representation Analysis in Dermatology Academic Materials
Girmaw Abebe, Celia Cintas, Roxana Daneshjou, Kush R. Varshney, Peter W. J. Staar, Skyler Speakman, Kenya S. Andrews, Chinyere Agunwa, Justin Jia, Elizabeth E. Bailey, Jules Lipoff, Ginikanwa Onyekaba, Veronica Rotemberg, Ademide Adelekun, James Zou 0001 |
AMIA | 1 |
| 2021 | DeepMI: Deep multi-lead ECG fusion for identifying myocardial infarction and its occurrence-time
Girmaw Abebe, Hamza A. Javed, Komminist Weldemariam, Jiyan Chen, Tingting Zhu 0001 |
Artif. Intell. Medicine | 1 |
| 2020 | Discriminant Knowledge Extraction from Electrocardiograms for Automated Diagnosis of Myocardial Infarction
Girmaw Abebe, Komminist Weldemariam, Hamza A. Javed, Jiyan Chen, Tingting Zhu 0001 |
PKAW | 1 |
| 2020 | PlethAugment: GAN-Based PPG Augmentation for Medical Diagnosis in Low-Resource SettingsabstractThe paucity of physiological time-series data collected from low-resource clinical settings limits the capabilities of modern machine learning algorithms in achieving high performance. Such performance is further hindered by class imbalance; datasets where a diagnosis is much more common than others. To overcome these two issues at low-cost while preserving privacy, data augmentation methods can be employed. In the time domain, the traditional method of time-warping could alter the underlying data distribution with detrimental consequences. This is prominent when dealing with physiological conditions that influence the frequency components of data. In this paper, we propose PlethAugment; three different conditional generative adversarial networks (CGANs) with an adapted diversity term for the generation of pathological photoplethysmogram (PPG) signals in order to boost medical classification performance. To evaluate and compare the GANs, we introduce a novel metric-agnostic method; the synthetic generalization curve. We validate this approach on two proprietary and two public datasets representing a diverse set of medical conditions. Compared to training on non-augmented class-balanced datasets, training on augmented datasets leads to an improvement of the AUROC by up to 29% when using cross validation. This illustrates the potential of the proposed CGANs to significantly improve classification performance. Dani Kiyasseh, Girmaw Abebe, Nhan Le Nguyen Thanh, Le Van Tan, Louise Thwaites, Tingting Zhu 0001, David A. Clifton |
IEEE J. Biomed. Health Informatics | 2 |
| 2020 | Multi-Modal Diagnosis of Infectious Diseases in the Developing WorldabstractIn low and middle income countries, infectious diseases continue to have a significant impact, particularly amongst the poorest in society. Tetanus and hand foot and mouth disease (HFMD) are two such diseases and, in both, death is associated with autonomic nervous system dysfunction (ANSD). Currently, photoplethysmogram or electrocardiogram monitoring is used to detect deterioration in these patients, however expensive clinical monitors are often required. In this study, we employ low-cost and mobile wearable devices to collect patient vital signs unobtrusively; and we develop machine learning algorithms for automatic and rapid triage of patients that provide efficient use of clinical resources. Existing methods are mainly dependent on the prior detection of clinical features with limited exploitation of multi-modal physiological data. Moreover, the latest developments in deep learning (e.g. cross-domain transfer learning) have not been sufficiently applied for infectious disease diagnosis. In this paper, we present a fusion of multi-modal physiological data to predict the severity of ANSD with a hierarchy of resource-aware decision making. First, an on-site triage process is performed using a simple classifier. Second, personalised longitudinal modelling is employed that takes the previous states of the patient into consideration. We have also employed a spectrogram representation of the physiological waveforms to exploit existing networks for cross-domain transfer learning, which avoids the laborious and data intensive process of training a network from scratch. Results show that the proposed framework has promising potential in supporting severity grading of infectious diseases in low-resources settings, such as in the developing world. Girmaw Abebe, Hamza A. Javed, Nhan Le Nguyen Thanh, Ha Thi Hai Duong, Le Van Tan, Louise Thwaites, David A. Clifton, Tingting Zhu 0001 |
IEEE J. Biomed. Health Informatics | 1 |
| 2017 | Hierarchical modeling for first-person vision activity recognition
Girmaw Abebe, Andrea Cavallaro |
Neurocomputing | 1 |
| 2016 | Robust multi-dimensional motion features for first-person vision activity recognitionabstractWe propose robust multi-dimensional motion features for human activity recognition from first-person videos. The proposed features encode information about motion magnitude, direction and variation, and combine them with virtual inertial data generated from the video itself. The use of grid flow representation, per-frame normalization and temporal feature accumulation enhances the robustness of our new representation. Results on multiple datasets demonstrate that the proposed feature representation outperforms existing motion features, and importantly it does so independently of the classifier. Moreover, the proposed multi-dimensional motion features are general enough to make them suitable for vision tasks beyond those related to wearable cameras. Girmaw Abebe, Andrea Cavallaro, Xavier Parra Llanas |
Comput. Vis. Image Underst. | 1 |