Girmaw Abebe

dblp:182/8656 · also Girmaw Abebe Tadesse · DBLP profile ↗
← Back
18ranked-venue papers
8as first author
13since 2021 · last 2026
0000-0002-2648-9102ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 5 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 4 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 Machine Learning for Sustainable Rice Production: Region-Scale Monitoring of Water-Saving Practices in Punjab, India
abstract
Rice cultivation supplies half the world's population with staple food, while also being a major driver of freshwater depletion--consuming roughly a quarter of global freshwater--and accounting for ~48% of greenhouse gas emissions from croplands. In regions like Punjab, India, where groundwater levels are plummeting at 41.6 cm/year, adopting water-saving rice farming practices is critical. Direct-Seeded Rice (DSR) and Alternate Wetting and Drying (AWD) can cut irrigation water use by 20–40% without hurting yields, yet lack of spatial data on adoption impedes effective adaptation policy and climate action. We present a machine learning framework to bridge this data gap by monitoring sustainable rice farming at scale. In collaboration with agronomy experts and a large-scale farmer training program, we obtained ground-truth data from ~1,400 fields across Punjab. Leveraging this partnership, we developed a novel dimensional classification approach that decouples sowing and irrigation practices, achieving F1 scores of 0.8 and 0.74 respectively, solely employing Sentinel-1 satellite imagery. Explainability analysis reveals that DSR classification is robust while AWD classification depends primarily on planting schedule differences, as Sentinel-1's 12-day revisit frequency cannot capture the higher frequency irrigation cycles characteristic of AWD practices. Applying this model across 3 million fields reveals spatial heterogeneity in adoption at the state level, highlighting gaps and opportunities for policy targeting. Our district-level adoption rates correlate well with government estimates (Spearman's=0.69 and Rank Biased Overlap=0.77). This study provides policymakers and sustainability programs a powerful tool to track practice adoption, inform targeted interventions, and drive data-driven policies for water conservation and climate mitigation at regional scale.
Ando Shah, Rajveer Singh, Akram Zaytar, Girmaw Abebe, Caleb Robinson, Negar Tafti, Stephen A. Wood, Rahul Dodhia, Juan M. Lavista Ferres
AAAI4
2024 Weak Labeling for Cropland Mapping in Africa
abstract
Cropland mapping can play a vital role in addressing environmental, agricultural, and food security challenges. However, in the context of Africa, practical applications are often hindered by the limited availability of high-resolution cropland maps. Such maps typically require extensive human labeling, thereby creating a scalability bottleneck. To address this, we propose an approach that utilizes unsupervised object clustering to refine existing weak labels, such as those obtained from global cropland maps. The refined labels, in conjunction with sparse human annotations, serve as training data for a semantic segmentation network designed to identify cropland areas. We conduct experiments to demonstrate the benefits of the improved weak labels generated by our method. In a scenario where we train our model with only 33 human-annotated labels, the F1score for the cropland category increases from 0.53 to 0.84 when we add the mined negative labels.
Gilles Quentin Hacheme, Akram Zaytar, Girmaw Abebe, Caleb Robinson, Rahul Dodhia, Juan M. Lavista Ferres, Stephen A. Wood
IGARSS3
2023 Spatially Constrained Adversarial Attack Detection and Localization in the Representation Space of Optical Flow Networks
abstract
Optical flow estimation have shown significant improvements with advances in deep neural networks. However, these flow networks have recently been shown to be vulnerable to patch-based adversarial attacks, which poses security risks in real-world applications, such as self-driving cars and robotics. We propose SADL, a Spatially constrained adversarial Attack Detection and Localization framework, to detect and localize these patch-based attack without requiring a dedicated training. The detection of an attacked input sequence is performed via iterative optimization on the features from the inner layers of flow networks, without any prior knowledge of the attacks. The novel spatially constrained optimization ensures that the detected anomalous subset of features comes from a local region. To this end, SADL provides a subset of nodes within a spatial neighborhood that contribute more to the detection, which will be utilized to localize the attack in the input sequence. The proposed SADL is validated across multiple datasets and flow networks. With patch attacks 4.8% of the size of the input image resolution on RAFT, our method successfully detects and localizes them with an average precision of 0.946 and 0.951 for KITTI-2015 and MPI-Sintel datasets, respectively. The results show that SADL consistently achieves higher detection rates than existing methods and provides new localization capabilities.
Hannah Kim 0002, Celia Cintas, Girmaw Abebe, Skyler Speakman
IJCAI3
2023 Quantifying the impact of COVID-19 on essential health services: a comparison of interrupted time series analysis using Prophet and Poisson regression models
abstract
BACKGROUND: Coronavirus disease 2019 (COVID-19) altered healthcare utilization patterns. However, there is a dearth of literature comparing methods for quantifying the extent to which the pandemic disrupted healthcare service provision in sub-Saharan African countries. OBJECTIVE: To compare interrupted time series analysis using Prophet and Poisson regression models in evaluating the impact of COVID-19 on essential health services. METHODS: We used reported data from Uganda's Health Management Information System from February 2018 to December 2020. We compared Prophet and Poisson models in evaluating the impact of COVID-19 on new clinic visits, diabetes clinic visits, and in-hospital deliveries between March 2020 to December 2020 and across the Central, Eastern, Northern, and Western regions of Uganda. RESULTS: The models generated similar estimates of the impact of COVID-19 in 10 of the 12 outcome-region pairs evaluated. Both models estimated declines in new clinic visits in the Central, Northern, and Western regions, and an increase in the Eastern Region. Both models estimated declines in diabetes clinic visits in the Central and Western regions, with no significant changes in the Eastern and Northern regions. For in-hospital deliveries, the models estimated a decline in the Western Region, no changes in the Central Region, and had different estimates in the Eastern and Northern regions. CONCLUSIONS: The Prophet and Poisson models are useful in quantifying the impact of interruptions on essential health services during pandemics but may result in different measures of effect. Rigor and multimethod triangulation are necessary to study the true effect of pandemics on essential health services.
William Ogallo, Irene Wanyana, Girmaw Abebe, Catherine Wanjiru, Victor Akinwande, Steven Kabwama, Sekou L. Remy, Charles Wachira, Sharon Okwako, Susan Kizito, Rhoda Wanyenze, Suzanne Kiwanuka, Aisha Walcott-Bryant
J. Am. Medical Informatics Assoc.3
2022 Principled Subpopulation Analysis of the BetterBirth Study and the Impact of WHO's Safe Childbirth Checklist Intervention
Girmaw Abebe, Megan Marx Delaney, Victor Akinwande, William Ogallo, Claire-Helene Mershon, Katherine Semrau, Skyler Speakman
AMIA1
2022 Model-free feature selection to facilitate automatic discovery of divergent subgroups in tabular data
abstract
Data-centric AI encourages the need for cleaning, evaluating, and understanding data in order to achieve trustworthy AI. Existing technologies, such as AutoML, make it easier to design and train models automatically, but there is a lack of a similar level of capability to extract data-centric insights. Manual stratification of tabular data per a given feature of interest (e.g., gender) is limited to scaling up for higher feature dimension, which could be addressed using automatic discovery of divergent/anomalous subgroups. Nonetheless, these automatic discovery techniques often search across potentially exponential combinations of features which could be simplified using a preceding feature selection step. Existing feature selection techniques for tabular data often involve fitting a particular model (e.g., XGBoost) in order to select important features. However, such model-based selection is prone to model-bias and spurious correlations in addition to requiring extra resources to design, fine-tune and train a model. In this paper, we propose a model-free and sparsity-based automatic feature selection (SAFS) framework to facilitate automatic discovery of divergent subgroups. Different to filter-based selection techniques, we exploit the sparsity of objective measures among feature values to rank and select features. We validated SAFS across two publicly available datasets (MIMIC-III and Allstate Claims) and compared it with six existing feature selection methods. SAFS achieves a reduction of the feature selection time by a factor of 81× and 104×, averaged cross the existing methods in the MIMIC-III and Claims datasets, respectively. SAFS-selected features are also shown to achieve competitive detection performance, e.g., 18.3% of features selected by SAFS detected similar divergent group compared to using the whole features, in the Claims dataset, with a Jaccard similarity of 0.95 but with a 16× reduction in detection time.
Girmaw Abebe, William Ogallo, Celia Cintas, Skyler Speakman
IEEE Big Data1
2022 Systematic Discovery of Bias in Data
abstract
Detecting bias in data is an integral component of trustworthy and responsible ML. For researchers and data scientists, investigating, detecting, and becoming aware of biases present in data is an important step to correcting and making better ML decisions. Bias exists in the form of subsets that deviate from global expectations. Typically, researchers begin with a set of pre-defined protected/sensitive attributes and use them as the basis upon which deviation from expectation is examined. For instance, a researcher may examine under- or over-representation of a particular gender or race and adjust ML models accordingly. While this works for most settings, it is suboptimal, because it does not cover the true scale of all possible enumerations of subsets in the data. In this paper, we argue for a different approach to bias discovery. Instead of performing stratification across a pre-defined set of features, we ask the more open-ended question — which subset has the highest deviation between observed and expected outcomes? To answer this question, we leverage subset scanning, which efficiently maximizes measures of divergence over exponentially many combinations of feature values. We demonstrate the capabilities and advantages of subset scanning over pre-defined stratification by analyzing scanning results on the Stanford Open Policing dataset. In so doing, we uncover anomalous subsets within the data which, to the best of our knowledge, have not been discovered before and show that it is impossible to uncover such anomalies by stratifying across a set of pre-defined features.
John Wamburu, Girmaw Abebe, Celia Cintas, Adebayo Oshingbesan, Tanya Akumu, Skyler Speakman
IEEE Big Data2
2022 Towards Creativity Characterization of Generative Models via Group-Based Subset Scanning
abstract
Deep generative models, such as Variational Autoencoders (VAEs) and Generative Adversarial Networks (GANs), have been employed widely in computational creativity research. However, such models discourage out-of-distribution generation to avoid spurious sample generation, thereby limiting their creativity. Thus, incorporating research on human creativity into generative deep learning techniques presents an opportunity to make their outputs more compelling and human-like. As we see the emergence of generative models directed toward creativity research, a need for machine learning-based surrogate metrics to characterize creative output from these models is imperative. We propose group-based subset scanning to identify, quantify, and characterize creative processes by detecting a subset of anomalous node-activations in the hidden layers of the generative models. Our experiments on the standard image benchmarks and their ``creatively generated'' variants reveal that the proposed subset scores distribution is more useful for detecting novelty in creative processes in the activation space rather than the pixel space. Further, we found that creative samples generate larger subsets of anomalies than normal or non-creative samples across datasets. The node activations highlighted during the creative decoding process are different from those responsible for the normal sample generation. Lastly, we assess if the images from the subsets selected by our method were also found creative by human evaluators, presenting a link between creativity perception in humans and node activations within deep neural nets.
Celia Cintas, Brian Quanz, Girmaw Abebe, Skyler Speakman
IJCAI4
2022 Pattern detection in the activation space for identifying synthesized content
Celia Cintas, Skyler Speakman, Girmaw Abebe, Victor Akinwande, Edward McFowland, Komminist Weldemariam
Pattern Recognit. Lett.3
2021 Data-Driven Sequential Uptake Pattern Discovery for Family Planning Studies
Celia Cintas, Victor Akinwande, Ramya Raghavendra, Girmaw Abebe, Aisha Walcott-Bryant, Charity Wayua, Fredrick Makumbi, Rhoda Wanyenze, Komminist Weldemariam
AMIA4
2021 Automatic Stratification of Tabular Health Data
Skyler Speakman, Girmaw Abebe, Victor Akinwande, William Ogallo, Claire-Helene Mershon, Nosa Orobaton, Daniel B. Neill
AMIA2
2021 Racial Representation Analysis in Dermatology Academic Materials
Girmaw Abebe, Celia Cintas, Roxana Daneshjou, Kush R. Varshney, Peter W. J. Staar, Skyler Speakman, Kenya S. Andrews, Chinyere Agunwa, Justin Jia, Elizabeth E. Bailey, Jules Lipoff, Ginikanwa Onyekaba, Veronica Rotemberg, Ademide Adelekun, James Zou 0001
AMIA1
2021 DeepMI: Deep multi-lead ECG fusion for identifying myocardial infarction and its occurrence-time
Girmaw Abebe, Hamza A. Javed, Komminist Weldemariam, Jiyan Chen, Tingting Zhu 0001
Artif. Intell. Medicine1
2020 Discriminant Knowledge Extraction from Electrocardiograms for Automated Diagnosis of Myocardial Infarction
Girmaw Abebe, Komminist Weldemariam, Hamza A. Javed, Jiyan Chen, Tingting Zhu 0001
PKAW1
2020 PlethAugment: GAN-Based PPG Augmentation for Medical Diagnosis in Low-Resource Settings
abstract
The paucity of physiological time-series data collected from low-resource clinical settings limits the capabilities of modern machine learning algorithms in achieving high performance. Such performance is further hindered by class imbalance; datasets where a diagnosis is much more common than others. To overcome these two issues at low-cost while preserving privacy, data augmentation methods can be employed. In the time domain, the traditional method of time-warping could alter the underlying data distribution with detrimental consequences. This is prominent when dealing with physiological conditions that influence the frequency components of data. In this paper, we propose PlethAugment; three different conditional generative adversarial networks (CGANs) with an adapted diversity term for the generation of pathological photoplethysmogram (PPG) signals in order to boost medical classification performance. To evaluate and compare the GANs, we introduce a novel metric-agnostic method; the synthetic generalization curve. We validate this approach on two proprietary and two public datasets representing a diverse set of medical conditions. Compared to training on non-augmented class-balanced datasets, training on augmented datasets leads to an improvement of the AUROC by up to 29% when using cross validation. This illustrates the potential of the proposed CGANs to significantly improve classification performance.
Dani Kiyasseh, Girmaw Abebe, Nhan Le Nguyen Thanh, Le Van Tan, Louise Thwaites, Tingting Zhu 0001, David A. Clifton
IEEE J. Biomed. Health Informatics2
2020 Multi-Modal Diagnosis of Infectious Diseases in the Developing World
abstract
In low and middle income countries, infectious diseases continue to have a significant impact, particularly amongst the poorest in society. Tetanus and hand foot and mouth disease (HFMD) are two such diseases and, in both, death is associated with autonomic nervous system dysfunction (ANSD). Currently, photoplethysmogram or electrocardiogram monitoring is used to detect deterioration in these patients, however expensive clinical monitors are often required. In this study, we employ low-cost and mobile wearable devices to collect patient vital signs unobtrusively; and we develop machine learning algorithms for automatic and rapid triage of patients that provide efficient use of clinical resources. Existing methods are mainly dependent on the prior detection of clinical features with limited exploitation of multi-modal physiological data. Moreover, the latest developments in deep learning (e.g. cross-domain transfer learning) have not been sufficiently applied for infectious disease diagnosis. In this paper, we present a fusion of multi-modal physiological data to predict the severity of ANSD with a hierarchy of resource-aware decision making. First, an on-site triage process is performed using a simple classifier. Second, personalised longitudinal modelling is employed that takes the previous states of the patient into consideration. We have also employed a spectrogram representation of the physiological waveforms to exploit existing networks for cross-domain transfer learning, which avoids the laborious and data intensive process of training a network from scratch. Results show that the proposed framework has promising potential in supporting severity grading of infectious diseases in low-resources settings, such as in the developing world.
Girmaw Abebe, Hamza A. Javed, Nhan Le Nguyen Thanh, Ha Thi Hai Duong, Le Van Tan, Louise Thwaites, David A. Clifton, Tingting Zhu 0001
IEEE J. Biomed. Health Informatics1
2017 Hierarchical modeling for first-person vision activity recognition
Girmaw Abebe, Andrea Cavallaro
Neurocomputing1
2016 Robust multi-dimensional motion features for first-person vision activity recognition
abstract
We propose robust multi-dimensional motion features for human activity recognition from first-person videos. The proposed features encode information about motion magnitude, direction and variation, and combine them with virtual inertial data generated from the video itself. The use of grid flow representation, per-frame normalization and temporal feature accumulation enhances the robustness of our new representation. Results on multiple datasets demonstrate that the proposed feature representation outperforms existing motion features, and importantly it does so independently of the classifier. Moreover, the proposed multi-dimensional motion features are general enough to make them suitable for vision tasks beyond those related to wearable cameras.
Girmaw Abebe, Andrea Cavallaro, Xavier Parra Llanas
Comput. Vis. Image Underst.1