VLDB 2026 Research / reviewers in the wild / expert
Celia Cintas
dblp:199/3961
· DBLP profile ↗
14ranked-venue papers
5as first author
10since 2021 · last 2024
0000-0002-8064-9189ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 4 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Using Causal Inference to Investigate Contraceptive Discontinuation in Sub-Saharan Africa
Victor Akinwande, Megan MacGregor, Celia Cintas, Ehud Karavani, Dennis Wei, Kush R. Varshney, Pablo A. Nepomnaschy |
IJCAI | 3 |
| 2023 | Spatially Constrained Adversarial Attack Detection and Localization in the Representation Space of Optical Flow NetworksabstractOptical flow estimation have shown significant improvements with advances in deep neural networks. However, these flow networks have recently been shown to be vulnerable to patch-based adversarial attacks, which poses security risks in real-world applications, such as self-driving cars and robotics. We propose SADL, a Spatially constrained adversarial Attack Detection and Localization framework, to detect and localize these patch-based attack without requiring a dedicated training. The detection of an attacked input sequence is performed via iterative optimization on the features from the inner layers of flow networks, without any prior knowledge of the attacks. The novel spatially constrained optimization ensures that the detected anomalous subset of features comes from a local region. To this end, SADL provides a subset of nodes within a spatial neighborhood that contribute more to the detection, which will be utilized to localize the attack in the input sequence. The proposed SADL is validated across multiple datasets and flow networks. With patch attacks 4.8% of the size of the input image resolution on RAFT, our method successfully detects and localizes them with an average precision of 0.946 and 0.951 for KITTI-2015 and MPI-Sintel datasets, respectively. The results show that SADL consistently achieves higher detection rates than existing methods and provides new localization capabilities. Hannah Kim 0002, Celia Cintas, Girmaw Abebe, Skyler Speakman |
IJCAI | 2 |
| 2023 | IberianVoxel: Automatic Completion of Iberian Ceramics for Cultural Heritage StudiesabstractAccurate completion of archaeological artifacts is a critical aspect in several archaeological studies, including documentation of variations in style, inference of chronological and ethnic groups, and trading routes trends, among many others. However, most available pottery is fragmented, leading to missing textural and morphological cues. Currently, the reassembly and completion of fragmented ceramics is a daunting and time-consuming task, done almost exclusively by hand, which requires the physical manipulation of the fragments. To overcome the challenges of manual reconstruction, reduce the materials' exposure and deterioration, and improve the quality of reconstructed samples, we present IberianVoxel, a novel 3D Autoencoder Generative Adversarial Network (3D AE-GAN) framework tested on an extensive database with complete and fragmented references. We generated a collection of 1001 3D voxelized samples and their fragmented references from Iberian wheel-made pottery profiles. The fragments generated are stratified into different size groups and across multiple pottery classes. Lastly, we provide quantitative and qualitative assessments to measure the quality of the reconstructed voxelized samples by our proposed method and archaeologists' evaluation. Celia Cintas, Manuel J. Lucena, José Manuel Fuertes, Antonio J. Rueda Ruiz, Rafael Jesús Segura, Carlos J. Ogáyar, Rolando González-José, Claudio Delrieux |
IJCAI | 2 |
| 2022 | Model-free feature selection to facilitate automatic discovery of divergent subgroups in tabular dataabstractData-centric AI encourages the need for cleaning, evaluating, and understanding data in order to achieve trustworthy AI. Existing technologies, such as AutoML, make it easier to design and train models automatically, but there is a lack of a similar level of capability to extract data-centric insights. Manual stratification of tabular data per a given feature of interest (e.g., gender) is limited to scaling up for higher feature dimension, which could be addressed using automatic discovery of divergent/anomalous subgroups. Nonetheless, these automatic discovery techniques often search across potentially exponential combinations of features which could be simplified using a preceding feature selection step. Existing feature selection techniques for tabular data often involve fitting a particular model (e.g., XGBoost) in order to select important features. However, such model-based selection is prone to model-bias and spurious correlations in addition to requiring extra resources to design, fine-tune and train a model. In this paper, we propose a model-free and sparsity-based automatic feature selection (SAFS) framework to facilitate automatic discovery of divergent subgroups. Different to filter-based selection techniques, we exploit the sparsity of objective measures among feature values to rank and select features. We validated SAFS across two publicly available datasets (MIMIC-III and Allstate Claims) and compared it with six existing feature selection methods. SAFS achieves a reduction of the feature selection time by a factor of 81× and 104×, averaged cross the existing methods in the MIMIC-III and Claims datasets, respectively. SAFS-selected features are also shown to achieve competitive detection performance, e.g., 18.3% of features selected by SAFS detected similar divergent group compared to using the whole features, in the Claims dataset, with a Jaccard similarity of 0.95 but with a 16× reduction in detection time. Girmaw Abebe, William Ogallo, Celia Cintas, Skyler Speakman |
IEEE Big Data | 3 |
| 2022 | Systematic Discovery of Bias in DataabstractDetecting bias in data is an integral component of trustworthy and responsible ML. For researchers and data scientists, investigating, detecting, and becoming aware of biases present in data is an important step to correcting and making better ML decisions. Bias exists in the form of subsets that deviate from global expectations. Typically, researchers begin with a set of pre-defined protected/sensitive attributes and use them as the basis upon which deviation from expectation is examined. For instance, a researcher may examine under- or over-representation of a particular gender or race and adjust ML models accordingly. While this works for most settings, it is suboptimal, because it does not cover the true scale of all possible enumerations of subsets in the data. In this paper, we argue for a different approach to bias discovery. Instead of performing stratification across a pre-defined set of features, we ask the more open-ended question — which subset has the highest deviation between observed and expected outcomes? To answer this question, we leverage subset scanning, which efficiently maximizes measures of divergence over exponentially many combinations of feature values. We demonstrate the capabilities and advantages of subset scanning over pre-defined stratification by analyzing scanning results on the Stanford Open Policing dataset. In so doing, we uncover anomalous subsets within the data which, to the best of our knowledge, have not been discovered before and show that it is impossible to uncover such anomalies by stratifying across a set of pre-defined features. John Wamburu, Girmaw Abebe, Celia Cintas, Adebayo Oshingbesan, Tanya Akumu, Skyler Speakman |
IEEE Big Data | 3 |
| 2022 | Towards Creativity Characterization of Generative Models via Group-Based Subset ScanningabstractDeep generative models, such as Variational Autoencoders (VAEs) and Generative Adversarial Networks (GANs), have been employed widely in computational creativity research. However, such models discourage out-of-distribution generation to avoid spurious sample generation, thereby limiting their creativity. Thus, incorporating research on human creativity into generative deep learning techniques presents an opportunity to make their outputs more compelling and human-like. As we see the emergence of generative models directed toward creativity research, a need for machine learning-based surrogate metrics to characterize creative output from these models is imperative. We propose group-based subset scanning to identify, quantify, and characterize creative processes by detecting a subset of anomalous node-activations in the hidden layers of the generative models. Our experiments on the standard image benchmarks and their ``creatively generated'' variants reveal that the proposed subset scores distribution is more useful for detecting novelty in creative processes in the activation space rather than the pixel space. Further, we found that creative samples generate larger subsets of anomalies than normal or non-creative samples across datasets. The node activations highlighted during the creative decoding process are different from those responsible for the normal sample generation. Lastly, we assess if the images from the subsets selected by our method were also found creative by human evaluators, presenting a link between creativity perception in humans and node activations within deep neural nets. Celia Cintas, Brian Quanz, Girmaw Abebe, Skyler Speakman |
IJCAI | 1 |
| 2022 | Pattern detection in the activation space for identifying synthesized content
Celia Cintas, Skyler Speakman, Girmaw Abebe, Victor Akinwande, Edward McFowland, Komminist Weldemariam |
Pattern Recognit. Lett. | 1 |
| 2021 | Towards effect estimation of COVID-19 Non-pharmaceutical Interventions
Vesna Resende Barros, Victor Akinwande, Itay Manes, Osnat Bar-Shira, Celia Cintas, Yishai Shimoni, Michal Rosen-Zvi |
AMIA | 5 |
| 2021 | Data-Driven Sequential Uptake Pattern Discovery for Family Planning Studies
Celia Cintas, Victor Akinwande, Ramya Raghavendra, Girmaw Abebe, Aisha Walcott-Bryant, Charity Wayua, Fredrick Makumbi, Rhoda Wanyenze, Komminist Weldemariam |
AMIA | 1 |
| 2021 | Racial Representation Analysis in Dermatology Academic Materials
Girmaw Abebe, Celia Cintas, Roxana Daneshjou, Kush R. Varshney, Peter W. J. Staar, Skyler Speakman, Kenya S. Andrews, Chinyere Agunwa, Justin Jia, Elizabeth E. Bailey, Jules Lipoff, Ginikanwa Onyekaba, Veronica Rotemberg, Ademide Adelekun, James Zou 0001 |
AMIA | 2 |
| 2020 | Preservation of Anomalous Subgroups On Variational Autoencoder Transformed DataabstractWe investigate the effect of variational autoencoder (VAE) based data anonymization and its ability to preserve anomalous subgroup properties. We present a Utility Guaranteed Deep Privacy (UGDP) system which casts existing anomalous pattern detection methods as a new utility measure for data synthesis. UGDP's approach shows that properties of an anomalous subset of records, identified in the original data set, are preserved through the anonymization of a VAE. This is despite the newly generated records being completely synthetic. More specifically, the Bias-Scan algorithm identifies a subgroup of records that are consistently over- (or under-) risked by a black-box classifier as an area of 'poor fit'. This scanning process is applied on both pre- and post- VAE synthesized data. The areas of poor fit (i.e. anomalous records) persist in both settings. We evaluate our approach using publicly available datasets from the financial industry. Our evaluation confirmed that the approach is able to produce synthetic datasets that preserved a high level of subgroup differentiation as identified initially in the original dataset. Such a distinction was maintained while having distinctly different records between the synthetic and original dataset. Samuel C. Maina, Reginald E. Bryant, William Ogallo, Kush R. Varshney, Skyler Speakman, Celia Cintas, Aisha Walcott-Bryant, Robert-Florian Samoilescu, Komminist Weldemariam |
ICASSP | 6 |
| 2020 | Decision Platform for Pattern Discovery and Causal Effect Estimation in Contraceptive DiscontinuationabstractContraceptive use improves the health of women and children in several ways, yet data shows high rates of discontinuation which is not well understood. We introduce an AI-based decision platform capable of analyzing event data to identify patterns of contraceptive uptake that are unique to a subpopulation of interest. These discriminatory patterns provide valuable, interpretable insights to policy-makers. The sequences then serve as a hypothesis for downstream causal analysis to estimate the effect of specific variables on discontinuation outcomes. Our platform presents a way to visualize, stratify, compare, and perform a causal analysis on covariates that determine contraceptive uptake behavior, and yet is general enough to be extended to a variety of applications. Celia Cintas, Ramya Raghavendra, Victor Akinwande, Aisha Walcott-Bryant, Charity Wayua, Komminist Weldemariam |
IJCAI | 1 |
| 2020 | Detecting Adversarial Attacks via Subset Scanning of Autoencoder Activations and Reconstruction ErrorabstractReliably detecting attacks in a given set of inputs is of high practical relevance because of the vulnerability of neural networks to adversarial examples. These altered inputs create a security risk in applications with real-world consequences, such as self-driving cars, robotics and financial services. We propose an unsupervised method for detecting adversarial attacks in inner layers of autoencoder (AE) networks by maximizing a non-parametric measure of anomalous node activations. Previous work in this space has shown AE networks can detect anomalous images by thresholding the reconstruction error produced by the final layer. Furthermore, other detection methods rely on data augmentation or specialized training techniques which must be asserted before training time. In contrast, we use subset scanning methods from the anomalous pattern detection domain to enhance detection power without labeled examples of the noise, retraining or data augmentation methods. In addition to an anomalous “score” our proposed method also returns the subset of nodes within the AE network that contributed to that score. This will allow future work to pivot from detection to visualisation and explainability. Our scanning approach shows consistently higher detection power than existing detection methods across several adversarial noise models and a wide range of perturbation strengths. Celia Cintas, Skyler Speakman, Victor Akinwande, William Ogallo, Komminist Weldemariam, Srihari Sridharan, Edward McFowland |
IJCAI | 1 |
| 2020 | Fairness of Classifiers Across Skin Tones in Dermatology
Newton M. Kinyanjui, Timothy Odonga, Celia Cintas, Noel Codella, Rameswar Panda, Prasanna Sattigeri, Kush R. Varshney |
MICCAI (6) | 3 |