EDBT 2026 Demo / reviewers in the wild / expert
Christopher Homan
dblp:h/VPless · also Christopher M. Homan, Christopher Michael Homan
· DBLP profile ↗
36ranked-venue papers
6as first author
15since 2021 · last 2026
0000-0003-1821-5125ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 20 · 13 since 2021Theory of computation · 9 · 5 first-authorGraphics, computer vision, multimedia, augmented reality and games · 7 · 6 since 2021Human-computer interaction and ubiquitous computing · 6 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 5Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Forest vs Tree: The (N, K) Trade-off in Reproducible ML EvaluationabstractReproducibility is a cornerstone of scientific validation and of the authority it confers on its results. Reproducibility in machine learning evaluations leads to greater trust, confidence, and value. However, the ground truth responses used in machine learning often necessarily come from humans, among whom disagreement is prevalent, and surprisingly little research has studied the impact of effectively ignoring disagreement in these responses, as is typically the case. One reason for the lack of research is that budgets for collecting human-annotated evaluation data are limited, and obtaining more samples from multiple raters for each example greatly increases the per-item annotation costs. We investigate the trade-off between the number of items (N) and the number of responses per item (K) needed for reliable machine learning evaluation. We analyze a diverse collection of categorical datasets for which multiple annotations per item exist, and simulated distributions fit to these datasets, to determine the optimal (N, K) configuration, given a fixed budget (N x K), for collecting evaluation data and reliably comparing the performance of machine learning models. Our findings show, first, that accounting for human disagreement may come with N x K at no more than 1000 (and often much lower) for every dataset tested on at least one metric. Moreover, this minimal N x K almost always occurred for K > 10. Furthermore, the nature of the tradeoff between K and N, or if one even existed, depends on the evaluation metric, with metrics that are more sensitive to the full distribution of responses performing better at higher levels of K. Our methods can be used to help ML practitioners get more effective test data by finding the optimal metrics and number of items and annotations per item to collect to get the most reliability for their budget. Deepak Pandita, Flip Korn, Christopher A. Welty, Christopher Homan |
AAAI | 4 |
| 2026 | ProRefine: Inference-Time Prompt Refinement with Textual Feedback (Student Abstract)abstractAgentic workflows, where multiple AI agents collaborate to accomplish complex tasks like reasoning or planning, play a substantial role in many cutting-edge commercial applications. These workflows depend critically on the prompts used to provide the roles models play in such workflows. Poorly designed prompts that fail even slightly to guide individual agents can lead to sub-optimal performance that may snowball within a system of agents, limiting their reliability and scalability. To address this important problem of inference-time prompt optimization, we introduce ProRefine, an innovative inference-time optimization method that uses an agentic loop of LLMs to generate and apply textual feedback. ProRefine dynamically refines prompts for multi-step reasoning tasks without additional training or ground truth labels. Evaluated on five benchmark mathematical reasoning datasets, ProRefine significantly surpasses zero-shot Chain-of-Thought baselines by 3 to 37 percentage points. This approach not only boosts accuracy but also allows smaller models to approach the performance of their larger counterparts. This highlights its potential for building cost-effective and powerful hybrid AI systems, thereby democratizing access to high-performing AI. Deepak Pandita, Tharindu Cyril Weerasooriya, Ankit Shah 0001, Isabelle Diana May-Xin Ng, Christopher Homan, Wei Wei 0019 |
AAAI | 5 |
| 2026 | SALAN: A Massive ASR Dataset for the Languages of Niger
Mamadou K. Keita, Christopher Homan, Emily Tucker Prud'hommeaux, Abdoulaye Sako, Seydou Diallo |
LREC | 2 |
| 2026 | Fostering Computer Science Theory Literacy: What makes a good conceptual model?abstractThe formulation of conceptual models is an important intermediate step in the mathematizing process. Conceptual modeling allows visualization of students' internal mental state as they work to express ideas in formal, mathematical terms. The objective of this Birds of a Feather session is to discuss, ''What makes a good conceptual model for understanding computer science theory?'' Students across all areas of computer science find the skills of computer science theory literacy, e.g. formal modeling and algorithms building, difficult to acquire. Building students' capacity to extensively build, critique and revise models, and doing so in groups, has great potential to help all CS majors achieve CST literacy and become valuable problem-solvers in their own subareas of computer science. In this BoF session we will examine examples of model-producing instructional prompts from a model-based unit in the course ''xTreme Theory'', part of an NSF-funded educational research project. The prompts were selected for breadth of accessibility regardless of CS area of expertise. We will also review the corresponding anonymized student submissions. Table groups will discuss similarities and differences in the student submissions and how they might assess the student-produced models for quality, then will contribute to a discussion of conceptual models and modeling in the teaching and learning of CSTheory. This session will support instructors who are interested in increasing students' capabilities for formal modeling and algorithms building and interested in connecting with colleagues in ways not afforded during the academic semester. Kimberly Fluet, Christopher Homan |
SIGCSE (2) | 2 |
| 2026 | Conceptual Models for Teaching and Learning Computer Science Theory
Kimberly Fluet, Lane A. Hemaspaandra, Christopher Homan |
SIGCSE (2) | 3 |
| 2025 | ARTICLE: Annotator Reliability Through In-Context LearningabstractEnsuring annotator quality in training and evaluation data is a key piece of machine learning in NLP. Tasks such as sentiment analysis and offensive speech detection are intrinsically subjective, creating a challenging scenario for traditional quality assessment approaches because it is hard to distinguish disagreement due to poor work from that due to differences of opinions between sincere annotators. With the goal of increasing diverse perspectives in annotation while ensuring consistency, we propose ARTICLE, an in-context learning (ICL) framework to estimate annotation quality through self-consistency. We evaluate this framework on two offensive speech datasets using multiple LLMs and compare its performance with traditional methods. Our findings indicate that ARTICLE can be used as a robust method for identifying reliable annotators, hence improving data quality. Sujan Dutta, Deepak Pandita, Tharindu Cyril Weerasooriya, Marcos Zampieri, Christopher Homan, Ashiqur R. KhudaBukhsh |
AAAI | 5 |
| 2025 | ARTICLE: Annotator Reliability Through In-Context Learning (Student Abstract)abstractEnsuring annotator quality in training and evaluation data is a key piece of machine learning in NLP. Tasks such as sentiment analysis and offensive speech detection are intrinsically subjective, creating a challenging scenario for traditional quality assessment approaches because it is hard to distinguish disagreement due to poor work from that due to differences of opinions between sincere annotators. With the goal of increasing diverse perspectives in annotation while ensuring consistency, we propose ARTICLE, an in-context learning (ICL) framework to estimate annotation quality through self-consistency. We evaluate this framework on two offensive speech datasets using multiple LLMs and compare its performance with traditional methods. Our findings indicate that ARTICLE can be used as a robust method for identifying reliable annotators, hence improving data quality. Sujan Dutta, Deepak Pandita, Tharindu Cyril Weerasooriya, Marcos Zampieri, Christopher Homan, Ashiqur R. KhudaBukhsh |
AAAI | 5 |
| 2025 | Hope vs. Hate: Understanding User Interactions with LGBTQ+ News Content in Mainstream US News Media through the Lens of Hope SpeechabstractThis paper makes three contributions.First, via a substantial corpus of 1,419,047 comments posted on 3,161 YouTube news videos of major US cable news outlets, we analyze how users engage with LGBTQ+ news content.Our analyses focus both on positive and negative content.In particular, we construct a hope speech classifier that detects positive (hope speech), negative, neutral, and irrelevant content.Second, in consultation with a public health expert specializing on LGBTQ+ health, we conduct an annotation study with a balanced and diverse political representation and release a dataset of 3,750 instances with crowd-sourced labels and detailed annotator demographic information.Finally, beyond providing a vital resource for the LGBTQ+ community, our annotation study and subsequent in-the-wild assessments reveal (1) strong association between rater political beliefs and how they rate content relevant to a marginalized community, (2) models trained on individual political beliefs exhibit considerable in-the-wild disagreement, and (3) zero-shot large language models (LLMs) align more with liberal raters.Trigger Warning: this paper contains offensive material that some may find upsetting.From crowdsourced storymapping projects sharing stories of love, loss, and a sense of belonging (Kirby et al., 2021) to safe, anonymous spaces to seek resources (McInroy et al., 2019) and dating platforms (Blackwell et al., 2015) -the internet and modern technologies play a positive role in the health and well-being of the LGBTQ+ community in various ways.However, cyberbullying (Abreu and Kenny, 2018), exposure to dehumanization through news media (Mendelsohn et al., 2020), and more recently, homophobic biases in large language models (LLMs) (Dutta et al., 2024a, 2025a) -* Ashiqur R. KhudaBukhsh is the corresponding author.are some of the modern technology perils the community grapples with. Jonathan Pofcher, Christopher Homan, Randall Sell, Ashiqur R. KhudaBukhsh |
EMNLP | 2 |
| 2025 | Bayelemabaga: Creating Resources for Bambara NLPabstractAllahsera Auguste Tapo, Kevin Assogba, Christopher M Homan, M. Mustafa Rafique, Marcos Zampieri. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Allahsera Tapo, Kevin Assogba, Christopher Homan, M. Mustafa Rafique, Marcos Zampieri |
NAACL (Long Papers) | 3 |
| 2024 | Leveraging Speech Data Diversity to Document Indigenous Heritage and Culture
Allahsera Tapo, Éric Le Ferrand, Zoey Liu, Christopher Homan, Emily Tucker Prud'hommeaux |
INTERSPEECH | 4 |
| 2024 | GRASP: A Disagreement Analysis Framework to Assess Group Associations in PerspectivesabstractVinodkumar Prabhakaran, Christopher Homan, Lora Aroyo, Aida Mostafazadeh Davani, Alicia Parrish, Alex Taylor, Mark Diaz, Ding Wang, Gregory Serapio-García. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Vinodkumar Prabhakaran, Christopher Homan, Lora Aroyo, Aida Mostafazadeh Davani, Alicia Parrish, Alex S. Taylor, Mark Diaz, Ding Wang 0006, Gregory Serapio-García |
NAACL-HLT | 2 |
| 2023 | Subjective Crowd Disagreements for Subjective Data: Uncovering Meaningful CrowdOpinion with Population-level LearningabstractTharindu Cyril Weerasooriya, Sarah Luger, Saloni Poddar, Ashiqur KhudaBukhsh, Christopher Homan. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Tharindu Cyril Weerasooriya, Sarah K. K. Luger, Saloni Poddar, Ashiqur R. KhudaBukhsh, Christopher Homan |
ACL (1) | 5 |
| 2023 | Vicarious Offense and Noise Audit of Offensive Speech Classifiers: Unifying Human and Machine Disagreement on What is OffensiveabstractTharindu Weerasooriya, Sujan Dutta, Tharindu Ranasinghe, Marcos Zampieri, Christopher Homan, Ashiqur KhudaBukhsh. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Tharindu Cyril Weerasooriya, Sujan Dutta, Tharindu Ranasinghe, Marcos Zampieri, Christopher Homan, Ashiqur R. KhudaBukhsh |
EMNLP | 5 |
| 2023 | DICES Dataset: Diversity in Conversational AI Evaluation for SafetyabstractMachine learning approaches often require training and evaluation datasets with a clear separation between positive and negative examples. This requirement overly simplifies the natural subjectivity present in many tasks, and obscures the inherent diversity in human perceptions and opinions about many content items. Preserving the variance in content and diversity in human perceptions in datasets is often quite expensive and laborious. This is especially troubling when building safety datasets for conversational AI systems, as safety is socio-culturally situated in this context. To demonstrate this crucial aspect of conversational AI safety, and to facilitate in-depth model performance analyses, we introduce the DICES (Diversity In Conversational AI Evaluation for Safety) dataset that contains fine-grained demographics information about raters, high replication of ratings per item to ensure statistical power for analyses, and encodes rater votes as distributions across different demographics to allow for in-depth explorations of different aggregation strategies. The DICES dataset enables the observation and measurement of variance, ambiguity, and diversity in the context of safety for conversational AI. We further describe a set of metrics that show how rater diversity influences safety perception across different geographic regions, ethnicity groups, age groups, and genders. The goal of the DICES dataset is to be used as a shared resource and benchmark that respects diverse perspectives during safety evaluation of conversational AI systems. Lora Aroyo, Alex S. Taylor, Mark Diaz, Christopher Homan, Alicia Parrish, Gregory Serapio-García, Vinodkumar Prabhakaran, Ding Wang 0006 |
NeurIPS | 4 |
| 2022 | Combining Simple but Novel Data Augmentation Methods for Improving Conformer ASR
Ronit Damania, Christopher Homan, Emily Tucker Prud'hommeaux |
INTERSPEECH | 2 |
| 2020 | Neighborhood-Based Pooling for Population-Level Label Distribution Learning
Tharindu Cyril Weerasooriya, Tong Liu 0010, Christopher Homan |
ECAI | 3 |
| 2019 | Learning to Predict Population-Level Label DistributionsabstractAs machine learning (ML) plays an ever increasing role in commerce, government, and daily life, reports of bias in ML systems against groups traditionally underrepresented in computing technologies have also increased. The problem appears to be extensive, yet it remains challenging even to fully assess the scope, let alone fix it. A fundamental reason is that ML systems are typically trained to predict one correct answer or set of answers; disagreements between the annotators who provide the training labels are resolved by either discarding minority opinions (which may correspond to demographic minorities or not) or presenting all opinions flatly, with no attempt to quantify how different answers might be distributed in society. Label distribution learning associates for each data item a probability distribution over the labels for that item. While such distributions may be representative of minority beliefs or not, they at least preserve diversities of opinion that conventional learning hides or ignores and represent a fundamental first step toward ML systems that can model diversity. We introduce a strategy for learning label distributions with only five-to-ten labels per item—a range that is typical of supervised learning datasets—by aggregating human-annotated labels over multiple, similarly rated data items. Our results suggest that specific label aggregation methods can help provide reliable, representative predictions at the population level. Tong Liu 0010, Akash Venkatachalam, Pratik Sanjay Bongale, Christopher Homan |
HCOMP | 4 |
| 2018 | Does Reciprocal Gratefulness in Twitter Predict Neighborhood Safety?: Comparing 911 Calls Where Users Reside or Use Social Media
Ann Marie White, Linxiao Bai, Christopher Homan, Melanie Funchess, Catherine Cerulli, Amen Ptah, Deepak Pandita, Henry A. Kautz |
ICWSM | 3 |
| 2017 | Modeling Information Sharing Behavior on Q&A Forums
Biru Cui, Shanchieh Jay Yang, Christopher Homan |
PAKDD (2) | 3 |
| 2016 | Understanding Discourse on Work and Job-Related Well-Being in Public Social MediaabstractWe construct a humans-in-the-loop supervised learning framework that integrates crowdsourcing feedback and local knowledge to detect job-related tweets from individual and business accounts. Using data-driven ethnography, we examine discourse about work by fusing language-based analysis with temporal, geospational, and labor statistics information. Tong Liu 0010, Christopher Homan, Cecilia O. Alm, Megan C. Lytle-Flint, Ann Marie White, Henry A. Kautz |
ACL (1) | 2 |
| 2016 | Analyzing Gender Bias in Student EvaluationsabstractUniversity students in the United States are routinely asked to provide feedback on the quality of the instruction they have received. Such feedback is widely used by university administrators to evaluate teaching ability, despite growing evidence that students assign lower numerical scores to women and people of color, regardless of the actual quality of instruction. In this paper, we analyze students’ written comments on faculty evaluation forms spanning eight years and five STEM disciplines in order to determine whether open-ended comments reflect these same biases. First, we apply sentiment analysis techniques to the corpus of comments to determine the overall affect of each comment. We then use this information, in combination with other features, to explore whether there is bias in how students describe their instructors. We show that while the gender of the evaluated instructor does not seem to affect students’ expressed level of overall satisfaction with their instruction, it does strongly influence the language that they use to describe their instructors and their experience in class. Andamlak Terkik, Emily Tucker Prud'hommeaux, Cecilia O. Alm, Christopher Homan, Scott Franklin 0001 |
COLING | 4 |
| 2015 | An Analysis of Domestic Abuse Discourse on RedditabstractDomestic abuse affects people of every race, class, age, and nation. There is sig-nificant research on the prevalence and ef-fects of domestic abuse; however, such re-search typically involves population-based surveys that have high financial costs. This work provides a qualitative analysis of do-mestic abuse using data collected from the social and news-aggregation website red-dit.com. We develop classifiers to detect submissions discussing domestic abuse, achieving accuracies of up to 92%, a sub-stantial error reduction over its baseline. Analysis of the top features used in detect-ing abuse discourse provides insight into the dynamics of abusive relationships. 1 Nicolas Schrading, Cecilia O. Alm, Raymond W. Ptucha, Christopher Homan |
EMNLP | 4 |
| 2015 | #WhyIStayed, #WhyILeft: Microblogging to Make Sense of Domestic AbuseabstractNicolas Schrading, Cecilia Ovesdotter Alm, Raymond Ptucha, Christopher Homan. Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2015. Nicolas Schrading, Cecilia O. Alm, Raymond W. Ptucha, Christopher Homan |
HLT-NAACL | 4 |
| 2015 | Dichotomy results for fixed point counting in boolean dynamical systems
Christopher Homan, Sven Kosub |
Theor. Comput. Sci. | 1 |
| 2014 | Non-independent Cascade Formation: Temporal and Spatial EffectsabstractDetermining cascade size and the factors affecting cascade size are two fundamental research problems in social network analysis. The commonly considered independent cascade model, when applied to social networks such as Digg, produces a phase-transition phenomenon where the cascade is either very small or very large. This phenomenon can be explained based on the concept of Giant Propagation Component (GPC). The GPC is defined as a maximally connected component, such that, by applying the independent cascade model, once any node of the component is infected, most of the remaining nodes in the component will eventually become infected with a high probability. While GPC exists in social networks, the phase-transition phenomenon, is not observed in the actual cascade size distribution when the information propagation is due to actions such as ``like'' or ``dig''. Biru Cui, Shanchieh Jay Yang, Christopher Homan |
CIKM | 3 |
| 2014 | Social structure and depression in TrevorSpaceabstractWe discover patterns related to depression in the social graph of an online community of approximately 20,000 lesbian, gay, and bisexual, transgender, and questioning youth. With survey data on fewer than two hundred community members and the network graph of the entire community (which is completely anonymous except for the survey responses), we detected statistically significant correlations between a number of graph properties and those TrevorSpace users showing a higher likelihood of depression, according to the Patient Healthcare Questionnaire-9, a standard instrument for estimating depression. Our results suggest that those who are less depressed are more deeply integrated into the social fabric of TrevorSpace than those who are more depressed. Our techniques may apply to other hard-to-reach online communities, like gay men on Facebook, where obtaining detailed information about individuals is difficult or expensive, but obtaining the social graph is not. Christopher Homan, Naiji Lu, Xin Tu, Megan C. Lytle-Flint, Vincent Silenzio |
CSCW | 1 |
| 2014 | Tuning the Diversity of Open-Ended Responses From the CrowdabstractCrowdsourcing can solve problems beyond the reach of state-of-the-art fully automated systems. A common pattern found in many such systems is for the workers to discover, in parallel, a number of candidate solutions and then vote on the best one to pass forward, often within a fixed amount of time. We present the propose-vote-abstain mechanism for eliciting from crowd workers the proper balance between solution discovery and selection. Each crowd worker is given a choice among proposing an answer, voting among the answers proposed so far, or abstaining, i.e., doing nothing. When a stopping condition is reached, the mechanism returns the answer with the most votes. Workers are paid a base amount, with bonuses if they propose or vote for the winning answer. Walter S. Lasecki, Christopher Homan, Jeffrey P. Bigham |
HCOMP | 2 |
| 2012 | On the approximability of Dodgson and Young elections
Ioannis Caragiannis, Jason A. Covey, Michal Feldman, Christopher Homan, Christos Kaklamanis, Nikos Karanikolas, Ariel D. Procaccia, Jeffrey S. Rosenschein |
Artif. Intell. | 4 |
| 2009 | On the approximability of Dodgson and Young electionsabstractThe voting rules proposed by Dodgson and Young are both designed to find the alternative closest to being a Condorcet winner, according to two different notions of proximity; the score of a given alternative is known to be hard to compute under either rule. In this paper, we put forward two algorithms for approximating the Dodgson score: an LP-based randomized rounding algorithm and a deterministic greedy algorithm, both of which yield an approximation ratio, where m is the number of alternatives; we observe that this result is asymptotically optimal, and further prove that our greedy algorithm is optimal up to a factor of 2, unless problems in have quasi-polynomial time algorithms. Although the greedy algorithm is computationally superior, we argue that the randomized rounding algorithm has an advantage from a social choice point of view. Further, we demonstrate that computing any reasonable approximation of the ranking produced by Dodgson's rule is -hard. This result provides a complexity-theoretic explanation of sharp discrepancies that have been observed in the Social Choice Theory literature when comparing Dodgson elections with simpler voting rules. Finally, we show that the problem of calculating the Young score is -hard to approximate by any factor. This leads to an inapproximability result for the Young ranking. Ioannis Caragiannis, Jason A. Covey, Michal Feldman, Christopher Homan, Christos Kaklamanis, Nikos Karanikolas, Ariel D. Procaccia, Jeffrey S. Rosenschein |
SODA | 4 |
| 2007 | Cluster computing and the power of edge recognition
Lane A. Hemaspaandra, Christopher Homan, Sven Kosub |
Inf. Comput. | 2 |
| 2007 | The Complexity of Computing the Size of an IntervalabstractGiven a p‐order A over a universe of strings (i.e., a transitive, reflexive, antisymmetric relation such that if $(x, y) \in A$, then $|x|$ is polynomially bounded by $|y|$), an interval size function of A returns, for each string x in the universe, the number of strings in the interval between strings $b(x)$ and $t(x)$ (with respect to A), where $b(x)$ and $t(x)$ are functions that are polynomial‐time computable in the length of x. By choosing sets of interval size functions based on feasibility requirements for their underlying p‐orders, we obtain new characterizations of complexity classes. We prove that the set of all interval size functions whose underlying p‐orders are polynomial‐time decidable is exactly #P. We show that the interval size functions for orders with polynomial‐time adjacency checks are closely related to the class FPSPACE(poly). Indeed, FPSPACE(poly) is exactly the class of all nonnegative functions that are an interval size function minus a polynomial‐time computable function. We study two important functions in relation to interval size functions. The function #DIV maps each natural number n to the number of nontrivial divisors of n. We show that #DIV is an interval size function of a polynomial‐time decidable partial p‐order with polynomial‐time adjacency checks. The function #MONSAT maps each monotone boolean formula F to the number of satisfying assignments of F. We show that #MONSAT is an interval size function of a polynomial‐time decidable total p‐order with polynomial‐time adjacency checks. Finally, we explore the related notion of cluster computation. Lane A. Hemaspaandra, Christopher Homan, Sven Kosub, Klaus W. Wagner |
SIAM J. Comput. | 2 |
| 2006 | Smoother Transitions Between Breadth-First-Spanning-Tree-Based Drawings
Christopher Homan, Andrew Pavlo, Jonathan Schull |
GD | 1 |
| 2006 | Guarantees for the Success Frequency of an Algorithm for Finding Dodgson-Election Winners
Christopher Homan, Lane A. Hemaspaandra |
MFCS | 1 |
| 2006 | Cluster Computing and the Power of Edge Recognition
Lane A. Hemaspaandra, Christopher Homan, Sven Kosub |
TAMC | 2 |
| 2004 | Tight lower bounds on the ambiguity of strong, total, associative, one-way functions
Christopher Homan |
J. Comput. Syst. Sci. | 1 |
| 2003 | One-way permutations and self-witnessing languages
Christopher Homan, Mayur Thakur |
J. Comput. Syst. Sci. | 1 |