EDBT 2026 Demo / reviewers in the wild / expert
Anthony Rios
dblp:133/1827
· DBLP profile ↗
24ranked-venue papers
7as first author
13since 2021 · last 2026
0000-0003-1781-3975ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 5 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 2 first-author · 4 since 2021Security and privacy · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Rethinking Access-Control Policy Authoring as a Multimodal Challenge [BlueSky Paper]abstractAccess control policies are rarely authored directly as machine-enforceable specifications. Instead, policy intent is developed through meetings in which requirements are communicated through spoken discussion, organizational charts, and policy diagrams. Although recent advances in natural language processing have improved rule extraction from text, current policy engineering pipelines largely ignore visual artifacts, leaving substantial policy information unstructured and difficult to translate into enforceable form. Sherifdeen Lawal, Xingmeng Zhao, Anthony Rios, Ram Krishnan |
SACMAT | 4 |
| 2026 | TRACER: Early Failure Detection for Task-Oriented DialogueabstractTask-oriented dialogue systems often fail before the final breakdown is obvious, but most evaluation only measures failure after the conversation has already gone wrong. We present TRACER, a method for early failure detection in task-oriented dialogue. TRACER predicts from a partial dialogue whether the full conversation will eventually fail by combining simple trajectory signals from belief-state changes with text representations of the evolving dialogue state. We evaluate the method in both oracle and generated belief-state settings, and test how well it works when only 25%, 50%, 75%, or 100% of the dialogue is visible. Across these settings, TRACER detects useful failure signals well before the end of the conversation and outperforms heuristic, classical, and single-stream baselines. These results suggest that early failure detection can provide a practical warning signal for dialogue systems before the interaction fully breaks down. Source code can be found here: https://github.com/erfan-nourbakhsh/TRACER. Erfan Nourbakhsh, Rocky Slavin, Ke Yang 0003, Anthony Rios |
SIGDIAL | 4 |
| 2025 | A Multi-Agent Framework for Mitigating Dialect Biases in Privacy Policy Question-Answering SystemsabstractĐorđe Klisura, Astrid R Bernaga Torres, Anna Karen Gárate-Escamilla, Rajesh Roshan Biswal, Ke Yang, Hilal Pataci, Anthony Rios. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Dorde Klisura, Astrid R. Bernaga Torres, Anna Karen Gárate-Escamilla, Rajesh Roshan Biswal, Ke Yang 0003, Hilal Pataci, Anthony Rios |
ACL (1) | 7 |
| 2025 | Charting the Future: Using Chart Question-Answering for Scalable Evaluation of LLM-Driven Data VisualizationsabstractWe propose a novel framework that leverages Visual Question Answering (VQA) models to automate the evaluation of LLM-generated data visualizations. Traditional evaluation methods often rely on human judgment, which is costly and unscalable, or focus solely on data accuracy, neglecting the effectiveness of visual communication. By employing VQA models, we assess data representation quality and the general communicative clarity of charts. Experiments were conducted using two leading VQA benchmark datasets, ChartQA and PlotQA, with visualizations generated by OpenAI’s GPT-3.5 Turbo and Meta’s Llama 3.1 70B-Instruct models. Our results indicate that LLM-generated charts do not match the accuracy of the original non-LLM-generated charts based on VQA performance measures. Moreover, while our results demonstrate that few-shot prompting significantly boosts the accuracy of chart generation, considerable progress remains to be made before LLMs can fully match the precision of human-generated graphs. This underscores the importance of our work, which expedites the research process by enabling rapid iteration without the need for human annotation, thus accelerating advancements in this field. James Ford, Xingmeng Zhao, Daniel Schumacher, Anthony Rios |
COLING | 4 |
| 2025 | Reflective Agreement: Combining Self-Mixture of Agents with a Sequence Tagger for Robust Event ExtractionabstractEvent Extraction (EE) involves automatically identifying and extracting structured information about events from unstructured text, including triggers, event types, and arguments. Traditional discriminative models demonstrate high precision but often exhibit limited recall, particularly for nuanced or infrequent events. Conversely, generative approaches leveraging Large Language Models (LLMs) provide higher semantic flexibility and recall but suffer from hallucinations and inconsistent predictions. To address these challenges, we propose Agreement-based Reflective Inference System (ARIS), a hybrid approach combining a Self Mixture of Agents with a discriminative sequence tagger. ARIS explicitly leverages structured model consensus, confidence-based filtering, and an LLM reflective inference module to reliably resolve ambiguities and enhance overall event prediction quality. We further investigate decomposed instruction fine-tuning for enhanced LLM event extraction understanding. Experiments demonstrate our approach outperforms existing state-of-the-art event extraction methods across three benchmark datasets. Fatemeh Haji, Mazal Bethany, C. Jason Chiang, Anthony Rios, Peyman Najafirad |
EMNLP | 4 |
| 2025 | Bike Frames: Understanding the Implicit Portrayal of Cyclists in the NewsabstractIncreasing cycling for transportation or recreation can boost public health and reduce the environmental impacts of vehicles. However, news agencies' ideologies and reporting styles often influence public perception of cycling. For example, if news agencies overly report cycling accidents, it may make people perceive cyclists as "dangerous," reducing the number of opting to cycle. Additionally, a decline in cycling can result in less government funding for safe infrastructure. In this paper, we develop a novel prompting method to detect the perceived perception of cyclists within news headlines. To support this, we introduce a new dataset called "Bike Frames," which contains 31,480 news headlines and 1,500 human annotations. Our analysis focuses on 11,385 headlines from the United States. We also propose the BikeFrame Chain-of-Code (CoC) framework, which predicts cyclist perception, identifies accident-related headlines, and determines fault. This framework uses structured pseudocode to represent logical reasoning steps and incorporates news agency bias to enhance prediction accuracy, outperforming traditional chain-of-thought methods used in large language models. Most importantly, we find that incorporating news bias information significantly impacts performance, improving the average F1 score from .739 to .815. Finally, we conduct a comprehensive case study on U.S. news headlines, revealing differences in reporting between mainstream news agencies and cycling-specific websites, as well as variations in coverage based on the gender of cyclists. WARNING: This paper contains descriptions of accidents and death. Xingmeng Zhao, Daniel Schumacher, Sashank Nalluri, Suhana Shrestha, Xavier Walton, Anthony Rios |
ICWSM | 6 |
| 2024 | Extracting Biomedical Entities from Noisy Audio TranscriptsabstractAutomatic Speech Recognition (ASR) technology is fundamental in transcribing spoken language into text, with considerable applications in the clinical realm, including streamlining medical transcription and integrating with Electronic Health Record (EHR) systems. Nevertheless, challenges persist, especially when transcriptions contain noise, leading to significant drops in performance when Natural Language Processing (NLP) models are applied. Named Entity Recognition (NER), an essential clinical task, is particularly affected by such noise, often termed the ASR-NLP gap. Prior works have primarily studied ASR’s efficiency in clean recordings, leaving a research gap concerning the performance in noisy environments. This paper introduces a novel dataset, BioASR-NER, designed to bridge the ASR-NLP gap in the biomedical domain, focusing on extracting adverse drug reactions and mentions of entities from the Brief Test of Adult Cognition by Telephone (BTACT) exam. Our dataset offers a comprehensive collection of almost 2,000 clean and noisy recordings. In addressing the noise challenge, we present an innovative transcript-cleaning method using GPT-4, investigating both zero-shot and few-shot methodologies. Our study further delves into an error analysis, shedding light on the types of errors in transcription software, corrections by GPT-4, and the challenges GPT-4 faces. This paper aims to foster improved understanding and potential solutions for the ASR-NLP gap, ultimately supporting enhanced healthcare documentation practices. Nima Ebadi, Kellen Morgan, Adrian Tan, Billy Linares, Sheri Osborn, Emma Majors, Jeremy Davis, Anthony Rios |
LREC/COLING | 8 |
| 2024 | A Comprehensive Study of Gender Bias in Chemical Named Entity Recognition ModelsabstractXingmeng Zhao, Ali Niazi, Anthony Rios. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Xingmeng Zhao, Ali Niazi, Anthony Rios |
NAACL-HLT | 3 |
| 2024 | Deciphering Textual Authenticity: A Generalized Strategy through the Lens of Large Language Semantics for Detecting Human vs. Machine-Generated Text
Mazal Bethany, Brandon Wherry, Emet Bethany, Nishant Vishwamitra, Anthony Rios, Peyman Najafirad |
USENIX Security Symposium | 5 |
| 2023 | A marker-based neural network system for extracting social determinants of healthabstractOBJECTIVE: The impact of social determinants of health (SDoH) on patients' healthcare quality and the disparity is well known. Many SDoH items are not coded in structured forms in electronic health records. These items are often captured in free-text clinical notes, but there are limited methods for automatically extracting them. We explore a multi-stage pipeline involving named entity recognition (NER), relation classification (RC), and text classification methods to automatically extract SDoH information from clinical notes. MATERIALS AND METHODS: The study uses the N2C2 Shared Task data, which were collected from 2 sources of clinical notes: MIMIC-III and University of Washington Harborview Medical Centers. It contains 4480 social history sections with full annotation for 12 SDoHs. In order to handle the issue of overlapping entities, we developed a novel marker-based NER model. We used it in a multi-stage pipeline to extract SDoH information from clinical notes. RESULTS: Our marker-based system outperformed the state-of-the-art span-based models at handling overlapping entities based on the overall Micro-F1 score performance. It also achieved state-of-the-art performance compared with the shared task methods. Our approach achieved an F1 of 0.9101, 0.8053, and 0.9025 for Subtasks A, B, and C, respectively. CONCLUSIONS: The major finding of this study is that the multi-stage pipeline effectively extracts SDoH information from clinical notes. This approach can improve the understanding and tracking of SDoHs in clinical settings. However, error propagation may be an issue and further research is needed to improve the extraction of entities with complex semantic meanings and low-frequency entities. We have made the source code available at https://github.com/Zephyr1022/SDOH-N2C2-UTSA. Xingmeng Zhao, Anthony Rios |
J. Am. Medical Informatics Assoc. | 2 |
| 2022 | Measuring Geographic Performance Disparities of Offensive Language ClassifiersabstractText classifiers are applied at scale in the form of one-size-fits-all solutions. Nevertheless, many studies show that classifiers are biased regarding different languages and dialects. When measuring and discovering these biases, some gaps present themselves and should be addressed. First, “Does language, dialect, and topical content vary across geographical regions?” and secondly “If there are differences across the regions, do they impact model performance?”. We introduce a novel dataset called GeoOLID with more than 14 thousand examples across 15 geographically and demographically diverse cities to address these questions. We perform a comprehensive analysis of geographical-related content and their impact on performance disparities of offensive language detection models. Overall, we find that current models do not generalize across locations. Likewise, we show that while offensive language models produce false positives on African American English, model performance is not correlated with each city’s minority population proportions. Warning: This paper contains offensive language. Brandon Lwowski, Peyman Najafirad, Anthony Rios |
COLING | 3 |
| 2022 | Turning Stocks into Memes: A Dataset for Understanding How Social Communities Can Drive Wall Street
Richard Alvarez, Paras Bhatt, Xingmeng Zhao, Anthony Rios |
ICWSM | 4 |
| 2021 | The risk of racial bias while tracking influenza-related content on social media using machine learningabstractOBJECTIVE: Machine learning is used to understand and track influenza-related content on social media. Because these systems are used at scale, they have the potential to adversely impact the people they are built to help. In this study, we explore the biases of different machine learning methods for the specific task of detecting influenza-related content. We compare the performance of each model on tweets written in Standard American English (SAE) vs African American English (AAE). MATERIALS AND METHODS: Two influenza-related datasets are used to train 3 text classification models (support vector machine, convolutional neural network, bidirectional long short-term memory) with different feature sets. The datasets match real-world scenarios in which there is a large imbalance between SAE and AAE examples. The number of AAE examples for each class ranges from 2% to 5% in both datasets. We also evaluate each model's performance using a balanced dataset via undersampling. RESULTS: We find that all of the tested machine learning methods are biased on both datasets. The difference in false positive rates between SAE and AAE examples ranges from 0.01 to 0.35. The difference in the false negative rates ranges from 0.01 to 0.23. We also find that the neural network methods generally has more unfair results than the linear support vector machine on the chosen datasets. CONCLUSIONS: The models that result in the most unfair predictions may vary from dataset to dataset. Practitioners should be aware of the potential harms related to applying machine learning to health-related social media data. At a minimum, we recommend evaluating fairness along with traditional evaluation metrics. Brandon Lwowski, Anthony Rios |
J. Am. Medical Informatics Assoc. | 2 |
| 2020 | FuzzE: Fuzzy Fairness Evaluation of Offensive Language Classifiers on African-American EnglishabstractHate speech and offensive language are rampant on social media. Machine learning has provided a way to moderate foul language at scale. However, much of the current research focuses on overall performance. Models may perform poorly on text written in a minority dialectal language. For instance, a hate speech classifier may produce more false positives on tweets written in African-American Vernacular English (AAVE). To measure these problems, we need text written in both AAVE and Standard American English (SAE). Unfortunately, it is challenging to curate data for all linguistic styles in a timely manner—especially when we are constrained to specific problems, social media platforms, or by limited resources. In this paper, we answer the question, “How can we evaluate the performance of classifiers across minority dialectal languages when they are not present within a particular dataset?” Specifically, we propose an automated fairness fuzzing tool called FuzzE to quantify the fairness of text classifiers applied to AAVE text using a dataset that only contains text written in SAE. Overall, we find that the fairness estimates returned by our technique moderately correlates with the use of real ground-truth AAVE text. Warning: Offensive language is displayed in this manuscript. Anthony Rios |
AAAI | 1 |
| 2020 | An Empirical Study of the Downstream Reliability of Pre-Trained Word EmbeddingsabstractWhile pre-trained word embeddings have been shown to improve the performance of downstream tasks, many questions remain regarding their reliability: Do the same pre-trained word embeddings result in the best performance with slight changes to the training data?Do the same pre-trained embeddings perform well with multiple neural network architectures?Do imputation strategies for unknown words impact reliability?In this paper, we introduce two new metrics to understand the downstream reliability of word embeddings.We find that downstream reliability of word embeddings depends on multiple factors, including, the evaluation metric, the handling of out-of-vocabulary words, and whether the embeddings are fine-tuned. Anthony Rios, Brandon Lwowski |
COLING | 1 |
| 2019 | Neural transfer learning for assigning diagnosis codes to EMRs
Anthony Rios, Ramakanth Kavuluru |
Artif. Intell. Medicine | 1 |
| 2019 | Cross-registry neural domain adaptation to extract mutational test results from pathology reports
Anthony Rios, Eric B. Durbin, Isaac Hands, Susanne M. Arnold, Darshil Shah, Stephen M. Schwartz, Bernardo H. L. Goulart, Ramakanth Kavuluru |
J. Biomed. Informatics | 1 |
| 2018 | Few-Shot and Zero-Shot Multi-Label Learning for Structured Label SpacesabstractLarge multi-label datasets contain labels that occur thousands of times (frequent group), those that occur only a few times (few-shot group), and labels that never appear in the training dataset (zero-shot group). Multi-label few- and zero-shot label prediction is mostly unexplored on datasets with large label spaces, especially for text classification. In this paper, we perform a fine-grained evaluation to understand how state-of-the-art methods perform on infrequent labels. Furthermore, we develop few- and zero-shot methods for multi-label text classification when there is a known structure over the label space, and evaluate them on two publicly available medical text datasets: MIMIC II and MIMIC III. For few-shot labels we achieve improvements of 6.2% and 4.8% in R@10 for MIMIC II and MIMIC III, respectively, over prior efforts; the corresponding R@10 improvements for zero-shot labels are 17.3% and 19%. Anthony Rios, Ramakanth Kavuluru |
EMNLP | 1 |
| 2018 | EMR Coding with Semi-Parametric Multi-Head Matching NetworksabstractCoding EMRs with diagnosis and procedure codes is an indispensable task for billing, secondary data analyses, and monitoring health trends. Both speed and accuracy of coding are critical. While coding errors could lead to more patient-side financial burden and mis-interpretation of a patient's well-being, timely coding is also needed to avoid backlogs and additional costs for the healthcare facility. In this paper, we present a new neural network architecture that combines ideas from few-shot learning matching networks, multi-label loss functions, and convolutional neural networks for text classification to significantly outperform other state-of-the-art models. Our evaluations are conducted using a well known deidentified EMR dataset (MIMIC) with a variety of multi-label performance measures. Anthony Rios, Ramakanth Kavuluru |
NAACL-HLT | 1 |
| 2018 | Generalizing biomedical relation classification with neural adversarial domain adaptationabstractMotivation: Creating large datasets for biomedical relation classification can be prohibitively expensive. While some datasets have been curated to extract protein-protein and drug-drug interactions (PPIs and DDIs) from text, we are also interested in other interactions including gene-disease and chemical-protein connections. Also, many biomedical researchers have begun to explore ternary relationships. Even when annotated data are available, many datasets used for relation classification are inherently biased. For example, issues such as sample selection bias typically prevent models from generalizing in the wild. To address the problem of cross-corpora generalization, we present a novel adversarial learning algorithm for unsupervised domain adaptation tasks where no labeled data are available in the target domain. Instead, our method takes advantage of unlabeled data to improve biased classifiers through learning domain-invariant features via an adversarial process. Finally, our method is built upon recent advances in neural network (NN) methods. Results: We experiment by extracting PPIs and DDIs from text. In our experiments, we show domain invariant features can be learned in NNs such that classifiers trained for one interaction type (protein-protein) can be re-purposed to others (drug-drug). We also show that our method can adapt to different source and target pairs of PPI datasets. Compared to prior convolutional and recurrent NN-based relation classification methods without domain adaptation, we achieve improvements as high as 30% in F1-score. Likewise, we show improvements over state-of-the-art adversarial methods. Availability and implementation: Experimental code is available at https://github.com/bionlproc/adversarial-relation-classification. Supplementary information: Supplementary data are available at Bioinformatics online. Anthony Rios, Ramakanth Kavuluru, Zhiyong Lu |
Bioinform. | 1 |
| 2018 | Data and systems for medication-related text classification and concept normalization from Twitter: insights from the Social Media Mining for Health (SMM4H)-2017 shared taskabstractObjective: We executed the Social Media Mining for Health (SMM4H) 2017 shared tasks to enable the community-driven development and large-scale evaluation of automatic text processing methods for the classification and normalization of health-related text from social media. An additional objective was to publicly release manually annotated data. Materials and Methods: We organized 3 independent subtasks: automatic classification of self-reports of 1) adverse drug reactions (ADRs) and 2) medication consumption, from medication-mentioning tweets, and 3) normalization of ADR expressions. Training data consisted of 15 717 annotated tweets for (1), 10 260 for (2), and 6650 ADR phrases and identifiers for (3); and exhibited typical properties of social-media-based health-related texts. Systems were evaluated using 9961, 7513, and 2500 instances for the 3 subtasks, respectively. We evaluated performances of classes of methods and ensembles of system combinations following the shared tasks. Results: Among 55 system runs, the best system scores for the 3 subtasks were 0.435 (ADR class F1-score) for subtask-1, 0.693 (micro-averaged F1-score over two classes) for subtask-2, and 88.5% (accuracy) for subtask-3. Ensembles of system combinations obtained best scores of 0.476, 0.702, and 88.7%, outperforming individual systems. Discussion: Among individual systems, support vector machines and convolutional neural networks showed high performance. Performance gains achieved by ensembles of system combinations suggest that such strategies may be suitable for operational systems relying on difficult text classification tasks (eg, subtask-1). Conclusions: Data imbalance and lack of context remain challenges for natural language processing of social media text. Annotated data from the shared task have been made available as reference standards for future studies (http://dx.doi.org/10.17632/rxwfb3tysd.1). Abeed Sarker, Maksim Belousov, Jasper Friedrichs, Kai Hakala, Svetlana Kiritchenko, Farrokh Mehryary, Sifei Han, Tung Tran 0001, Anthony Rios, Ramakanth Kavuluru, Berry de Bruijn, Filip Ginter, Debanjan Mahata, Saif M. Mohammad, Goran Nenadic, Graciela Gonzalez-Hernandez |
J. Am. Medical Informatics Assoc. | 9 |
| 2015 | Automatic Assignment of Non-Leaf MeSH Terms to Biomedical Articles
Ramakanth Kavuluru, Anthony Rios |
AMIA | 2 |
| 2015 | An empirical evaluation of supervised learning approaches in assigning diagnosis codes to electronic medical records
Ramakanth Kavuluru, Anthony Rios |
Artif. Intell. Medicine | 2 |
| 2014 | A Knowledge-Based Collaborative Clinical Case Mining Framework
Ramakanth Kavuluru, Anthony Rios, Brandon Kulengowski, Patrick McNamara |
AMIA | 2 |