Bryan Li

dblp:243/6637 · DBLP profile ↗
← Back
10ranked-venue papers
3as first author
8since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
YearPublicationVenuePosition
2026 Incorporating Q&A Nuggets Into Retrieval-Augmented Generation
Laura Dietz, Bryan Li, Gabrielle K. Liu, Jia-Huei Ju, Eugene Yang 0001, Dawn J. Lawrie, William Gantt Walden, James Mayfield
ECIR (2)2
2026 Insider Knowledge: How Much Can RAG Systems Gain from Evaluation Secrets?
Laura Dietz, Bryan Li, Eugene Yang 0001, Dawn J. Lawrie, William Gantt Walden, James Mayfield
ECIR (1)2
2026 Investigating Retrieval-Augmented Generation Systems on Unanswerable, Uncheatable, Realistic, Multi-hop Queries
Gabrielle K. Liu, Bryan Li, Arman Cohan, William Gantt Walden, Eugene Yang 0001
ECIR (2)2
2026 Auto-ARGUE: LLM-Based Report Generation Evaluation
abstract
Generation of citation-backed reports is a primary use case for retrieval-augmented generation (RAG) systems. While open-source evaluation tools exist for various RAG tasks, tools designed for report generation are lacking. Accordingly, we introduce Auto-ARGUE, a robust LLM-based implementation of the recently proposed ARGUE framework for report generation evaluation. We present analysis of Auto-ARGUE on the report generation pilot task from the TREC 2024 NeuCLIR track and on two tasks from the TREC 2024 RAG track, showing good system-level correlations with human judgments. Additionally, we release ARGUE-viz, a web app for visualization and fine-grained analysis of Auto-ARGUE judgments and scores1.
William Gantt Walden, Marc Mason, Orion Weller, Laura Dietz, John M. Conroy, Neil P. Molino, Hannah Recknor, Bryan Li, Gabrielle K. Liu, Dawn J. Lawrie, James Mayfield, Eugene Yang 0001
SIGIR8
2024 Eliciting Better Multilingual Structured Reasoning from LLMs through Code
abstract
The development of large language models (LLM) has shown progress on reasoning, though studies have largely considered either English or simple reasoning tasks.To address this, we introduce a multilingual structured reasoning and explanation dataset, termed xSTREET, that covers four tasks across six languages.xSTREET exposes a gap in base LLM performance between English and non-English reasoning tasks. 1 We then propose two methods to remedy this gap, building on the insight that LLMs trained on code are better reasoners.First, at training time, we augment a code dataset with multilingual comments using machine translation while keeping program code as-is.Second, at inference time, we bridge the gap between training and inference by employing a prompt structure that incorporates step-by-step code primitives to derive new facts and find a solution.Our methods show improved multilingual performance on xSTREET, most notably on the scientific commonsense reasoning subtask.Furthermore, the models show no regression on non-reasoning tasks, thus demonstrating our techniques maintain general-purpose abilities.
Bryan Li, Tamer Alkhouli, Daniele Bonadiman, Nikolaos Pappas 0004, Saab Mansour
ACL (1)1
2024 This Land is Your, My Land: Evaluating Geopolitical Bias in Language Models through Territorial Disputes
abstract
Bryan Li, Samar Haider, Chris Callison-Burch. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Bryan Li, Samar Haider, Chris Callison-Burch
NAACL-HLT1
2024 Retrospective for the Dynamic Sensorium Competition for predicting large-scale mouse primary visual cortex activity from videos
abstract
Understanding how biological visual systems process information is challenging because of the nonlinear relationship between visual input and neuronal responses. Artificial neural networks allow computational neuroscientists to create predictive models that connect biological and machine vision.Machine learning has benefited tremendously from benchmarks that compare different models on the same task under standardized conditions. However, there was no standardized benchmark to identify state-of-the-art dynamic models of the mouse visual system.To address this gap, we established the SENSORIUM 2023 Benchmark Competition with dynamic input, featuring a new large-scale dataset from the primary visual cortex of ten mice. This dataset includes responses from 78,853 neurons to 2 hours of dynamic stimuli per neuron, together with behavioral measurements such as running speed, pupil dilation, and eye movements.The competition ranked models in two tracks based on predictive performance for neuronal responses on a held-out test set: one focusing on predicting in-domain natural stimuli and another on out-of-distribution (OOD) stimuli to assess model generalization.As part of the NeurIPS 2023 Competition Track, we received more than 160 model submissions from 22 teams. Several new architectures for predictive models were proposed, and the winning teams improved the previous state-of-the-art model by 50\%. Access to the dataset as well as the benchmarking infrastructure will remain online at www.sensorium-competition.net.
Polina Turishcheva, Paul G. Fahey, Michaela Vystrcilová, Laura Hansel, Rachel Froebe, Kayla Ponder, Yongrong Qiu, Konstantin Willeke, Mohammad Bashiri, Ruslan Baikulov, Yu Zhu 0008, Lei Ma 0008, Tiejun Huang 0001, Bryan Li, Wolf De Wulf, Nina Kudryashova, Matthias H. Hennig, Nathalie Rochefort, Arno Onken, Eric Y. Wang, Zhiwei Ding, Andreas S. Tolias, Fabian H. Sinz, Alexander S. Ecker
NeurIPS15
2023 Bidirectional Language Models Are Also Few-shot Learners
Ajay Patel, Bryan Li, Mohammad Sadegh Rasooli, Noah Constant, Colin Raffel, Chris Callison-Burch
ICLR2
2020 Exploring Content Selection in Summarization of Novel Chapters
abstract
We present a new summarization task, generating summaries of novel chapters using summary/chapter pairs from online study guides.This is a harder task than the news summarization task, given the chapter length as well as the extreme paraphrasing and generalization found in the summaries.We focus on extractive summarization, which requires the creation of a gold-standard set of extractive summaries.We present a new metric for aligning reference summary sentences with chapter sentences to create gold extracts and also experiment with different alignment methods.Our experiments demonstrate significant improvement over prior alignment approaches for our task as shown through automatic metrics and a crowd-sourced pyramid analysis.
Faisal Ladhak, Bryan Li, Yaser Al-Onaizan, Kathy McKeown
ACL2
2019 Acoustic and Lexical Sentiment Analysis for Customer Service Calls
abstract
We describe the development of a sentiment analysis system for customer service calls, starting with the data acquisition and labeling, and proceeding to the algorithmic information extraction and modeling process from both spoken words and their acoustic expression. The proposed system is based on the combination of multiple acoustic and lexical models in a late fusion approach. Acoustic aspects of sentiment are captured by utterance-level features based on aggregated openSMILE and raw cepstral features, and further augmented with an energy contour model. Lexical aspects are captured by back-off n-gram language models. These models are found to combine effectively, showing different strengths as pertains to positive and negative sentiment detection.
Bryan Li, Dimitrios Dimitriadis, Andreas Stolcke
ICASSP1