EDBT 2026 Demo / reviewers in the wild / expert
Xiang Li 0069
dblp:40/1491-69 · also Xiang Lorraine Li
· DBLP profile ↗
22ranked-venue papers
4as first author
16since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 21 · 3 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MMBERT: Scaled Mixture-of-Experts Multimodal BERT for Robust Chinese Hate Speech Detection Under Cloaking PerturbationsabstractHate speech detection on Chinese social networks presents distinct challenges, particularly due to the widespread use of cloaking techniques designed to evade conventional text-based detection systems. Although large language models (LLMs) have recently improved hate speech detection capabilities, the majority of existing work has concentrated on English datasets, with limited attention given to multimodal strategies in the Chinese context. In this study, we propose MMBERT, a novel BERT-based multimodal framework that integrates textual, speech, and visual modalities through a Mixture-of-Experts (MoE) architecture. To address the instability associated with directly integrating MoE into BERT-based models, we develop a progressive three-stage training paradigm. MMBERT incorporates modality-specific experts, a shared self-attention mechanism, and a router-based expert allocation strategy to enhance robustness against adversarial perturbations. Empirical results in several Chinese hate speech datasets show that MMBERT significantly surpasses fine-tuned BERT-based encoder models, fine-tuned LLMs, and LLMs utilizing in-context learning approaches. Qiyao Xue, Yuchen Dou, Zheyuan Shi, Xiang Li 0069 |
AAAI | 4 |
| 2026 | Neuron-Aware Active Few-Shot Learning for LLMsabstractActive Few-Shot Learning (AFSL) adapts LLMs to specialized domains by identifying the most valuable unlabeled samples for annotation and use as few-shot demonstrations, effectively reducing human annotation costs while promoting high performance.However, existing methods typically rely on output-level signals for the sample identification, such as predictive entropy or semantic similarities with test-time data based on external embeddings, which often overlook models' internal dynamics which could pinpoint specific knowledge gaps.To bridge this gap, we propose NEUFS, a Neuron-Aware Active Few-Shot Learning framework that shifts the selection paradigm from output-level proxies to models' internal dynamics.NEUFS utilizes neuron activation patterns to represent sample directly, and includes a dual-criteria selection strategy that: (1) ensures few-shot sample diversity with neuron patterns for broader example coverage, while (2) prioritizing on identifying informative and challenging few-shot samples LLMs tend to hallucinate by quantifying neuron consensus.Experiments on three datasets demonstrate that NEUFS excels in both reasoning and text classification tasks, outperforming existing AFSL baselines.Ablation studies further highlight that internal neuron activations provide a more principled and effective selection signal than external embeddings, validating the superiority of the proposed NEUFS. Zhuowei Chen, Christian D. Schunn, Raquel Coelho, Xiang Li 0069 |
ACL (1) | 5 |
| 2025 | Think Globally, Group Locally: Evaluating LLMs Using Multi-Lingual Word Grouping GamesabstractLarge language models (LLMs) can exhibit biases in reasoning capabilities due to linguistic modality, performing better on tasks in one language versus another, even with similar content.Most previous works evaluate this through reasoning tasks where reliance on strategies or knowledge can ensure success, such as in commonsense or math tasks.However, abstract reasoning is vital to reasoning for everyday life, where people apply "out-of-the-box thinking" to identify and use patterns for solutions, without a reliance on formulaic approaches.Comparatively, little work has evaluated linguistic biases in this task type.In this paper, we propose a task inspired by the New York Times Connections: GLOBALGROUP, that evaluates models in an abstract reasoning task across several languages.We constructed a game benchmark with five linguistic backgrounds -English, Spanish, Chinese, Hindi, and Arabic -in both the native language and an English translation for comparison.We also proposed game difficulty measurements to evaluate models on games with similar difficulty, enabling a more controlled comparison, which is particularly important in reasoning evaluations.Through experimentation, we find English modalities largely lead to better performance in this abstract reasoning task, and performance disparities between open-and closed-source models. César Guerra-Solano, Zhuochun Li, Xiang Li 0069 |
EMNLP | 3 |
| 2025 | Leveraging Large Models to Evaluate Novel Content: A Case Study on Advertisement CreativityabstractEvaluating creativity is challenging, even for humans, not only because of its subjectivity but also because it involves complex cognitive processes. Inspired by work in marketing, we attempt to break down visual advertisement creativity into atypicality and originality. With fine-grained human annotations on these dimensions, we propose a suite of tasks specifically for such a subjective problem. We also evaluate the alignment between state-of-the-art (SoTA) vision language models (VLMs) and humans on our proposed benchmark, demonstrating both the promises and challenges of using VLMs for automatic creativity assessment. Zhaoyi Hou, Adriana Kovashka, Xiang Li 0069 |
EMNLP | 3 |
| 2024 | Every Answer Matters: Evaluating Commonsense with Probabilistic MeasuresabstractQi Cheng, Michael Boratko, Pranay Kumar Yelugam, Tim O’Gorman, Nalini Singh, Andrew McCallum, Xiang Li. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Michael Boratko, Pranay Kumar Yelugam, Tim O'Gorman, Nalini Singh, Andrew McCallum, Xiang Li 0069 |
ACL (1) | 7 |
| 2024 | Incorporating Geo-Diverse Knowledge into Prompting for Increased Geographical Robustness in Object RecognitionabstractExisting object recognition models have been shown to lack robustness in diverse geographical scenarios due to domain shifts in design and context. Class representations need to be adapted to more accurately reflect an object concept under these shifts. In the absence of training data from target geographies, we hypothesize that geographically diverse descriptive knowledge of categories can enhance robustness. For this purpose, we explore the feasibility of probing a large language model for geography-based object knowledge, and we examine the effects of integrating knowledge into zero-shot and learnable soft prompting with CLIP. Within this exploration, we propose geog-raphy knowledge regularization to ensure that soft prompts trained on a source set of geographies generalize to an un-seen target set. Accuracy gains over prompting baselines on DollarStreet while training only on Europe data are up to +2.8/1.2/1.6 on target data from Africa/Asia/Americas, and +4.6 overall on the hardest classes. Competitive performance is shown vs. few-shot target training, and analysis is provided to direct future study of geographical robustness. Kyle Buettner, Sina Malakouti, Xiang Li 0069, Adriana Kovashka |
CVPR | 3 |
| 2024 | In Search of the Long-Tail: Systematic Generation of Long-Tail Inferential Knowledge via Logical Rule Guided SearchabstractHuihan Li, Yuting Ning, Zeyi Liao, Siyuan Wang, Xiang Lorraine Li, Ximing Lu, Wenting Zhao, Faeze Brahman, Yejin Choi, Xiang Ren. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Huihan Li 0001, Yuting Ning, Zeyi Liao, Xiang Li 0069, Ximing Lu, Faeze Brahman, Yejin Choi 0001, Xiang Ren 0001 |
EMNLP | 5 |
| 2024 | PlaSma: Procedural Knowledge Models for Language-based Planning and Re-PlanningabstractProcedural planning, which entails decomposing a high-level goal into a sequence of temporally ordered steps, is an important yet intricate task for machines. It involves integrating common-sense knowledge to reason about complex and often contextualized situations, e.g. ``scheduling a doctor's appointment without a phone''. While current approaches show encouraging results using large language models (LLMs), they are hindered by drawbacks such as costly API calls and reproducibility issues. In this paper, we advocate planning using smaller language models. We present PlaSma, a novel two-pronged approach to endow small language models with procedural knowledge and (constrained) language-based planning capabilities. More concretely, we develop *symbolic procedural knowledge distillation* to enhance the commonsense knowledge in small language models and an *inference-time algorithm* to facilitate more structured and accurate reasoning. In addition, we introduce a new related task, *Replanning*, that requires a revision of a plan to cope with a constrained situation. In both the planning and replanning settings, we show that orders-of-magnitude smaller models (770M-11B parameters) can compete and often surpass their larger teacher models' capabilities. Finally, we showcase successful application of PlaSma in an embodied environment, VirtualHome. Faeze Brahman, Chandra Bhagavatula, Valentina Pyatkin, Jena D. Hwang, Xiang Li 0069, Hirona Jacqueline Arai, Soumya Sanyal 0001, Keisuke Sakaguchi, Xiang Ren 0001, Yejin Choi 0001 |
ICLR | 5 |
| 2024 | UNcommonsense Reasoning: Abductive Reasoning about Uncommon SituationsabstractWenting Zhao, Justin T. Chiu, Jena Hwang, Faeze Brahman, Jack Hessel, Sanjiban Choudhury, Yejin Choi, Xiang Lorraine Li, Alane Suhr. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Justin T. Chiu, Jena D. Hwang, Faeze Brahman, Jack Hessel, Sanjiban Choudhury, Yejin Choi 0001, Xiang Li 0069, Alane Suhr |
NAACL-HLT | 8 |
| 2023 | Editing Common Sense in TransformersabstractEditing model parameters directly in Transformers makes updating open-source transformer-based models possible without re-training (Meng et al., 2023).However, these editing methods have only been evaluated on statements about encyclopedic knowledge with a single correct answer.Commonsense knowledge with multiple correct answers, e.g., an apple can be green or red but not transparent, has not been studied but is as essential for enhancing transformers' reliability and usefulness.In this paper, we investigate whether commonsense judgments are causally associated with localized, editable parameters in Transformers, and we provide an affirmative answer.We find that directly applying the MEMIT editing algorithm results in sub-par performance, and propose to improve it for the commonsense domain by varying edit tokens and improving the layer selection strategy, i.e., MEMIT CSK .GPT-2 Large and XL models edited using MEMIT CSK outperform best-fine-tuned baselines by 10.97% and 10.73% F1 scores on PEP3k and 20Q datasets.In addition, we propose a novel evaluation dataset, PROBE SET, that contains unaffected and affected neighborhoods, affected paraphrases, and affected reasoning challenges.MEMIT CSK performs well across the metrics while fine-tuning baselines show significant trade-offs between unaffected and affected metrics.These results suggest a compelling future direction for incorporating feedback about common sense into Transformers through direct model editing. 1 * Co-first and last authors.Lorraine's work done at AI2. 1 Code and datasets for all experiments are available at https://github.com/anshitag/memit_csk Anshita Gupta, Debanjan Mondal, Akshay Krishna Sheshadri, Wenlong Zhao 0001, Xiang Li 0069, Sarah Wiegreffe, Niket Tandon |
EMNLP | 5 |
| 2023 | Faith and Fate: Limits of Transformers on CompositionalityabstractTransformer large language models (LLMs) have sparked admiration for their exceptional performance on tasks that demand intricate multi-step reasoning. Yet, these models simultaneously show failures on surprisingly trivial problems.
This begs the question: Are these errors incidental, or do they signal more substantial limitations?
In an attempt to demystify transformer LLMs, we investigate the limits of these models across three representative compositional tasks---multi-digit multiplication, logic grid puzzles, and a classic dynamic programming problem. These tasks require breaking problems down into sub-steps and synthesizing these steps into a precise answer. We formulate compositional tasks as computation graphs to systematically quantify the level of complexity, and break down reasoning steps into intermediate sub-procedures.
Our empirical findings suggest that transformer LLMs solve compositional tasks by reducing multi-step compositional reasoning into linearized subgraph matching, without necessarily developing systematic problem-solving skills. To round off our empirical study, we provide theoretical arguments on abstract multi-step reasoning problems that highlight how autoregressive generations' performance can rapidly decay with increased task complexity. Nouha Dziri, Ximing Lu, Melanie Sclar, Xiang Li 0069, Bill Y. Lin, Sean Welleck, Peter West, Chandra Bhagavatula, Ronan Le Bras 0001, Jena D. Hwang, Soumya Sanyal 0001, Xiang Ren 0001, Allyson Ettinger, Zaïd Harchaoui, Yejin Choi 0001 |
NeurIPS | 4 |
| 2022 | Word2Box: Capturing Set-Theoretic Semantics of Words using Box EmbeddingsabstractShib Dasgupta, Michael Boratko, Siddhartha Mishra, Shriya Atmakuri, Dhruvesh Patel, Xiang Li, Andrew McCallum. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Shib Sankar Dasgupta, Michael Boratko, Siddhartha Mishra, Shriya Atmakuri, Dhruvesh Patel, Xiang Li 0069, Andrew McCallum |
ACL (1) | 6 |
| 2022 | A Systematic Investigation of Commonsense Knowledge in Large Language ModelsabstractXiang Lorraine Li, Adhiguna Kuncoro, Jordan Hoffmann, Cyprien de Masson d’Autume, Phil Blunsom, Aida Nematzadeh. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. Xiang Li 0069, Adhiguna Kuncoro, Jordan Hoffmann, Cyprien de Masson d'Autume, Phil Blunsom, Aida Nematzadeh |
EMNLP | 1 |
| 2021 | Answering Complex Open-Domain Questions with Multi-Hop Dense Retrieval
Wenhan Xiong, Xiang Li 0069, Srinivasan Iyer 0001, Jingfei Du, Patrick S. H. Lewis, William Yang Wang, Yashar Mehdad, Scott Yih, Sebastian Riedel 0001, Douwe Kiela, Barlas Oguz |
ICLR | 2 |
| 2021 | Probabilistic Box Embeddings for Uncertain Knowledge Graph ReasoningabstractXuelu Chen, Michael Boratko, Muhao Chen, Shib Sankar Dasgupta, Xiang Lorraine Li, Andrew McCallum. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Xuelu Chen, Michael Boratko, Muhao Chen 0001, Shib Sankar Dasgupta, Xiang Li 0069, Andrew McCallum |
NAACL-HLT | 5 |
| 2021 | Looking Beyond Sentence-Level Natural Language Inference for Question Answering and Text SummarizationabstractAnshuman Mishra, Dhruvesh Patel, Aparna Vijayakumar, Xiang Lorraine Li, Pavan Kapanipathi, Kartik Talamadupula. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Anshuman Mishra, Dhruvesh Patel, Aparna Vijayakumar, Xiang Li 0069, Pavan Kapanipathi, Kartik Talamadupula |
NAACL-HLT | 4 |
| 2020 | ProtoQA: A Question Answering Dataset for Prototypical Common-Sense ReasoningabstractGiven questions regarding some prototypical situation -such as Name something that people usually do before they leave the house for work?-a human can easily answer them via acquired experiences.There can be multiple right answers for such questions, with some more common for a situation than others.This paper introduces a new question answering dataset for training and evaluating common sense reasoning capabilities of artificial intelligence systems in such prototypical situations.The training set is gathered from an existing set of questions played in a longrunning international game show -FAMILY-FEUD.The hidden evaluation set is created by gathering answers for each question from 100 crowd-workers.We also propose a generative evaluation task where a model has to output a ranked list of answers, ideally covering all prototypical answers for a question.After presenting multiple competitive baseline models, we find that human performance still exceeds model scores on all evaluation metrics with a meaningful gap, supporting the challenging nature of the task. * Equal contribution.(i) Name something that people usually do before they leave for work?Ask 100 crowd-workers + manual clustering Michael Boratko, Xiang Li 0069, Tim O'Gorman, Rajarshi Das, Dan Le, Andrew McCallum |
EMNLP (1) | 2 |
| 2020 | Improving Local Identifiability in Probabilistic Box EmbeddingsabstractGeometric embeddings have recently received attention for their natural ability to represent transitive asymmetric relations via containment. Box embeddings, where objects are represented by n-dimensional hyperrectangles, are a particularly promising example of such an embedding as they are closed under intersection and their volume can be calculated easily, allowing them to naturally represent calibrated probability distributions. The benefits of geometric embeddings also introduce a problem of local identifiability, however, where whole neighborhoods of parameters result in equivalent loss which impedes learning. Prior work addressed some of these issues by using an approximation to Gaussian convolution over the box parameters, however this intersection operation also increases the sparsity of the gradient. In this work we model the box parameters with min and max Gumbel distributions, which were chosen such that the space is still closed under the operation of intersection. The calculation of the expected intersection volume involves all parameters, and we demonstrate experimentally that this drastically improves the ability of such models to learn. Shib Sankar Dasgupta, Michael Boratko, Luke Vilnis, Xiang Li 0069, Andrew McCallum |
NeurIPS | 5 |
| 2019 | Smoothing the Geometry of Probabilistic Box Embeddings
Xiang Li 0069, Luke Vilnis, Michael Boratko, Andrew McCallum |
ICLR | 1 |
| 2018 | Probabilistic Embedding of Knowledge Graphs with Box Lattice MeasuresabstractEmbedding methods which enforce a partial order or lattice structure over the concept space, such as Order Embeddings (OE) (Vendrov et al., 2016), are a natural way to model transitive relational data (e.g.entailment graphs).However, OE learns a deterministic knowledge base, limiting expressiveness of queries and the ability to use uncertainty for both prediction and learning (e.g.learning from expectations).Probabilistic extensions of OE (Lai and Hockenmaier, 2017) have provided the ability to somewhat calibrate these denotational probabilities while retaining the consistency and inductive bias of ordered models, but lack the ability to model the negative correlations found in real-world knowledge.In this work we show that a broad class of models that assign probability measures to OE can never capture negative correlation, which motivates our construction of a novel box lattice and accompanying probability measure to capture anticorrelation and even disjoint concepts, while still providing the benefits of probabilistic modeling, such as the ability to perform rich joint and conditional queries over arbitrary sets of concepts, and both learning from and predicting calibrated uncertainty.We show improvements over previous approaches in modeling the Flickr and WordNet entailment graphs, and investigate the power of the model. * Equal contribution. Luke Vilnis, Xiang Li 0069, Shikhar Murty, Andrew McCallum |
ACL (1) | 2 |
| 2016 | Commonsense Knowledge Base CompletionabstractWe enrich a curated resource of commonsense knowledge by formulating the problem as one of knowledge base completion (KBC). Most work in KBC focuses on knowledge bases like Freebase that relate entities drawn from a fixed set. However, the tuples in ConceptNet (Speer and Havasi, 2012) define relations between an unbounded set of phrases. We develop neural network models for scoring tuples on arbitrary phrases and evaluate them by their ability to distinguish true held-out tuples from false ones. We find strong performance from a bilinear model using a simple additive architecture to model phrases. We manually evaluate our trained model’s ability to assign quality scores to novel tuples, finding that it can propose tuples at the same quality level as mediumconfidence tuples from ConceptNet. Xiang Li 0069, Aynaz Taheri, Lifu Tu, Kevin Gimpel |
ACL (1) | 1 |
| 2016 | Interactive provenance summaries for reproducible scienceabstractRecorded provenance facilitates reproducible science. Provenance metadata can help determine how data were possibly transformed, processed, and derived from original sources. While provenance is crucial for verification and validation, there remains the issue of the granularity - detail at which provenance data must be provided to a user, especially for conducting reproducible science. When data are reproduced successfully the need for detailed provenance is minimal and an essence of the recorded provenance suffices. However, when data are not reproduced correctly users want to quickly drill down into fine-grained provenance to understand causes for failure. In this paper, we describe a drill-up/drill-down method for exploring provenance traces. The drill-up method summarizes the trace by grouping nodes and edges of the trace that have same derivation histories. The method preserves provenance data flow semantics. The drill-down method compares summary groups and ranks groups that may have information about the errors. Both the methods are implemented in an efficient manner using light-weight data structures so as to be suitable for reproducible science. We conduct a thorough experimental analysis to show how the operators perform in compressing and expanding real provenance graphs. Xiang Li 0069, Tanu Malik |
eScience | 1 |