VLDB 2026 Research / reviewers in the wild / expert
Xintong Li 0001
dblp:21/8501-1
· DBLP profile ↗
8ranked-venue papers
3as first author
8since 2021 · last 2026
0009-0007-7839-2357ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 3 first-author · 8 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
8 papers |
Language models and text generation · 18% Information extraction and text analysis · 14% Vision and language · 12% | |
| Databases, data mining, and information retrieval
1 paper |
Machine learning and data management · 100% |
Topics — the 24 heaviest of 26, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Vision and language
multimodal reasoning |
1.0 | 1 | 2026 | SceneAlign: Aligning Multimodal Reasoning to Scene Graphs in Complex Visual Scenes · ACL (1) 2026 |
Computer vision › Segmentation and scene understanding
scene graph |
1.0 | 1 | 2026 | SceneAlign: Aligning Multimodal Reasoning to Scene Graphs in Complex Visual Scenes · ACL (1) 2026 |
Machine learning › Efficient and distributed learning
active learning |
0.9 | 1 | 2025 | From Selection to Generation: A Survey of LLM-based Active Learning · ACL (1) 2025 |
Natural language and speech › Language models and text generation
chain-of-thought reasoning |
0.9 | 1 | 2025 | OCEAN: Offline Chain-of-thought Evaluation and Alignment in Large Language Models · ICLR 2025 |
Natural language and speech › Language models and text generation › large language model reasoning
implicit reasoning |
0.9 | 1 | 2025 | Toward Multi-Session Personalized Conversation: A Large-Scale Dataset and Hierarchical Tree Framework for Implicit Reasoning · EMNLP 2025 |
Natural language and speech › Language models and text generation
instruction tuning |
0.9 | 1 | 2025 | CoMMIT: Coordinated Multimodal Instruction Tuning · EMNLP 2025 |
Knowledge, reasoning and agents › Knowledge representation and reasoning
knowledge graph reasoning |
0.9 | 1 | 2025 | OCEAN: Offline Chain-of-thought Evaluation and Alignment in Large Language Models · ICLR 2025 |
Computer vision › Vision and language › vision-language model › multimodal large language model
multimodal instruction tuning |
0.9 | 1 | 2025 | CoMMIT: Coordinated Multimodal Instruction Tuning · EMNLP 2025 |
Machine learning › Reinforcement learning
off-policy evaluation |
0.9 | 1 | 2025 | OCEAN: Offline Chain-of-thought Evaluation and Alignment in Large Language Models · ICLR 2025 |
Natural language and speech › Question answering and dialogue systems
personalized dialogue |
0.9 | 1 | 2025 | Toward Multi-Session Personalized Conversation: A Large-Scale Dataset and Hierarchical Tree Framework for Implicit Reasoning · EMNLP 2025 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › nonmonotonic reasoning › preference handling
preference modeling |
0.9 | 1 | 2025 | OCEAN: Offline Chain-of-thought Evaluation and Alignment in Large Language Models · ICLR 2025 |
Natural language and speech › Information extraction and text analysis › text classification
multi-label text classification |
0.8 | 1 | 2024 | Open-world Multi-label Text Classification with Extremely Weak Supervision · EMNLP 2024 |
Natural language and speech › Information extraction and text analysis
text classification |
0.8 | 1 | 2024 | Open-world Multi-label Text Classification with Extremely Weak Supervision · EMNLP 2024 |
Machine learning › Learning paradigms
weakly supervised learning |
0.8 | 1 | 2024 | Open-world Multi-label Text Classification with Extremely Weak Supervision · EMNLP 2024 |
Machine learning › Transfer learning and domain adaptation › domain adaptation › representation learning for domain adaptation
geometry-aware domain adaptation |
0.7 | 1 | 2023 | Geometry-Aware Adaptation for Pretrained Models · NeurIPS 2023 |
Machine learning › Learning theory
sample complexity |
0.7 | 1 | 2023 | Geometry-Aware Adaptation for Pretrained Models · NeurIPS 2023 |
Machine learning › Transfer learning and domain adaptation
zero-shot learning |
0.7 | 1 | 2023 | Geometry-Aware Adaptation for Pretrained Models · NeurIPS 2023 |
Machine learning and data management › weak supervision
labeling functions |
0.6 | 1 | 2022 | AutoWS-Bench-101: Benchmarking Automated Weak Supervision with 100 Labels · NeurIPS 2022 |
Machine learning and data management › weak supervision
programmatic weak supervision |
0.6 | 1 | 2022 | AutoWS-Bench-101: Benchmarking Automated Weak Supervision with 100 Labels · NeurIPS 2022 |
Machine learning and data management
weak supervision |
0.6 | 1 | 2022 | AutoWS-Bench-101: Benchmarking Automated Weak Supervision with 100 Labels · NeurIPS 2022 |
Machine learning › Reinforcement learning › off-policy evaluation
inverse propensity scoring |
0.3 | 1 | 2025 | OCEAN: Offline Chain-of-thought Evaluation and Alignment in Large Language Models · ICLR 2025 |
Machine learning › Transfer learning and domain adaptation
pre-trained models |
0.2 | 1 | 2023 | Geometry-Aware Adaptation for Pretrained Models · NeurIPS 2023 |
Machine learning › Transfer learning and domain adaptation
few-shot learning |
0.2 | 1 | 2022 | AutoWS-Bench-101: Benchmarking Automated Weak Supervision with 100 Labels · NeurIPS 2022 |
Machine learning › Deep learning architectures and training
foundation model |
0.2 | 1 | 2022 | AutoWS-Bench-101: Benchmarking Automated Weak Supervision with 100 Labels · NeurIPS 2022 |
Methods — techniques the papers use, named apart from their topics
scene graph alignment · 1.0multimodal alignment · 1.0reinforcement learning · 0.9multimodal learning · 0.9markov decision process · 0.9level-based retrieval · 0.9inverse propensity scoring · 0.9hierarchical tree framework · 0.9coordinated tuning · 0.9clustering · 0.8zero-shot foundation models · 0.6supervised learning baselines · 0.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SceneAlign: Aligning Multimodal Reasoning to Scene Graphs in Complex Visual ScenesabstractChuhan Wang, Xintong Li, Jennifer Yuntong Zhang, Junda Wu, Chengkai Huang, Lina Yao, Julian McAuley, Jingbo Shang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Xintong Li 0001, Jennifer Yuntong Zhang, Junda Wu, Chengkai Huang, Lina Yao 0001, Julian J. McAuley, Jingbo Shang |
ACL (1) | 2 |
| 2025 | From Selection to Generation: A Survey of LLM-based Active LearningabstractYu Xia, Subhojyoti Mukherjee, Zhouhang Xie, Junda Wu, Xintong Li, Ryan Aponte, Hanjia Lyu, Joe Barrow, Hongjie Chen, Franck Dernoncourt, Branislav Kveton, Tong Yu, Ruiyi Zhang, Jiuxiang Gu, Nesreen K. Ahmed, Yu Wang, Xiang Chen, Hanieh Deilamsalehy, Sungchul Kim, Zhengmian Hu, Yue Zhao, Nedim Lipka, Seunghyun Yoon, Ting-Hao Kenneth Huang, Zichao Wang, Puneet Mathur, Soumyabrata Pal, Koyel Mukherjee, Zhehao Zhang, Namyong Park, Thien Huu Nguyen, Jiebo Luo, Ryan A. Rossi, Julian McAuley. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Yu Xia 0007, Subhojyoti Mukherjee, Zhouhang Xie, Junda Wu, Xintong Li 0001, Ryan Aponte, Hanjia Lyu, Joe Barrow, Hongjie Chen 0003, Franck Dernoncourt, Branislav Kveton, Tong Yu 0001, Ruiyi Zhang 0002, Jiuxiang Gu, Nesreen K. Ahmed, Yu Wang 0160, Xiang Chen 0010, Hanieh Deilamsalehy, Sungchul Kim, Zhengmian Hu, Yue Zhao 0016, Nedim Lipka, Seunghyun Yoon 0002, Ting-Hao 'Kenneth' Huang, Zichao Wang 0001, Puneet Mathur, Soumyabrata Pal, Koyel Mukherjee 0001, Zhehao Zhang 0001, Namyong Park 0001, Thien Huu Nguyen, Jiebo Luo 0001, Ryan Rossi, Julian J. McAuley |
ACL (1) | 5 |
| 2025 | Toward Multi-Session Personalized Conversation: A Large-Scale Dataset and Hierarchical Tree Framework for Implicit ReasoningabstractThere has been a surge in the use of large language models (LLM) conversational agents to generate responses based on long-term history from multiple sessions.However, existing long-term open-domain dialogue datasets lack complex, real-world personalization and fail to capture implicit reasoning-where relevant information is embedded in subtle, syntactic, or semantically distant connections rather than explicit statements.In such cases, traditional retrieval methods fail to capture relevant context, and long-context modeling also becomes inefficient due to numerous complicated personarelated details.To address this gap, we introduce IMPLEXCONV, a large-scale long-term dataset with 2,500 examples, each containing approximately 100 conversation sessions, designed to study implicit reasoning in personalized dialogues.Additionally, we propose TAC-ITREE, a novel hierarchical tree framework that structures conversation history into multiple levels of summarization.Instead of brute-force searching all data, TACITREE enables an efficient, level-based retrieval process where models refine their search by progressively selecting relevant details.Our experiments demonstrate that TACITREE significantly improves the ability of LLMs to reason over long-term conversations with implicit contextual dependencies.I love playing sports like basketball and swimming.I'm going on a sports stadium tour!Aug 15, 2024 I broke my leg in a car accident... Xintong Li 0001, Jalend Bantupalli, Ria Dharmani, Yuwei Zhang 0001, Jingbo Shang |
EMNLP | 1 |
| 2025 | CoMMIT: Coordinated Multimodal Instruction TuningabstractXintong Li, Junda Wu, Tong Yu, Rui Wang, Yu Wang, Xiang Chen, Jiuxiang Gu, Lina Yao, Julian McAuley, Jingbo Shang. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Xintong Li 0001, Junda Wu, Tong Yu 0001, Rui Wang 0088, Xiang Chen 0010, Jiuxiang Gu, Lina Yao 0001, Julian J. McAuley, Jingbo Shang |
EMNLP | 1 |
| 2025 | OCEAN: Offline Chain-of-thought Evaluation and Alignment in Large Language ModelsabstractOffline evaluation of LLMs is crucial in understanding their capacities, though current methods remain underexplored in existing research. In this work, we focus on the offline evaluation of the chain-of-thought capabilities and show how to optimize LLMs based on the proposed evaluation method. To enable offline feedback with rich knowledge and reasoning paths, we use knowledge graphs (KGs) (e.g., Wikidata5M) to provide feedback on the generated chain of thoughts. Due to the heterogeneity between LLM reasoning and KG structures, direct interaction and feedback from knowledge graphs on LLM behavior are challenging, as they require accurate entity linking and grounding of LLM-generated chains of thought in the KG. To address the above challenge, we propose an offline chain-of-thought evaluation framework, OCEAN, which models chain-of-thought reasoning in LLMs as a Markov Decision Process (MDP), and evaluate the policy’s alignment with KG preference modeling. To overcome the reasoning heterogeneity and grounding problems, we leverage on-policy KG exploration and reinforcement learning to model a KG policy that generates token-level likelihood distributions for LLM-generated chain-of-thought reasoning paths, simulating KG reasoning preference. Then we incorporate the knowledge-graph feedback on the validity and alignment of the generated reasoning paths into inverse propensity scores and propose KG-IPS estimator. Theoretically, we prove the unbiasedness of the proposed KG-IPS estimator and provide a lower bound on its variance. With the off-policy evaluated value function, we can directly enable off-policy optimization to further enhance chain-of-thought alignment. Our empirical study shows that OCEAN can be efficiently optimized for generating chain-of-thought reasoning paths with higher estimated values without affecting LLMs’ general abilities in downstream tasks or their internal knowledge. Junda Wu, Xintong Li 0001, Ruoyu Wang 0038, Yu Xia 0007, Yuxin Xiong, Jianing Wang 0002, Tong Yu 0001, Xiang Chen 0010, Branislav Kveton, Lina Yao 0001, Jingbo Shang, Julian J. McAuley |
ICLR | 2 |
| 2024 | Open-world Multi-label Text Classification with Extremely Weak SupervisionabstractWe study open-world multi-label text classification under extremely weak supervision (XWS), where the user only provides a brief description for classification objectives without any labels or ground-truth label space.Similar single-label XWS settings have been explored recently, however, these methods cannot be easily adapted for multi-label.We observe that (1) most documents have a dominant class covering the majority of content and (2) long-tail labels would appear in some documents as a dominant class.Therefore, we first utilize the user description to prompt a large language model (LLM) for dominant keyphrases of a subset of raw documents, and then construct a (initial) label space via clustering.We further apply a zero-shot multi-label classifier to locate the documents with small top predicted scores, so we can revisit their dominant keyphrases for more long-tail labels.We iterate this process to discover a comprehensive label space and construct a multi-label classifier as a novel method, X-MLClass.X-MLClass exhibits a remarkable increase in ground-truth label space coverage on various datasets, for example, a 40% improvement on the AAPD dataset over topic modeling and keyword extraction methods.Moreover, X-MLClass achieves the best end-to-end multi-label classification accuracy. B Prompt Templates for Generating KeyphrasesCode 2 provides an example of the prompt used to generate keyphrases for a selected chunk of the Amazon-531 dataset.Users can help us define the objective with examples.For example, the coarse-grained objectives look like "games" and "animals", while the corresponding fine-grained objectives are "trading_card_games" and "reptiles". Xintong Li 0001, Jinya Jiang, Ria Dharmani, Jayanth Srinivasa, Gaowen Liu, Jingbo Shang |
EMNLP | 1 |
| 2023 | Geometry-Aware Adaptation for Pretrained ModelsabstractMachine learning models---including prominent zero-shot models---are often trained on datasets whose labels are only a small proportion of a larger label space. Such spaces are commonly equipped with a metric that relates the labels via distances between them. We propose a simple approach to exploit this information to adapt the trained model to reliably predict new classes---or, in the case of zero-shot prediction, to improve its performance---without any additional training. Our technique is a drop-in replacement of the standard prediction rule, swapping $\text{argmax}$ with the Fréchet mean. We provide a comprehensive theoretical analysis for this approach, studying (i) learning-theoretic results trading off label space diameter, sample complexity, and model dimension, (ii) characterizations of the full range of scenarios in which it is possible to predict any unobserved class, and (iii) an optimal active learning-like next class selection procedure to obtain optimal training classes for when it is not possible to predict the entire range of unobserved classes. Empirically, using easily-available external metrics, our proposed approach, Loki, gains up to 29.7% relative improvement over SimCLR on ImageNet and scales to hundreds of thousands of classes. When no such metric is available, Loki can use self-derived metrics from class embeddings and obtains a 10.5% improvement on pretrained zero-shot models such as CLIP. Nicholas Carl Roberts, Xintong Li 0001, Dyah Adila, Sonia Cromp, Tzu-Heng Huang, Jitian Zhao, Frederic Sala |
NeurIPS | 2 |
| 2022 | AutoWS-Bench-101: Benchmarking Automated Weak Supervision with 100 LabelsabstractWeak supervision (WS) is a powerful method to build labeled datasets for training supervised models in the face of little-to-no labeled data. It replaces hand-labeling data with aggregating multiple noisy-but-cheap label estimates expressed by labeling functions (LFs). While it has been used successfully in many domains, weak supervision's application scope is limited by the difficulty of constructing labeling functions for domains with complex or high-dimensional features. To address this, a handful of methods have proposed automating the LF design process using a small set of ground truth labels. In this work, we introduce AutoWS-Bench-101: a framework for evaluating automated WS (AutoWS) techniques in challenging WS settings---a set of diverse application domains on which it has been previously difficult or impossible to apply traditional WS techniques. While AutoWS is a promising direction toward expanding the application-scope of WS, the emergence of powerful methods such as zero-shot foundation models reveal the need to understand how AutoWS techniques compare or cooperate with modern zero-shot or few-shot learners. This informs the central question of AutoWS-Bench-101: given an initial set of 100 labels for each task, we ask whether a practitioner should use an AutoWS method to generate additional labels or use some simpler baseline, such as zero-shot predictions from a foundation model or supervised learning. We observe that it is necessary for AutoWS methods to incorporate signal from foundation models if they are to outperform simple few-shot baselines, and AutoWS-Bench-101 promotes future research in this direction. We conclude with a thorough ablation study of AutoWS methods. Nicholas Carl Roberts, Xintong Li 0001, Tzu-Heng Huang, Dyah Adila, Spencer Schoenberg, Cheng-Yu Liu 0002, Lauren Pick, Aws Albarghouthi, Frederic Sala |
NeurIPS | 2 |