VLDB 2026 Research / reviewers in the wild / expert
Chunpu Xu
dblp:251/9631
· DBLP profile ↗
14ranked-venue papers
5as first author
11since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 5 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SCALE: Selective Resource Allocation for Overcoming Performance Bottlenecks in Mathematical Test-time ScalingabstractTest-time compute scaling has emerged as a powerful paradigm for enhancing mathematical reasoning in large language models (LLMs) by allocating additional computational resources during inference. However, current methods employ uniform resource distribution across all reasoning sub-problems, creating fundamental bottlenecks where challenging sub-problems receive insufficient attention while routine operations consume disproportionate resources. This uniform allocation creates performance bottlenecks where additional computational resources yield diminishing returns. Inspired by dual-process theory, we propose SCALE (Selective Resource Allocation), a framework that selectively allocates computational resources based on sub-problem difficulty. SCALE operates through four stages: (1) problem decomposition into sequential reasoning sub-problems, (2) difficulty assessment of each sub-problem to distinguish between routine operations and computationally challenging sub-problems, (3) selective processing mode assignment between System 1 for simple sub-problems and System 2 for complex ones, and (4) sequential execution with context propagation. By concentrating resources on challenging sub-problems while processing routine operations efficiently, SCALE achieves substantial performance improvements with superior resource utilization. Extensive experiments demonstrate that SCALE significantly outperforms uniform scaling baselines, achieving accuracy improvements of up to 13.75 percentage points (57.50% to 71.25% on AIME25) while reducing computational costs by 33-53%, representing a major advance in test-time scaling that addresses fundamental limitations of current approaches. Chunpu Xu, Ruifeng Yuan, Jessie Wang 0004, Wenjie Li 0002, Pengfei Liu 0003 |
AAAI | 2 |
| 2026 | Foresight Optimization for Strategic Reasoning in Large Language ModelsabstractJessie Wang, Jiawen Duan, Jian Wang, Kaitao Song, Chunpu Xu, Johnny K. W. Ho, YU Fenggang, Johan F. Hoorn, Wenjie Li. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Jessie Wang 0004, Jiawen Duan, Jian Wang 0054, Kaitao Song, Chunpu Xu, Johnny K. W. Ho, Fenggang Yu, Johan F. Hoorn, Wenjie Li 0002 |
ACL (1) | 5 |
| 2025 | Towards Dynamic Theory of Mind: Evaluating LLM Adaptation to Temporal Evolution of Human StatesabstractAs Large Language Models (LLMs) increasingly participate in human-AI interactions, evaluating their Theory of Mind (ToM) capabilities - particularly their ability to track dynamic mental states - becomes crucial. While existing benchmarks assess basic ToM abilities, they predominantly focus on static snapshots of mental states, overlooking the temporal evolution that characterizes real-world social interactions. We present DynToM, a novel benchmark specifically designed to evaluate LLMs’ ability to understand and track the temporal progression of mental states across interconnected scenarios. Through a systematic four-step framework, we generate 1,100 social contexts encompassing 5,500 scenarios and 78,100 questions, each validated for realism and quality. Our comprehensive evaluation of ten state-of-the-art LLMs reveals that their average performance underperforms humans by 44.7%, with performance degrading significantly when tracking and reasoning about the shift of mental states. This performance gap highlights fundamental limitations in current LLMs’ ability to model the dynamic nature of human mental states. Jiashuo Wang, Qiancheng Xu, Changhe Song, Chunpu Xu, Wenjie Li 0002, Pengfei Liu 0003 |
ACL (1) | 5 |
| 2025 | MIO: A Foundation Model on Multimodal TokensabstractZekun Moore Wang, King Zhu, Chunpu Xu, Wangchunshu Zhou, Jiaheng Liu, Yibo Zhang, Jessie Wang, Ning Shi, Siyu Li, Yizhi Li, Haoran Que, Zhaoxiang Zhang, Yuanxing Zhang, Ge Zhang, Ke Xu, Jie Fu, Wenhao Huang. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Zekun Moore Wang, King Zhu, Chunpu Xu, Wangchunshu Zhou, Jessie Jiashuo Wang, Ning Shi, Haoran Que, Zhaoxiang Zhang 0001, Yuanxing Zhang, Ge Zhang 0009, Ke Xu 0001, Jie Fu 0001, Wenhao Huang 0001 |
EMNLP | 3 |
| 2025 | LIMOPro: Reasoning Refinement for Efficient and Effective Test-time ScalingabstractLarge language models (LLMs) have demonstrated remarkable reasoning capabilities through test-time scaling approaches, particularly when fine-tuned with chain-of-thought (CoT) data distilled from more powerful large reasoning models (LRMs). However, these reasoning chains often contain verbose elements that mirror human problem-solving, categorized as progressive reasoning (the essential solution development path) and functional elements (verification processes, alternative solution approaches, and error corrections). While progressive reasoning is crucial, the functional elements significantly increase computational demands during test-time inference. We introduce PIR (Perplexity-based Importance Refinement), a principled framework that quantitatively evaluates the importance of each reasoning step based on its impact on answer prediction confidence. PIR systematically identifies and selectively prunes only low-importance functional steps while preserving all progressive reasoning components, creating optimized training data that maintains the integrity of the core solution path while reducing verbosity. Models fine-tuned on PIR-optimized data exhibit superior test-time scaling properties, generating more concise reasoning chains while achieving improved accuracy (+0.9\% to +6.6\%) with significantly reduced token usage (-3\% to -41\%) across challenging reasoning benchmarks (AIME, AMC, and GPQA Diamond). Our approach demonstrates strong generalizability across different model sizes, data sources, and token budgets, offering a practical solution for deploying reasoning-capable LLMs in scenarios where efficient test-time scaling, response time, and computational efficiency are valuable constraints. Code and dataset are available at the [LIMOPro GitHub repository.](https://github.com/GAIR-NLP/LIMOPro) Jiashuo Wang, Ruifeng Yuan, Chunpu Xu, Kaishuai Xu, Wenjie Li 0002, Pengfei Liu 0003 |
NeurIPS | 4 |
| 2025 | HICL: Hashtag-Driven In-Context Learning for Social Media Natural Language UnderstandingabstractNatural language understanding (NLU) is integral to various social media applications. However, the existing NLU models rely heavily on context for semantic learning, resulting in compromised performance when faced with short and noisy social media content. To address this issue, we leverage in-context learning (ICL), wherein language models learn to make inferences by conditioning on a handful of demonstrations to enrich the context and propose a novel hashtag-driven ICL (HICL) framework. Concretely, we pretrain a model #Encoder, which employs #hashtags (user-annotated topic labels) to drive BERT-based pretraining through contrastive learning. Our objective here is to enable #Encoder to gain the ability to incorporate topic-related semantic information, which allows it to retrieve topic-related posts to enrich contexts and enhance social media NLU with noisy contexts. To further integrate the retrieved context with the source text, we employ a gradient-based method to identify trigger terms useful in fusing information from both sources. For empirical studies, we collected 45 M tweets to set up an in-context NLU benchmark, and the experimental results on seven downstream tasks show that HICL substantially advances the previous state-of-the-art results. Furthermore, we conducted an extensive analysis and found that the following hold: 1) combining source input with a top-retrieved post from #Encoder is more effective than using semantically similar posts and 2) trigger words can largely benefit in merging context from the source and retrieved posts. Hanzhuo Tan, Chunpu Xu, Jing Li 0049, Yuqun Zhang, Zeyang Fang, Baohua Lai |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2024 | PopALM: Popularity-Aligned Language Models for Social Media Trendy Response PredictionabstractSocial media platforms are daily exhibiting millions of events. To preliminarily predict the mainstream public reaction to these events, we study trendy response prediction to automatically generate top-liked user replies to social media events. While previous works focus on generating responses without factoring in popularity, we propose Popularity-Aligned Language Models (PopALM) to distinguish responses liked by a larger audience through reinforcement learning. Recognizing the noisy labels from user “likes”, we tailor-make curriculum learning in proximal policy optimization (PPO) to help models capture the essential samples for easy-to-hard training. In experiments, we build a large-scale Weibo dataset for trendy response prediction, and its results show that PopALM can help boost the performance of advanced language models. Erxin Yu, Jing Li 0049, Chunpu Xu |
LREC/COLING | 3 |
| 2022 | Borrowing Human Senses: Comment-Aware Self-Training for Social Media Multimodal ClassificationabstractSocial media is daily creating massive multimedia content with paired image and text, presenting the pressing need to automate the vision and language understanding for various multimodal classification tasks.Compared to the commonly researched visual-lingual data, social media posts tend to exhibit more implicit image-text relations.To better glue the crossmodal semantics therein, we capture hinting features from user comments, which are retrieved via jointly leveraging visual and lingual similarity.Afterwards, the classification tasks are explored via self-training in a teacherstudent framework, motivated by the usually limited labeled data scales in existing benchmarks.Substantial experiments are conducted on four multimodal social media benchmarks for image-text relation classification, sarcasm detection, sentiment classification, and hate speech detection.The results show that our method further advances the performance of previous state-of-the-art models, which do not employ comment modeling or self-training. Chunpu Xu, Jing Li 0049 |
EMNLP | 1 |
| 2021 | Imagine, Reason and Write: Visual Storytelling with Graph Knowledge and Relational ReasoningabstractVisual storytelling is a task of creating a short story based on photo streams. Different from visual captions, stories contain not only factual descriptions, but also imaginary concepts that do not appear in the images. In this paper, we propose a novel imagine-reason-write generation framework (IRW) for visual storytelling, inspired by the logic of humans when they write the story. First, an imagine module is leveraged to learn the imaginative storyline explicitly, improving the coherence and reasonability of the generated story. Second, we employ a reason module to fully exploit the external knowledge (commonsense knowledge base) and task-specific knowledge (scene graph and event graph) with relational reasoning method based on the storyline. In this way, we can effectively capture the most informative commonsense and visual relationships among objects in images, which enhances the diversity and informativeness of the generated story. Finally, we integrate the imaginary concepts and relational knowledge to generate human-like story based on the original semantics of images. Extensive experiments on a benchmark dataset (i.e., VIST) demonstrate that the proposed IRW framework significantly outperforms the state-of-the-art methods across multiple evaluation metrics. Chunpu Xu, Min Yang 0007, Chengming Li 0004, Ying Shen 0001, Xiang Ao 0001, Ruifeng Xu 0001 |
AAAI | 1 |
| 2021 | #HowYouTagTweets: Learning User Hashtagging Preferences via Personalized Topic AttentionabstractMillions of hashtags are created on social media every day to cross-refer messages concerning similar topics.To help people find the topics they want to discuss, this paper characterizes a user's hashtagging preferences via predicting how likely they will post with a hashtag.It is hypothesized that one's interests in a hashtag are related to what they said before (user history) and the existing posts present the hashtag (hashtag contexts).These factors are married in the deep semantic space built with a pre-trained BERT and a neural topic model via joint training.In this way, user interests learned from the past can be customized to match future hashtags, which is beyond the capability of existing methods assuming unchanged hashtag semantics.Furthermore, we propose a novel personalized topic attention to capture salient contents to personalize hashtag contexts.Experiments on a large-scale Twitter dataset show that our model significantly outperforms the state-of-the-art recommendation approach without exploiting latent topics. 1 * Jing Li is the corresponding author.† This work was mainly conducted before Ziyan Jiang joined Amazon.1 Our dataset and code are publicly available in https://github.com/polyusmart/ Personalized-Hashtag-Preferences Sample tweets in H's hashtag contexts.Love thriller and mystery?Check out: URL She is writing the end for a long time.85-5 star review!Why not visit for Sunday share?Sample tweets in U 's user history.Darpocalypse is now available as an ebook!Read book 2 in the epic living dead series.Zombie Thriller Apocalypse Reader: Zeke is a skilled lover, so easy to fall for.But Yuji Zhang 0002, Chunpu Xu, Jing Li 0049, Ziyan Jiang, Baolin Peng |
EMNLP (1) | 3 |
| 2021 | Retrieval-enhanced adversarial training with dynamic memory-augmented attention for image paragraph captioning
Chunpu Xu, Min Yang 0007, Xiang Ao 0001, Ying Shen 0001, Ruifeng Xu 0001, Jinwen Tian |
Knowl. Based Syst. | 1 |
| 2020 | Interactive Dual Generative Adversarial Networks for Image CaptioningabstractImage captioning is usually built on either generation-based or retrieval-based approaches. Both ways have certain strengths but suffer from their own limitations. In this paper, we propose an Interactive Dual Generative Adversarial Network (IDGAN) for image captioning, which mutually combines the retrieval-based and generation-based methods to learn a better image captioning ensemble. IDGAN consists of two generators and two discriminators, where the generation- and retrieval-based generators mutually benefit from each other's complementary targets that are learned from two dual adversarial discriminators. Specifically, the generation- and retrieval-based generators provide improved synthetic and retrieved candidate captions with informative feedback signals from the two respective discriminators that are trained to distinguish the generated captions from the true captions and assign top rankings to true captions respectively, thus featuring the merits of both retrieval-based and generation-based approaches. Extensive experiments on MSCOCO dataset demonstrate that the proposed IDGAN model significantly outperforms the compared methods for image captioning. Junhao Liu 0001, Kai Wang 0036, Chunpu Xu, Zhou Zhao 0001, Ruifeng Xu 0001, Ying Shen 0001, Min Yang 0007 |
AAAI | 3 |
| 2020 | Interactive Key-Value Memory-augmented Attention for Image Paragraph CaptioningabstractImage paragraph captioning (IPC) aims to generate a fine-grained paragraph to describe the visual content of an image.Significant progress has been made by deep neural networks, in which the attention mechanism plays an essential role.However, conventional attention mechanisms tend to ignore the past alignment information, which often results in problems of repetitive captioning and incomplete captioning.In this paper, we propose an Interactive key-value Memoryaugmented Attention model for image Paragraph captioning (IMAP) to keep track of the attention history (salient objects coverage information) along with the update-chain of the decoder state and therefore avoid generating repetitive or incomplete image descriptions.In addition, we employ an adaptive attention mechanism to realize adaptive alignment from image regions to caption words, where an image region can be mapped to an arbitrary number of caption words while a caption word can also attend to an arbitrary number of image regions.Extensive experiments on a benchmark dataset (i.e., Stanford) demonstrate the effectiveness of our IMAP model. Chunpu Xu, Chengming Li 0004, Xiang Ao 0001, Min Yang 0007, Jinwen Tian |
COLING | 1 |
| 2019 | A Unified Generation-Retrieval Framework for Image CaptioningabstractRecent image captioning approaches are typically trained on generation-based or retrieval-based approaches. Both methods have their advantages but limited by the disadvantages. In this paper, we propose a Unified Generation-Retrieval framework for Image Captioning (UGRIC) by using adversarial learning. Different from previous methods, the proposed UGRIC model leverages the informative contents of N-best response candidates provided by the retrieval-based model to enhance the generation-based method. In addition, to further improve the informativeness of the generated caption, we employ copying mechanism to choose words from the retrieved candidate captions and put them into proper positions of the output sequence. Experiments on MSCOCO dataset demonstrate the effectiveness of the UGRIC model through various evaluation metrics.\footnoteCode and data are available at: \urlhttp://tinyurl.com/y6z2x6ho. Chunpu Xu, Wei Zhao 0033, Min Yang 0007, Xiang Ao 0001, Wangrong Cheng, Jinwen Tian |
CIKM | 1 |