EDBT 2026 Demo / reviewers in the wild / expert
Xiaojun Meng
dblp:79/9935
· DBLP profile ↗
26ranked-venue papers
6as first author
18since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 8 since 2021Human-computer interaction and ubiquitous computing · 8 · 5 first-authorDatabases, data management, data science and information retrieval · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Difficulty Is Not Enough: Curriculum Learning for LLMs Fine-tuning Must Consider UtilityabstractFine-tuning plays an essential role in improving the performance of large language models (LLMs) on specific tasks. A central challenge lies in designing data-efficient strategy to achieve better fine-tuning performance. Curriculum learning, which organizes data from easy to hard, has become a widely adopted technique in LLMs training. However, existing methods for curriculum learning focus only on the difficulty of samples, while neglecting their contribution to improving model performance, making them vulnerable when applied to fine-tuning LLMs. To address this, we propose Difficulty-Utility Curriculum Learning (DUCL), a curriculum learning framework that jointly considers difficulty and utility. DUCL introduces a novel scoring method, Difficulty-Utility Evaluation (DUE), and a soft scheduling strategy called Window Ordering, which together promote efficient and effective fine-tuning. Our method not only improves convergence and final performance with negligible computational overhead, but is also broadly applicable across a wide range of tasks, making it a practical and scalable solution for LLMs fine-tuning. Zishang Jiang, Jinyi Han, Tingyun Li, Sihang Jiang 0001, Xiaojun Meng, Jiansheng Wei, Jiaqing Liang, Yanghua Xiao |
AAAI | 6 |
| 2026 | MMIFEvol: Towards Evolutionary Multimodal Instruction FollowingabstractMultimodal Instruction Following serves as a fundamental capability of multimodal language models, involving accurate comprehension and execution of user-provided instructions. However, existing multimodal instruction-following datasets and benchmarks face the shortcomings outlined below: (a) Lack of Difficulty Stratification, they collect diverse instruction categories but neglect the stratification of difficulty levels across these categories, which leads to overlap, bias, and low interpretability. (b) Lack of Fine-Grained Metrics, they conflate the model's ability to ``solve tasks" and ``follow constraints" into a single metric, which fails to accurately reflect its instruction-following capability. (c) Lack of Multi-Task Instructions, they overlook the fact that real-world user instructions often consist of multiple combined tasks. This paper proposes MMIFEvol, a framework for multimodal instruction evolving and benchmarking. First, we define the essential components of a carefully curated multimodal instruction set and establish corresponding difficulty levels, based on which we synthesize diverse instruction data. Next, we decouple the evaluation criteria for the instruction following into three different metrics to construct a high-quality benchmark and assess existing models. Experimental results demonstrate that current models still struggle with following complex instructions, while fine-tuning using MMIFEvol data effectively improves models' responsiveness to multimodal instructions. Sihang Jiang 0001, Xiangru Zhu, Yuyan Chen, Xiaojun Meng, Jiansheng Wei, Yanghua Xiao |
AAAI | 5 |
| 2026 | Metaphor Reasoning is Meta-reasoningabstractQianyu He, Junting Lu, Yikai Zhang, Siyu Yuan, Xiaojun Meng, Jiansheng Wei, Jiaqing Liang, Yanghua Xiao. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Qianyu He, Junting Lu, Yikai Zhang 0004, Xiaojun Meng, Jiansheng Wei, Jiaqing Liang, Yanghua Xiao |
ACL (1) | 5 |
| 2026 | Immediate Inference: The Missing Foundation in Large Language Model Logical ReasoningabstractSihang Jiang, Zhiyu Lu, Keyi Wang, Jiaqing Liang, Yanghua Xiao, Xiaojun Meng, Jiansheng Wei. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Sihang Jiang 0001, Jiaqing Liang, Yanghua Xiao, Xiaojun Meng, Jiansheng Wei |
ACL (1) | 6 |
| 2026 | The GaoYao Benchmark: A Comprehensive Framework for Evaluating Multilingual and Multicultural Abilities of Large Language ModelsabstractYilun Liu, Chunguang Zhao, Mengyao Piao, Lingqi Miao, Shimin Tao, Minggui HE, Chenxin Liu, Zhang Li, Mahongxia, Jiaxin Guo, Chen Liu, Liqun Deng, Jiansheng Wei, Xiaojun Meng, Fanyi Du, Daimeng Wei, Yanghua Xiao. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Yilun Liu 0001, Chunguang Zhao, Mengyao Piao, Lingqi Miao, Shimin Tao, Minggui He, Chenxin Liu, Hongxia Ma, Liqun Deng, Jiansheng Wei, Xiaojun Meng, Fanyi Du, Daimeng Wei, Yanghua Xiao |
ACL (1) | 14 |
| 2026 | Code LLMs Still Fall Short of Top Programmers: Evaluating Algorithmic Code Generation Through Computational ThinkingabstractEvaluating the coding capabilities of models through algorithmic code generation is challenging, as it requires deep problem understanding and complex algorithm design. Current benchmarks suffer from a narrow focus on final execution results (such as pass@k), neglecting the crucial reasoning and problem-solving processes inherent in code generation. To address this limitation, we introduce a multi-phase algorithmic code generation benchmark, MUPA, structured around human computational thinking. MUPA dissects the evaluation into four distinct phases: example understanding, algorithm selection, solution description, and code generation. This framework facilitates a comprehensive assessment by providing insights into the model's intermediate problem-solving steps, rather than just the final code. We manually curated 197 high-quality competitive programming problems from Codeforces. Utilizing an LLM-as-a-judge paradigm with specialized prompts, our rigorous evaluation of several existing code generation LLMs reveals significant across-the-board challenges. Notably, we establish a positive correlation, indicating that proficiency in an earlier phase directly impacts performance in subsequent phases, underscoring the interdependency of these algorithmic skills. The benchmark is publicly available at https://github.com/cheniison/MUPA. Shisong Chen, Ziyu Zhou 0019, Zhixu Li, Yanghua Xiao, Xin Lin 0001, Xiaojun Meng, Jiansheng Wei, Kuien Liu |
WSDM | 8 |
| 2025 | Friends-MMC: A Dataset for Multi-modal Multi-party Conversation UnderstandingabstractMulti-modal multi-party conversation (MMC) is a less studied yet important topic of research due to that it well fits real-world scenarios and thus potentially has more widely-used applications. Compared with the traditional multi-modal conversations, MMC requires stronger character-centered understanding abilities as there are many interlocutors appearing in both the visual and textual context. To facilitate the study of this problem, we present Friends-MMC in this paper, an MMC dataset that contains 24,000+ unique utterances paired with video context. To explore the character-centered understanding of the dialogue, we also annotate the speaker of each utterance, the names and bounding bboxes of faces that appear in the video. Based on this Friends-MMC dataset, we further study two fundamental MMC tasks: conversation speaker identification and conversation response prediction, both of which have the multi-party nature with the video or image as visual context. For conversation speaker identification, we demonstrate the inefficiencies of existing methods such as pre-trained models, and propose a simple yet effective baseline method that leverages an optimization solver to utilize the context of two modalities to achieve better performance. For conversation response prediction, we fine-tune generative dialogue models on Friend-MMC, and analyze the benefits of speaker information. The code and dataset will be publicly available, and thus we call for more attention on modelling speaker information when understanding conversations. Yueqian Wang, Xiaojun Meng, Yuxuan Wang 0004, Jianxin Liang, Qun Liu 0001, Dongyan Zhao 0001 |
AAAI | 2 |
| 2025 | Sparsing Law: Towards Large Language Models with Greater Activation SparsityabstractActivation sparsity denotes the existence of substantial weakly-contributed neurons within feed-forward networks of large language models (LLMs), providing wide potential benefits such as computation acceleration. However, existing works lack thorough quantitative studies on this useful property, in terms of both its measurement and influential factors. In this paper, we address three underexplored research questions: (1) How can activation sparsity be measured more accurately? (2) How is activation sparsity affected by the model architecture and training process? (3) How can we build a more sparsely activated and efficient LLM? Specifically, we develop a generalizable and performance-friendly metric, named CETT-PPL-1%, to measure activation sparsity. Based on CETT-PPL-1%, we quantitatively study the influence of various factors and observe several important phenomena, such as the convergent power-law relationship between sparsity and training data amount, the higher competence of ReLU activation than mainstream SiLU activation, the potential sparsity merit of a small width-depth ratio, and the scale insensitivity of activation sparsity. Finally, we provide implications for building sparse and effective LLMs, and demonstrate the reliability of our findings by training a 2.4B model with a sparsity ratio of 93.52%, showing 4.1$\times$ speedup compared with its dense version. The codes and checkpoints are available at https://github.com/thunlp/SparsingLaw/. Yuqi Luo, Xu Han 0007, Yingfa Chen, Chaojun Xiao, Xiaojun Meng, Liqun Deng, Jiansheng Wei, Zhiyuan Liu 0001, Maosong Sun 0001 |
ICML | 6 |
| 2025 | ReasVQA: Advancing VideoQA with Imperfect Reasoning ProcessabstractJianxin Liang, Xiaojun Meng, Huishuai Zhang, Yueqian Wang, Jiansheng Wei, Dongyan Zhao. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Jianxin Liang, Xiaojun Meng, Huishuai Zhang, Yueqian Wang, Jiansheng Wei, Dongyan Zhao 0001 |
NAACL (Long Papers) | 2 |
| 2025 | End-to-End VideoQA with Frame Scoring Mechanisms and Adaptive Sampling
Jianxin Liang, Xiaojun Meng, Yueqian Wang, Chang Liu 0076, Qun Liu 0001, Dongyan Zhao 0001 |
NLPCC (2) | 2 |
| 2025 | Unsupervised Domain Adaptive Visual Question Answering in the Era of Multi-Modal Large Language ModelsabstractUnsupervised domain adaptation (UDA) for visual question answering (VQA) has attracted research interest. However, with Multi-modal Large Language Models (MLLMs) showing great performance on VQA datasets, UDA for VQA based on MLLMs remains unexplored. To fill this gap, we propose the first systematic approach to Unsupervised Domain Adaptation VQA based on MLLMs (UDAM). First, we introduce semantic context feature alignment and domain query feature alignment, which utilize a single token embedding for each modality to capture contextual domain information from unimodal inputs and conduct coarsegrained feature alignment on it, thus alleviating domain shifts in the unimodal feature space. Second, we propose the novel semantics-guided query feature alignment, which differentiates important domain-specific queries from learnable query outputs and conducts fine-grained feature alignment controlled by a semantics-guided weight map to reduce domain shifts in the cross-modal feature space. Third, we devise a pair-wise domain-aware prompt strategy, which aids UDA by prompting MLLMs to discern the commonality of tasks and the distinctiveness of domains in multi-modal inputs. Extensive experiments demonstrate UDAM's effectiveness in adapting MLLMs to unlabeled new domains. Weixi Weng, Xiaojun Meng, Jieming Zhu, Qun Liu 0001, Chun Yuan 0003 |
WACV | 3 |
| 2025 | Enhancing inter-sentence coherence of extractive summarization with multitask learning
Renlong Jie, Xiaojun Meng, Lifeng Shang, Xin Jiang 0002, Qun Liu 0001 |
J. Intell. Inf. Syst. | 2 |
| 2024 | Unsupervised Extractive Summarization with Learnable Length Control StrategiesabstractUnsupervised extractive summarization is an important technique in information extraction and retrieval. Compared with supervised method, it does not require high-quality human-labelled summaries for training and thus can be easily applied for documents with different types, domains or languages. Most of existing unsupervised methods including TextRank and PACSUM rely on graph-based ranking on sentence centrality. However, this scorer can not be directly applied in end-to-end training, and the positional-related prior assumption is often needed for achieving good summaries. In addition, less attention is paid to length-controllable extractor, where users can decide to summarize texts under particular length constraint. This paper introduces an unsupervised extractive summarization model based on a siamese network, for which we develop a trainable bidirectional prediction objective between the selected summary and the original document. Different from the centrality-based ranking methods, our extractive scorer can be trained in an end-to-end manner, with no other requirement of positional assumption. In addition, we introduce a differentiable length control module by approximating 0-1 knapsack solver for end-to-end length-controllable extracting. Experiments show that our unsupervised method largely outperforms the centrality-based baseline using a same sentence encoder. In terms of length control ability, via our trainable knapsack module, the performance consistently outperforms the strong baseline without utilizing end-to-end training. Human evaluation further evidences that our method performs the best among baselines in terms of relevance and consistency. Renlong Jie, Xiaojun Meng, Xin Jiang 0002, Qun Liu 0001 |
AAAI | 2 |
| 2023 | Wukong-Reader: Multi-modal Pre-training for Fine-grained Visual Document UnderstandingabstractHaoli Bai, Zhiguang Liu, Xiaojun Meng, Li Wentao, Shuang Liu, Yifeng Luo, Nian Xie, Rongfu Zheng, Liangwei Wang, Lu Hou, Jiansheng Wei, Xin Jiang, Qun Liu. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Haoli Bai, Xiaojun Meng, Yifeng Luo, Nian Xie, Rongfu Zheng, Liangwei Wang 0004, Lu Hou 0002, Jiansheng Wei, Xin Jiang 0002, Qun Liu 0001 |
ACL (1) | 3 |
| 2023 | Lexicon-injected Semantic Parsing for Task-Oriented DialogabstractRecently, semantic parsing using hierarchical representations for dialog systems has captured substantial attention. Task-Oriented Parse (TOP), a tree representation with intents and slots as labels of nested tree nodes, has been proposed for parsing user utterances. Previous TOP parsing methods are limited on leveraging lexicon resources, which are often used to guide the real dialog system. To mitigate this issue, we first propose a novel span-splitting representation for span-based parser that outperforms existing methods. Then we present a novel lexicon-injected semantic parser, which collects slot labels of tree representation as a lexicon, and injects lexical features to the span representation of parser. An additional slot disambiguation technique is involved to remove inappropriate span match occurrences from the lexicon. Experiments show that our best parser produces a new state-of-the-art result (87.62%) on the TOP dataset, and also confirm the effectiveness of our proposed lexicon-injected parser and slot disambiguation model. Xiaojun Meng, Wenlin Dai, Yasheng Wang, Baojun Wang, Zhiyong Wu 0003, Xin Jiang 0002, Qun Liu 0001 |
ICASSP | 1 |
| 2023 | Learning Summary-Worthy Visual Representation for Abstractive Summarization in VideoabstractMultimodal abstractive summarization for videos (MAS) requires generating a concise textual summary to describe the highlights of a video according to multimodal resources, in our case, the video content and its transcript. Inspired by the success of the large-scale generative pre-trained language model (GPLM) in generating high-quality textual content (e.g., summary), recent MAS methods have proposed to adapt the GPLM to this task by equipping it with the visual information, which is often obtained through a general-purpose visual feature extractor. However, the generally extracted visual features may overlook some summary-worthy visual information, which impedes model performance. In this work, we propose a novel approach to learning the summary-worthy visual representation that facilitates abstractive summarization. Our method exploits the summary-worthy information from both the cross-modal transcript data and the knowledge that distills from the pseudo summary. Extensive experiments on three public multimodal datasets show that our method outperforms all competing baselines. Furthermore, with the advantages of summary-worthy visual information, our model can have a significant improvement on small datasets or even datasets with limited training data. Zenan Xu, Xiaojun Meng, Yasheng Wang, Qinliang Su, Zexuan Qiu, Xin Jiang 0002, Qun Liu 0001 |
IJCAI | 2 |
| 2022 | UniMS: A Unified Framework for Multimodal Summarization with Knowledge DistillationabstractWith the rapid increase of multimedia data, a large body of literature has emerged to work on multimodal summarization, the majority of which target at refining salient information from textual and image modalities to output a pictorial summary with the most relevant images. Existing methods mostly focus on either extractive or abstractive summarization and rely on the presence and quality of image captions to build image references. We are the first to propose a Unified framework for Multimodal Summarization grounding on BART, UniMS, that integrates extractive and abstractive objectives, as well as selecting the image output. Specially, we adopt knowledge distillation from a vision-language pretrained model to improve image selection, which avoids any requirement on the existence and quality of image captions. Besides, we introduce a visual guided decoder to better integrate textual and visual modalities in guiding abstractive text generation. Results show that our best model achieves a new state-of-the-art result on a large-scale benchmark dataset. The newly involved extractive objective as well as the knowledge distillation technique are proven to bring a noticeable improvement to the multimodal summarization task. Zhengkun Zhang, Xiaojun Meng, Yasheng Wang, Xin Jiang 0002, Qun Liu 0001, Zhenglu Yang |
AAAI | 2 |
| 2022 | Wukong: A 100 Million Large-scale Chinese Cross-modal Pre-training BenchmarkabstractVision-Language Pre-training (VLP) models have shown remarkable performance on various downstream tasks. Their success heavily relies on the scale of pre-trained cross-modal datasets. However, the lack of large-scale datasets and benchmarks in Chinese hinders the development of Chinese VLP models and broader multilingual applications. In this work, we release a large-scale Chinese cross-modal dataset named Wukong, which contains 100 million Chinese image-text pairs collected from the web. Wukong aims to benchmark different multi-modal pre-training methods to facilitate the VLP research and community development. Furthermore, we release a group of models pre-trained with various image encoders (ViT-B/ViT-L/SwinT) and also apply advanced pre-training techniques into VLP such as locked-image text tuning, token-wise similarity in contrastive learning, and reduced-token interaction. Extensive experiments and a benchmarking of different downstream tasks including a new largest human-verified image-text test dataset are also provided. Experiments show that Wukong can serve as a promising Chinese pre-training dataset and benchmark for different cross-modal learning methods. For the zero-shot image classification task on 10 datasets, $Wukong_\text{ViT-L}$ achieves an average accuracy of 73.03%. For the image-text retrieval task, it achieves a mean recall of 71.6% on AIC-ICC which is 12.9% higher than WenLan 2.0. Also, our Wukong models are benchmarked on downstream tasks with other variants on multiple datasets, e.g., Flickr8K-CN, Flickr-30K-CN, COCO-CN, et al. More information can be referred to https://wukong-dataset.github.io/wukong-dataset/. Jiaxi Gu, Xiaojun Meng, Guansong Lu, Lu Hou 0002, Niu Minzhe, Xiaodan Liang, Lewei Yao, Runhui Huang, Wei Zhang 0196, Xin Jiang 0002, Chunjing Xu, Hang Xu 0004 |
NeurIPS | 2 |
| 2019 | ScaffoMapping: Assisting Concept Mapping for Video Learners
Shan Zhang 0006, Xiaojun Meng, Can Liu 0003, Shengdong Zhao 0001, Vibhor Sehgal, Morten Fjeld |
INTERACT (2) | 2 |
| 2017 | ReTool: Interactive Microtask and Workflow Design through DemonstrationabstractIn addition to simple form filling, there is an increasing need for crowdsourcing workers to perform freeform interactions directly on content in microtask crowdsourcing (e.g. proofreading articles or specifying object boundary in an image). Such microtasks are often organized within well-designed workflows to optimize task quality and workload distribution. However, designing and implementing the interface and workflow for such microtasks is challenging because it typically requires programming knowledge and tedious manual effort. We present ReTool, a web-based tool for requesters to design and publish interactive microtasks and workflows by demonstrating the microtasks for text and image content. We evaluated ReTool against a task-design tool from a popular crowdsourcing platform and showed the advantages of ReTool over the existing approach. Xiaojun Meng, Shengdong Zhao 0001, Morten Fjeld |
CHI | 2 |
| 2017 | NexP: A Beginner Friendly Toolkit for Designing and Conducting Controlled Experiments
Xiaojun Meng, Pin Sym Foong, Simon T. Perrault, Shengdong Zhao 0001 |
INTERACT (3) | 1 |
| 2016 | 5-Step Approach to Designing Controlled ExperimentsabstractControlled experiment, an approach that has been adopted from research methods in psychology, is now widely used in HCI. Design an effective controlled experiment is not necessarily easy, even for experienced researchers. Existing controlled experiment designing tools focused more on the process of designing, but they are often not intuitive and guided enough for less experienced users to use. In this demo paper, we introduce NexP (Next Experiment Tool), a web-based open-source tool for designing controlled experiments. NexP introduces a 5-step approach to guide users through the experimental design process, and helped them to better understand the experimental design process, making it both a useful and educational tool. Xiaojun Meng, Pin Sym Foong, Simon T. Perrault, Shengdong Zhao 0001 |
AVI | 1 |
| 2016 | HyNote: Integrated Concept Mapping and NotetakingabstractNotes can be taken in a linear or nonlinear way. Previous work suggests that nonlinear note taking is advantageous in terms of sense-making and long-term recall. However, previous studies also reveal that the combination of divided attention and time pressure make realtime notetaking a challenge. In this paper, we propose a new hybrid workflow inheriting advantages from both linear and nonlinear notetaking approaches. Our resulting HyNote (Hybrid Notetaking) system uses statistical parsing of linear raw notes to facilitate concepts mapping, allowing users to smoothly switch between linear and nonlinear approaches with low effort and time costs. Results from our preliminary study of HyNote show that users can easily map concepts in realtime and achieve superior understanding of lecture contents in a video learning task compared with using the traditional linear Notepad application. Xiaojun Meng, Shengdong Zhao 0001, Darren Edge |
AVI | 1 |
| 2015 | CozyMaps: Real-time Collaboration on a Shared Map with Multiple DisplaysabstractWith the use of several tablet devices and a shared large display, CozyMaps is a multi-display system that supports real-time collocated collaboration on a shared map. This paper builds on existing works and introduces rich user interactions by proposing awareness, notification, and view sharing techniques, to enable seamless information sharing and integration in map-based applications. Based on our exploratory study, we demonstrated that participants are satisfied with these new proposed interactions. We found that view sharing techniques should be location-focused rather than user-focused. Our results provide implications for the design of interactive techniques in collaborative multi-display map systems. Kelvin Cheng 0002, Liang He 0005, Xiaojun Meng, David A. Shamma, Anbarasan Thangapalam |
MobileHCI | 3 |
| 2014 | WADE: simplified GUI add-on development for third-party softwareabstractWe present the WADE Integrated Development Environment (IDE), which simplifies interface and functionality modification of existing third-party software without access to source code. WADE clones the Graphical User Interface (GUI) of a host program through dynamic-link library (DLL) injection, enabling modifications to (1) the GUI in a WYSIWYG fashion and (2) software functionality. We compare WADE with an alternative state-of-the-art runtime toolkit overloading approach in a user-study, whose results demonstrate that WADE significantly simplifies the task of GUI-based add-on development. Xiaojun Meng, Shengdong Zhao 0001, James R. Eagan, Subramanian Ramanathan |
CHI | 1 |
| 2010 | Task-oriented BOM conversion and view mappingabstractTo increase the productivity as much as possible and accommodate to the requirements of mass production for the production of multi-variety, multi-type, multi-batch and changeable lot-size, to rapidly respond the urgent task of important production, the technology of task-oriented BOM conversion and view mapping based on UDS (Uniform Data Source) is presented. The method to unify the data source by classifying the parts and collecting the common parts was established. The theory of task-oriented BOM conversion and view mapping is analyzed. The step of implementing task-oriented BOM view mapping is given. Finally, an interpretive, qualitative case was conducted. Xiaojun Meng, Ruxin Ning, Aimin Wang 0002 |
CSCWD | 1 |