EDBT 2026 Demo / reviewers in the wild / expert
Mukai Li
dblp:279/3018
· DBLP profile ↗
13ranked-venue papers
2as first author
13since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 2 first-author · 13 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | EconProver: Towards More Economical Test-Time Scaling for Automated Theorem ProvingabstractMukai Li, Linfeng Song, Zhenwen Liang, Jiahao Xu, Shansan Gong, Qi Liu, Haitao Mi, Dong Yu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Mukai Li, Linfeng Song, Zhenwen Liang, Shansan Gong, Qi Liu 0049, Haitao Mi, Dong Yu 0001 |
ACL (1) | 1 |
| 2026 | OS-Sentinel: Towards Safety-Enhanced Mobile GUI Agents via Hybrid Validation in Realistic WorkflowsabstractQiushi Sun, Mukai Li, Zhoumianze Liu, Zhihui Xie, Fangzhi Xu, Zhangyue Yin, Kanzhi Cheng, Zehao Li, Zichen Ding, Qi Liu, Zhiyong Wu, Zhuosheng Zhang, Ben Kao, Lingpeng Kong. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Qiushi Sun, Mukai Li, Zhoumianze Liu, Zhihui Xie 0002, Fangzhi Xu, Zhangyue Yin, Kanzhi Cheng, Zichen Ding 0002, Qi Liu 0049, Zhiyong Wu 0003, Zhuosheng Zhang 0001, Ben Kao, Lingpeng Kong |
ACL (1) | 2 |
| 2025 | Design Choices for Extending the Context Length of Visual Language ModelsabstractVisual Language Models (VLMs) demonstrate impressive capabilities in processing multimodal inputs, yet applications such as visual agents, which require handling multiple images and high-resolution videos, demand enhanced long-range modeling.Moreover, existing opensource VLMs lack systematic exploration into extending their context length, and commercial models often provide limited details.To tackle this, we aim to establish an effective solution that enhances long context performance of VLMs while preserving their capacities in short context scenarios.Towards this goal, we make the best design choice through extensive experiment settings from data curation to context window extending and utilizing: ( 1) we analyze data sources and length distributions to construct ETVLM -a data recipe to balance the performance across scenarios; (2) we examine existing position extending methods, identify their limitations and propose M-RoPE++ as an enhanced approach; we also choose to solely instruction-tune the backbone with mixed-source data; (3) we discuss how to better utilize extended context windows and propose hybrid-resolution training.Built on the Qwen-VL series model, we propose GI-RAFFE, which is effectively extended to 128K lengths.Evaluated on extensive long context VLM benchmarks such as VideoMME and Viusal Haystacks, our GIRAFFE achieves stateof-the-art performance among similarly sized open-source long VLMs and is competitive with commercial model GPT-4V. 1 Mukai Li, Lei Li 0039, Shansan Gong, Qi Liu 0049 |
ACL (1) | 1 |
| 2025 | Scaling Diffusion Language Models via Adaptation from Autoregressive ModelsabstractDiffusion Language Models (DLMs) have emerged as a promising new paradigm for text generative modeling, potentially addressing limitations of autoregressive (AR) models. However, current DLMs have been studied at a smaller scale compared to their AR counterparts and lack fair comparison on language modeling benchmarks. Additionally, training diffusion models from scratch at scale remains challenging. Given the prevalence of open-source AR language models, we propose adapting these models to build text diffusion models. We demonstrate connections between AR and diffusion modeling objectives and introduce a simple continual pre-training approach for training diffusion models. Through systematic evaluation on language modeling, reasoning, and commonsense benchmarks, we show that we can convert AR models ranging from 127M to 7B parameters (GPT2 and LLaMA) into diffusion models DiffuGPT and DiffuLLaMA, using less than 200B tokens for training. Our experimental results reveal that these models outperform earlier DLMs and are competitive with their AR counterparts. We release a suite of DLMs (127M-355M-7B) capable of generating fluent text, performing in-context learning, filling in the middle without prompt re-ordering, and following instructions. Shansan Gong, Shivam Agarwal, Yizhe Zhang 0002, Jiacheng Ye, Mukai Li, Chenxin An, Peilin Zhao, Wei Bi, Jiawei Han 0001, Hao Peng 0009, Lingpeng Kong |
ICLR | 6 |
| 2024 | L-Eval: Instituting Standardized Evaluation for Long Context Language ModelsabstractChenxin An, Shansan Gong, Ming Zhong, Xingjian Zhao, Mukai Li, Jun Zhang, Lingpeng Kong, Xipeng Qiu. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Chenxin An, Shansan Gong, Ming Zhong 0005, Xingjian Zhao, Mukai Li, Jun Zhang 0003, Lingpeng Kong, Xipeng Qiu |
ACL (1) | 5 |
| 2024 | VLFeedback: A Large-Scale AI Feedback Dataset for Large Vision-Language Models AlignmentabstractLei Li, Zhihui Xie, Mukai Li, Shunian Chen, Peiyi Wang, Liang Chen, Yazheng Yang, Benyou Wang, Lingpeng Kong, Qi Liu. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Lei Li 0039, Zhihui Xie 0002, Mukai Li, Shunian Chen, Peiyi Wang, Liang Chen 0024, Yazheng Yang, Benyou Wang, Lingpeng Kong, Qi Liu 0049 |
EMNLP | 3 |
| 2023 | DiffuSeq: Sequence to Sequence Text Generation with Diffusion Models
Shansan Gong, Mukai Li, Jiangtao Feng, Zhiyong Wu 0003, Lingpeng Kong |
ICLR | 2 |
| 2023 | LAMM: Language-Assisted Multi-Modal Instruction-Tuning Dataset, Framework, and BenchmarkabstractLarge language models have emerged as a promising approach towards achieving general-purpose AI agents. The thriving open-source LLM community has greatly accelerated the development of agents that support human-machine dialogue interaction through natural language processing. However, human interaction with the world extends beyond only text as a modality, and other modalities such as vision are also crucial. Recent works on multi-modal large language models, such as GPT-4V and Bard, have demonstrated their effectiveness in handling visual modalities. However, the transparency of these works is limited and insufficient to support academic research. To the best of our knowledge, we present one of the very first open-source endeavors in the field, LAMM, encompassing a Language-Assisted Multi-Modal instruction tuning dataset, framework, and benchmark. Our aim is to establish LAMM as a growing ecosystem for training and evaluating MLLMs, with a specific focus on facilitating AI agents capable of bridging the gap between ideas and execution, thereby enabling seamless human-AI interaction. Our main contribution is three-fold: 1) We present a comprehensive dataset and benchmark, which cover a wide range of vision tasks for 2D and 3D vision. Extensive experiments validate the effectiveness of our dataset and benchmark. 2) We outline the detailed methodology of constructing multi-modal instruction tuning datasets and benchmarks for MLLMs, enabling rapid scaling and extension of MLLM research to diverse domains, tasks, and modalities. 3) We provide a primary but potential MLLM training framework optimized for modality extension. We also provide baseline models, comprehensive experimental observations, and analysis to accelerate future research. Our baseline model is trained within 24 A100 GPU hours, framework supports training with V100 and RTX3090 is available thanks to the open-source society. Codes and data are now available at https://openlamm.github.io. Zhenfei Yin, Jianjian Cao, Zhelun Shi, Dingning Liu, Mukai Li, Xiaoshui Huang, Zhiyong Wang 0001, Lu Sheng, Lei Bai 0001, Wanli Ouyang |
NeurIPS | 6 |
| 2022 | Prompt for Extraction? PAIE: Prompting Argument Interaction for Event Argument ExtractionabstractIn this paper, we propose an effective yet efficient model PAIE for both sentence-level and document-level Event Argument Extraction (EAE), which also generalizes well when there is a lack of training data.On the one hand, PAIE utilizes prompt tuning for extractive objectives to take the best advantages of Pre-trained Language Models (PLMs).It introduces two span selectors based on the prompt to select start/end tokens among input texts for each role.On the other hand, it captures argument interactions via multi-role prompts and conducts joint optimization with optimal span assignments via a bipartite matching loss.Also, with a flexible prompt design, PAIE can extract multiple arguments with the same role instead of conventional heuristic threshold tuning.We have conducted extensive experiments on three benchmarks, including both sentenceand document-level EAE.The results present promising improvements from PAIE (3.5% and 2.3% F1 gains in average on three benchmarks, for PAIE-base and PAIE-large respectively).Further analysis demonstrates the efficiency, generalization to few-shot settings, and effectiveness of different extractive prompt tuning strategies.Our code is available at https: //github.com/mayubo2333/PAIE. Yubo Ma, Yixin Cao 0002, Mukai Li, Meiqi Chen 0001, Kun Wang 0056 |
ACL (1) | 4 |
| 2022 | ERGO: Event Relational Graph Transformer for Document-level Event Causality IdentificationabstractDocument-level Event Causality Identification (DECI) aims to identify event-event causal relations in a document. Existing works usually build an event graph for global reasoning across multiple sentences. However, the edges between events have to be carefully designed through heuristic rules or external tools. In this paper, we propose a novel Event Relational Graph TransfOrmer (ERGO) framework for DECI, to ease the graph construction and improve it over the noisy edge issue. Different from conventional event graphs, we define a pair of events as a node and build a complete event relational graph without any prior knowledge or tools. This naturally formulates DECI as a node classification problem, and thus we capture the causation transitivity among event pairs via a graph transformer. Furthermore, we design a criss-cross constraint and an adaptive focal loss for the imbalanced classification, to alleviate the issues of false positives and false negatives. Extensive experiments on two benchmark datasets show that ERGO greatly outperforms previous state-of-the-art (SOTA) methods (12.8% F1 gains on average). Meiqi Chen 0001, Yixin Cao 0002, Kunquan Deng, Mukai Li, Kun Wang 0056, Yan Zhang 0004 |
COLING | 4 |
| 2021 | Hidden Killer: Invisible Textual Backdoor Attacks with Syntactic TriggerabstractFanchao Qi, Mukai Li, Yangyi Chen, Zhengyan Zhang, Zhiyuan Liu, Yasheng Wang, Maosong Sun. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Fanchao Qi, Mukai Li, Yangyi Chen, Zhengyan Zhang, Zhiyuan Liu 0001, Yasheng Wang, Maosong Sun 0001 |
ACL/IJCNLP (1) | 2 |
| 2021 | ONION: A Simple and Effective Defense Against Textual Backdoor AttacksabstractBackdoor attacks are a kind of emergent training-time threat to deep neural networks (DNNs).They can manipulate the output of DNNs and possess high insidiousness.In the field of natural language processing, some attack methods have been proposed and achieve very high attack success rates on multiple popular models.Nevertheless, there are few studies on defending against textual backdoor attacks.In this paper, we propose a simple and effective textual backdoor defense named ONION, which is based on outlier word detection and, to the best of our knowledge, is the first method that can handle all the textual backdoor attack situations.Experiments demonstrate the effectiveness of our model in defending BiLSTM and BERT against five different backdoor attacks.All the code and data of this paper can be obtained at https: //github.com/thunlp/ONION. Fanchao Qi, Yangyi Chen, Mukai Li, Yuan Yao 0013, Zhiyuan Liu 0001, Maosong Sun 0001 |
EMNLP (1) | 3 |
| 2021 | Mind the Style of Text! Adversarial and Backdoor Attacks Based on Text Style TransferabstractAdversarial attacks and backdoor attacks are two common security threats that hang over deep learning.Both of them harness taskirrelevant features of data in their implementation.Text style is a feature that is naturally irrelevant to most NLP tasks, and thus suitable for adversarial and backdoor attacks.In this paper, we make the first attempt to conduct adversarial and backdoor attacks based on text style transfer, which is aimed at altering the style of a sentence while preserving its meaning.We design an adversarial attack method and a backdoor attack method, and conduct extensive experiments to evaluate them.Experimental results show that popular NLP models are vulnerable to both adversarial and backdoor attacks based on text style transfer-the attack success rates can exceed 90% without much effort.It reflects the limited ability of NLP models to handle the feature of text style that has not been widely realized.In addition, the style transfer-based adversarial and backdoor attack methods show superiority to baselines in many aspects.All the code and data of this paper can be obtained at https:// github.com/thunlp/StyleAttack. Fanchao Qi, Yangyi Chen, Xurui Zhang, Mukai Li, Zhiyuan Liu 0001, Maosong Sun 0001 |
EMNLP (1) | 4 |