VLDB 2026 Research / reviewers in the wild / expert
Zekun Wang 0001
dblp:181/5089-1
· DBLP profile ↗
13ranked-venue papers
4as first author
12since 2021 · last 2026
0000-0003-0151-5367ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 4 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Graph Reasoning Paradigm: Structured and Symbolic Reasoning with Topology-Aware Reinforcement Learning for Large Language ModelsabstractRunxuan Liu, Xianhao Ou, Xinyan Ma, Jiyuan Wang, Jiafeng Liang, Jiaqi Li, Tao He, Zheng Chu, Rongchuan Mu, Zekun Wang, Baoxin Wang, Dayong Wu, Ming Liu, Shijin Wang, Guoping Hu, Bing Qin. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Runxuan Liu, Xianhao Ou, Xinyan Ma, Jiafeng Liang, Jiaqi Li 0004, Tao He 0014, Rongchuan Mu, Zekun Wang 0001, Baoxin Wang, Dayong Wu, Ming Liu 0004, Shijin Wang 0001, Bing Qin 0001 |
ACL (1) | 10 |
| 2026 | PARIF: Pushing the Pareto Frontier of Instruction Following and Reasoning with Curriculum Reinforcement LearningabstractRongchuan Mu, Zexin Wang, Qianyu Wang, MingHua Ma, Zekun Wang, Ming Liu, Bing Qin. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Rongchuan Mu, Minghua Ma, Zekun Wang 0001, Ming Liu 0004, Bing Qin 0001 |
ACL (1) | 5 |
| 2026 | APSam: An Aggregating-Then-Pruning Sampler for Question-Conditional DenoisingabstractVideo question answering (VideoQA) necessitates simultaneous understanding of visual and linguistic information, requiring both in-depth analysis of individual modality features and the establishment of cross-modal correlations to achieve precise reasoning. However, VideoQA models often struggle with irrelevant temporal and spatial noise due to the dense events and concepts in real-world complex video contents. Previous works reduce noise by only sampling a fixed number of visual tokens at the patch level, overlooking the variation in the required granularities of features and quantities of visual cues across different question conditions. To address these, we propose an Aggregating-then-Pruning Sampler (APSam), which diversifies feature granularities and adaptively denoises on a per-question basis. Specifically, we propose a conditional token aggregator to obtain multi-granularity visual semantics by merging similar question-relevant tokens. Then, we propose a conditional token pruner, which restricts noise tokens through a variable-capacity receptive field determined by the inputs. Experimental results show that APSam achieves significant performance on three challenging complex VideoQA datasets,i.e., AGQAv2, NExT-QA, and STAR. Further analyses reveal that the APSam also exhibits high reasoning capability and interpretability. Jiafeng Liang, Shixin Jiang, Wei Tang 0015, Ning Wang 0020, Zekun Wang 0001, Xun Mao, Ming Liu 0004, Bing Qin 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | EffiVLM-BENCH: A Comprehensive Benchmark for Evaluating Training-Free Acceleration in Large Vision-Language ModelsabstractZekun Wang, MingHua Ma, Zexin Wang, Rongchuan Mu, Liping Shan, Ming Liu, Bing Qin. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Zekun Wang 0001, Minghua Ma, Rongchuan Mu, Liping Shan, Ming Liu 0004, Bing Qin 0001 |
ACL (1) | 1 |
| 2025 | CFSP: An Efficient Structured Pruning Framework for LLMs with Coarse-to-Fine Activation InformationabstractThe colossal parameters and computational overhead of Large Language Models (LLMs) challenge their real-world applications. Network pruning, which targets unstructured or structured sparsity by removing redundant parameters, has recently been explored for LLM acceleration. Existing LLM pruning works focus on unstructured pruning, which typically requires special hardware support for a practical speed-up. In contrast, structured pruning can reduce latency on general devices. However, it remains a challenge to perform structured pruning efficiently and maintain performance, especially at high sparsity ratios. To this end, we introduce an efficient structured pruning framework named CFSP, which leverages both Coarse (interblock) and Fine-grained (intrablock) activation information as an importance criterion to guide pruning. The pruning is highly efficient, as it only requires one forward pass to compute feature activations. Specifically, we first allocate the sparsity budget across blocks based on their importance and then retain important weights within each block. In addition, we introduce a recovery fine-tuning strategy that adaptively allocates training overhead based on coarse-grained importance to further improve performance. Experimental results demonstrate that CFSP outperforms existing methods on diverse models across various sparsity budgets. Our code will be available at https://github.com/wyxscir/CFSP. Yuxin Wang 0002, Minghua Ma, Zekun Wang 0001, Jingchang Chen, Liping Shan, Qing Yang 0033, Dongliang Xu, Ming Liu 0004, Bing Qin 0001 |
COLING | 3 |
| 2025 | Improved Diffusion-based Generative Model with Better Adversarial RobustnessabstractDiffusion Probabilistic Models (DPMs) have achieved significant success in generative tasks. However, their training and sampling processes suffer from the issue of distribution mismatch. During the denoising process, the input data distributions differ between the training and inference stages, potentially leading to inaccurate data generation. To obviate this, we analyze the training objective of DPMs and theoretically demonstrate that this mismatch can be alleviated through Distributionally Robust Optimization (DRO), which is equivalent to performing robustness-driven Adversarial Training (AT) on DPMs. Furthermore, for the recently proposed Consistency Model (CM), which distills the inference process of the DPM, we prove that its training objective also encounters the mismatch issue. Fortunately, this issue can be mitigated by AT as well. Based on these insights, we propose to conduct efficient AT on both DPM and CM. Finally, extensive empirical studies validate the effectiveness of AT in diffusion-based models. The code is available at https://github.com/kugwzk/AT_Diff. Zekun Wang 0001, Mingyang Yi, Shuchen Xue, Zhenguo Li, Ming Liu 0004, Bing Qin 0001, Zhiming Ma |
ICLR | 1 |
| 2025 | Exploring & exploiting high-order graph structure for sparse knowledge graph completion
Tao He 0014, Ming Liu 0004, Yixin Cao 0002, Zekun Wang 0001, Bing Qin 0001 |
Frontiers Comput. Sci. | 4 |
| 2024 | SmartTrim: Adaptive Tokens and Attention Pruning for Efficient Vision-Language ModelsabstractDespite achieving remarkable performance on various vision-language tasks, Transformer-based Vision-Language Models (VLMs) suffer from redundancy in inputs and parameters, significantly hampering their efficiency in real-world applications. Moreover, the degree of redundancy in token representations and model parameters, such as attention heads, varies significantly for different inputs. In light of the challenges, we propose SmartTrim, an adaptive acceleration framework for VLMs, which adjusts the computational overhead per instance. Specifically, we integrate lightweight modules into the original backbone to identify and prune redundant token representations and attention heads within each layer. Furthermore, we devise a self-distillation strategy to enhance the consistency between the predictions of the pruned model and its fully-capacity counterpart. Experimental results across various vision-language tasks consistently demonstrate that SmartTrim accelerates the original model by 2-3 times with minimal performance degradation, highlighting the effectiveness and efficiency compared to previous approaches. Code will be available at https://github.com/kugwzk/SmartTrim. Zekun Wang 0001, Jingchang Chen, Wangchunshu Zhou, Jiafeng Liang, Liping Shan, Ming Liu 0004, Dongliang Xu, Qing Yang 0033, Bing Qin 0001 |
LREC/COLING | 1 |
| 2024 | GUIDE: A Guideline-Guided Dataset for Instructional Video Comprehension
Jiafeng Liang, Shixin Jiang, Zekun Wang 0001, Haojie Pan, Zerui Chen, Ming Liu 0004, Ruiji Fu, Zhongyuan Wang 0006, Bing Qin 0001 |
IJCAI | 3 |
| 2024 | Divide-and-Conquer Meets Consensus: Unleashing the Power of Functions in Code GenerationabstractDespite recent progress made by large language models in code generation, they still struggle with programs that meet complex requirements. Recent work utilizes plan-and-solve decomposition to decrease the complexity and leverage self-tests to refine the generated program. Yet, planning deep-inside requirements in advance can be challenging, and the tests need to be accurate to accomplish self-improvement. To this end, we propose FunCoder, a code generation framework incorporating the divide-and-conquer strategy with functional consensus. Specifically, FunCoder recursively branches off sub-functions as smaller goals during code generation, represented by a tree hierarchy. These sub-functions are then composited to attain more complex objectives. Additionally, we designate functions via a consensus formed by identifying similarities in program behavior, mitigating error propagation. FunCoder outperforms state-of-the-art methods by +9.8% on average in HumanEval, MBPP, xCodeEval and MATH with GPT-3.5 and GPT-4. Moreover, our method demonstrates superiority on smaller models: With FunCoder, StableCode-3b surpasses GPT-3.5 by +18.6% and achieves 97.7% of GPT-4's performance on HumanEval. Further analysis reveals that our proposed dynamic function decomposition is capable of handling complex requirements, and the functional consensus prevails over self-testing in correctness evaluation. Jingchang Chen, Hongxuan Tang, Qianglong Chen, Zekun Wang 0001, Ming Liu 0004, Bing Qin 0001 |
NeurIPS | 5 |
| 2023 | GTR: A Grafting-Then-Reassembling Framework for Dynamic Scene Graph GenerationabstractDynamic scene graph generation aims to identify visual relationships (subject-predicate-object) in frames based on spatio-temporal contextual information in the video. Previous work implicitly models the spatio-temporal interaction simultaneously, which leads to entanglement of spatio-temporal contextual information. To this end, we propose a Grafting-Then-Reassembling framework (GTR), which explicitly extracts intra-frame spatial information and inter-frame temporal information in two separate stages to decouple spatio-temporal contextual information. Specifically, we first graft a static scene graph generation model to generate static visual relationships within frames. Then we propose the temporal dependency model to extract the temporal dependencies across frames, and explicitly reassemble static visual relationships into dynamic scene graphs. Experimental results show that GTR achieves the state-of-the-art performance on Action Genome dataset. Further analyses reveal that the reassembling stage is crucial to the success of our framework. Jiafeng Liang, Yuxin Wang 0002, Zekun Wang 0001, Ming Liu 0004, Ruiji Fu, Zhongyuan Wang 0006, Bing Qin 0001 |
IJCAI | 3 |
| 2022 | Distilled Dual-Encoder Model for Vision-Language UnderstandingabstractOn vision-language understanding (VLU) tasks, fusion-encoder vision-language models achieve superior results but sacrifice efficiency because of the simultaneous encoding of images and text.On the contrary, the dual-encoder model that separately encodes images and text has the advantage in efficiency, while failing on VLU tasks due to the lack of deep cross-modal interactions.To get the best of both worlds, we propose DIDE 1 , a framework that distills the knowledge of the fusion-encoder teacher model into the dual-encoder student model.Since the cross-modal interaction is the key to the superior performance of teacher model but is absent in the student model, we encourage the student not only to mimic the predictions of teacher, but also to calculate the cross-modal attention distributions and align with the teacher.Experimental results demonstrate that DIDE is competitive with the fusion-encoder teacher model in performance (only a 1% drop) while enjoying 4× faster inference.Further analyses reveal that the proposed cross-modal attention distillation is crucial to the success of our framework. Zekun Wang 0001, Wenhui Wang 0003, Ming Liu 0004, Bing Qin 0001, Furu Wei |
EMNLP | 1 |
| 2020 | Molweni: A Challenge Multiparty Dialogues-based Machine Reading Comprehension Dataset with Discourse StructureabstractResearch into the area of multiparty dialog has grown considerably over recent years.We present the Molweni dataset 1 , a machine reading comprehension (MRC) dataset with discourse structure built over multiparty dialog.Molweni's source samples from the Ubuntu Chat Corpus, including 10,000 dialogs comprising 88,303 utterances.We annotate 30,066 questions on this corpus, including both answerable and unanswerable questions.Molweni also uniquely contributes discourse dependency annotations in a modified Segmented Discourse Representation Theory (SDRT; (Asher et al., 2016)) style for all of its multiparty dialogs, contributing large-scale (78,245 annotated discourse relations) data to bear on the task of multiparty dialog discourse parsing.Our experiments show that Molweni is a challenging dataset for current MRC models: BERT-wwm, a current, strong SQuAD 2.0 performer, achieves only 67.7% F 1 on Molweni's questions, a 20+% significant drop as compared against its SQuAD 2.0 performance. Jiaqi Li 0004, Ming Liu 0004, Min-Yen Kan, Zekun Wang 0001, Wenqiang Lei, Ting Liu 0001, Bing Qin 0001 |
COLING | 5 |