VLDB 2026 Research / reviewers in the wild / expert
Yiming Ju
dblp:301/9136
· DBLP profile ↗
10ranked-venue papers
5as first author
10since 2021 · last 2026
0009-0000-0188-7385ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 5 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Scaling Towards the Information Boundary of Instructions through Data SynthesizingabstractHigh-quality instructions are crucial for aligning pretrained models to improve their performance on downstream tasks. Although current instruction datasets have reached tens of millions of samples, models finetuned on them may still struggle with complex instruction following and tasks in rare domains. This is primarily due to limited expansion in both "coverage" (coverage of task types and knowledge areas) and "depth" (instruction complexity) of the instruction set. To address this issue, we propose a systematic instruction data construction framework, which integrates a hierarchical labeling system, an informative seed selection algorithm, an evolutionary data synthesis process, and a model deficiency diagnosis with targeted data generation. These components form an iterative closed-loop to continuously enhance the coverage and depth of instruction data. Based on this framework, we construct Infinity Instruct Subject, a high-quality dataset containing approximately 1.5 million instructions. Experiments on multiple foundation models and benchmark tasks demonstrate its effectiveness in improving instruction-following capabilities. Further analyses suggest that InfinityInstruct-Subject shows enlarged coverage and depth compared to comparable synthesized instruction datasets. Yiming Ju, Tengfei Pan |
AAAI | 3 |
| 2025 | Beyond IID: Optimizing Instruction Finetuning from the Perspective of Instruction Interaction and DependencyabstractWith the availability of various instruction datasets, a pivotal challenge is how to effectively select and integrate these instructions to fine-tune large language models (LLMs). Previous research mainly focuses on selecting individual high-quality instructions. However, these works overlooked the joint interactions and dependencies between different categories of instructions, leading to suboptimal selection strategies. Moreover, the nature of these interaction patterns remains largely unexplored, let alone optimize the instruction set with regard to them. To fill these gaps, in this paper, we: (1) systemically investigate interaction and dependency patterns between different categories of instructions, (2) manage to optimize the instruction set concerning the interaction patterns using a linear programming-based method, and optimize the learning schema of SFT using an instruction dependency taxonomy guided curriculum learning. Experimental results across different LLMs demonstrate improved performance over strong baselines on widely adopted benchmarks. Yiming Ju, Tengfei Pan |
AAAI | 3 |
| 2025 | Exploiting Contextual Knowledge in LLMs through V-usable Information based Layer EnhancementabstractXiaowei Yuan, Zhao Yang, Ziyang Huang, Yequan Wang, Siqi Fan, Yiming Ju, Jun Zhao, Kang Liu. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Xiaowei Yuan, Zhao Yang 0004, Ziyang Huang 0005, Yequan Wang, Siqi Fan 0001, Yiming Ju, Jun Zhao 0001, Kang Liu 0001 |
ACL (1) | 6 |
| 2024 | Mitigating Training Imbalance in LLM Fine-Tuning via Selective Parameter MergingabstractSupervised fine-tuning (SFT) is crucial for adapting Large Language Models (LLMs) to specific tasks.In this work, we demonstrate that the order of training data can lead to significant training imbalances, potentially resulting in performance degradation.Consequently, we propose to mitigate this imbalance by merging SFT models fine-tuned with different data orders, thereby enhancing the overall effectiveness of SFT.Additionally, we introduce a novel technique, "parameter-selection merging," which outperforms traditional weightedaverage methods on five datasets.Further, through analysis and ablation studies, we validate the effectiveness of our method and identify the sources of performance improvements. Yiming Ju, Ziyi Ni, Xingrun Xing, Zhixiong Zeng, Siqi Fan 0001, Zheng Zhang 0006 |
EMNLP | 1 |
| 2024 | SpikeLM: Towards General Spike-Driven Language Modeling via Elastic Bi-Spiking MechanismsabstractTowards energy-efficient artificial intelligence similar to the human brain, the bio-inspired spiking neural networks (SNNs) have advantages of biological plausibility, event-driven sparsity, and binary activation. Recently, large-scale language models exhibit promising generalization capability, making it a valuable issue to explore more general spike-driven models. However, the binary spikes in existing SNNs fail to encode adequate semantic information, placing technological challenges for generalization. This work proposes the first fully spiking mechanism for general language tasks, including both discriminative and generative ones. Different from previous spikes with 0,1 levels, we propose a more general spike formulation with bi-directional, elastic amplitude, and elastic frequency encoding, while still maintaining the addition nature of SNNs. In a single time step, the spike is enhanced by direction and amplitude information; in spike frequency, a strategy to control spike firing rate is well designed. We plug this elastic bi-spiking mechanism in language modeling, named SpikeLM. It is the first time to handle general language tasks with fully spike-driven models, which achieve much higher accuracy than previously possible. SpikeLM also greatly bridges the performance gap between SNNs and ANNs in language modeling. Our code is available at https://github.com/Xingrun-Xing/SpikeLM. Xingrun Xing, Zheng Zhang 0006, Ziyi Ni, Shitao Xiao, Yiming Ju, Siqi Fan 0001, Yequan Wang |
ICML | 5 |
| 2024 | KLoB: a Benchmark for Assessing Knowledge Localization Methods in Language Models
Yiming Ju, Xingrun Xing, Zhixiong Zeng |
PRICAI (2) | 1 |
| 2024 | Explanation Guided Knowledge Distillation for Pre-trained Language Model CompressionabstractKnowledge distillation is widely used in pre-trained language model compression, which can transfer knowledge from a cumbersome model to a lightweight one. Though knowledge distillation based model compression has achieved promising performance, we observe that explanations between the teacher model and the student model are not consistent. We argue that the student model should study not only the predictions of the teacher model but also the internal reasoning process. To this end, we propose Explanation Guided Knowledge Distillation (EGKD) in this article, which utilizes explanations to represent the thinking process and improve knowledge distillation. To obtain explanations in our distillation framework, we select three typical explanation methods rooted in different mechanisms, namely gradient-based , perturbation-based , and feature selection methods. Then, to improve computational efficiency, we propose different optimization strategies to utilize the explanations obtained by these three different explanation methods, which could provide the student model with better learning guidance. Experimental results on GLUE demonstrate that leveraging explanations can improve the performance of the student model. Moreover, our EGKD could also be applied to model compression with different architectures. Zhao Yang 0004, Yuanzhe Zhang, Dianbo Sui, Yiming Ju, Jun Zhao 0001, Kang Liu 0001 |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 4 |
| 2022 | Logic Traps in Evaluating Attribution ScoresabstractModern deep learning models are notoriously opaque, which has motivated the development of methods for interpreting how deep models predict.This goal is usually approached with attribution method, which assesses the influence of features on model predictions.As an explanation method, the evaluation criteria of attribution methods is how accurately it reflects the actual reasoning process of the model (faithfulness).Meanwhile, since the reasoning process of deep models is inaccessible, researchers design various evaluation methods to demonstrate their arguments.However, some crucial logic traps in these evaluation methods are ignored in most works, causing inaccurate evaluation and unfair comparison.This paper systematically reviews existing methods for evaluating attribution scores and summarizes the logic traps in these methods.We further conduct experiments to demonstrate the existence of each logic trap.Through both theoretical and experimental analysis, we hope to increase attention on the inaccurate evaluation of attribution scores.Moreover, with this paper, we suggest stopping focusing on improving performance under unreliable evaluation systems and starting efforts on reducing the impact of proposed logic traps. Yiming Ju, Yuanzhe Zhang, Zhao Yang 0004, Zhongtao Jiang, Kang Liu 0001, Jun Zhao 0001 |
ACL (1) | 1 |
| 2022 | CMQA: A Dataset of Conditional Question Answering with Multiple-Span AnswersabstractForcing the answer of the Question Answering (QA) task to be a single text span might be restrictive since the answer can be multiple spans in the context. Moreover, we found that multi-span answers often appear with two characteristics when building the QA system for a real-world application. First, multi-span answers might be caused by users lacking domain knowledge and asking ambiguous questions, which makes the question need to be answered with conditions. Second, there might be hierarchical relations among multiple answer spans. Some recent span-extraction QA datasets include multi-span samples, but they only contain unconditional and parallel answers, which cannot be used to tackle this problem. To bridge the gap, we propose a new task: conditional question answering with hierarchical multi-span answers, where both the hierarchical relations and the conditions need to be extracted. Correspondingly, we introduce CMQA, a Conditional Multiple-span Chinese Question Answering dataset to study the new proposed task. The final release of CMQA consists of 7,861 QA pairs and 113,089 labels, where all samples contain multi-span answers, 50.4% of samples are conditional, and 56.6% of samples are hierarchical. CMQA can serve as a benchmark to study the new proposed task and help study building QA systems for real-world applications. The low performance of models drawn from related literature shows that the new proposed task is challenging for the community to solve. Yiming Ju, Weikang Wang 0005, Yuanzhe Zhang, Suncong Zheng, Kang Liu 0001, Jun Zhao 0001 |
COLING | 1 |
| 2021 | Enhancing Multiple-choice Machine Reading Comprehension by Punishing Illogical InterpretationsabstractMachine Reading Comprehension (MRC), which requires a machine to answer questions given the relevant documents, is an important way to test machines' ability to understand human language.Multiple-choice MRC is one of the most studied tasks in MRC due to the convenience of evaluation and the flexibility of answer format.Post-hoc interpretation aims to explain a trained model and reveal how the model arrives at the prediction.One of the most important interpretation forms is to attribute model decisions to input features.Based on post-hoc interpretation methods, we assess attributions of paragraphs in multiplechoice MRC and improve the model by punishing the illogical attributions.Our method can improve model performance without any external information and model structure change.Furthermore, we also analyze how and why such a self-training method works. Yiming Ju, Yuanzhe Zhang, Zhixing Tian, Kang Liu 0001, Xiaohuan Cao, Wenting Zhao 0006, Jun Zhao 0001 |
EMNLP (1) | 1 |