Jiansheng Wei

dblp:20/10032 · DBLP profile ↗
← Back
22ranked-venue papers
7as first author
17since 2021 · last 2026
0000-0001-8518-4088ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 1 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 7 since 2021Systems, architecture and hardware · 4 · 4 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Difficulty Is Not Enough: Curriculum Learning for LLMs Fine-tuning Must Consider Utility
abstract
Fine-tuning plays an essential role in improving the performance of large language models (LLMs) on specific tasks. A central challenge lies in designing data-efficient strategy to achieve better fine-tuning performance. Curriculum learning, which organizes data from easy to hard, has become a widely adopted technique in LLMs training. However, existing methods for curriculum learning focus only on the difficulty of samples, while neglecting their contribution to improving model performance, making them vulnerable when applied to fine-tuning LLMs. To address this, we propose Difficulty-Utility Curriculum Learning (DUCL), a curriculum learning framework that jointly considers difficulty and utility. DUCL introduces a novel scoring method, Difficulty-Utility Evaluation (DUE), and a soft scheduling strategy called Window Ordering, which together promote efficient and effective fine-tuning. Our method not only improves convergence and final performance with negligible computational overhead, but is also broadly applicable across a wide range of tasks, making it a practical and scalable solution for LLMs fine-tuning.
Zishang Jiang, Jinyi Han, Tingyun Li, Sihang Jiang 0001, Xiaojun Meng, Jiansheng Wei, Jiaqing Liang, Yanghua Xiao
AAAI7
2026 MMIFEvol: Towards Evolutionary Multimodal Instruction Following
abstract
Multimodal Instruction Following serves as a fundamental capability of multimodal language models, involving accurate comprehension and execution of user-provided instructions. However, existing multimodal instruction-following datasets and benchmarks face the shortcomings outlined below: (a) Lack of Difficulty Stratification, they collect diverse instruction categories but neglect the stratification of difficulty levels across these categories, which leads to overlap, bias, and low interpretability. (b) Lack of Fine-Grained Metrics, they conflate the model's ability to ``solve tasks" and ``follow constraints" into a single metric, which fails to accurately reflect its instruction-following capability. (c) Lack of Multi-Task Instructions, they overlook the fact that real-world user instructions often consist of multiple combined tasks. This paper proposes MMIFEvol, a framework for multimodal instruction evolving and benchmarking. First, we define the essential components of a carefully curated multimodal instruction set and establish corresponding difficulty levels, based on which we synthesize diverse instruction data. Next, we decouple the evaluation criteria for the instruction following into three different metrics to construct a high-quality benchmark and assess existing models. Experimental results demonstrate that current models still struggle with following complex instructions, while fine-tuning using MMIFEvol data effectively improves models' responsiveness to multimodal instructions.
Sihang Jiang 0001, Xiangru Zhu, Yuyan Chen, Xiaojun Meng, Jiansheng Wei, Yanghua Xiao
AAAI6
2026 Metaphor Reasoning is Meta-reasoning
abstract
Qianyu He, Junting Lu, Yikai Zhang, Siyu Yuan, Xiaojun Meng, Jiansheng Wei, Jiaqing Liang, Yanghua Xiao. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Qianyu He, Junting Lu, Yikai Zhang 0004, Xiaojun Meng, Jiansheng Wei, Jiaqing Liang, Yanghua Xiao
ACL (1)6
2026 Immediate Inference: The Missing Foundation in Large Language Model Logical Reasoning
abstract
Sihang Jiang, Zhiyu Lu, Keyi Wang, Jiaqing Liang, Yanghua Xiao, Xiaojun Meng, Jiansheng Wei. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Sihang Jiang 0001, Jiaqing Liang, Yanghua Xiao, Xiaojun Meng, Jiansheng Wei
ACL (1)7
2026 The GaoYao Benchmark: A Comprehensive Framework for Evaluating Multilingual and Multicultural Abilities of Large Language Models
abstract
Yilun Liu, Chunguang Zhao, Mengyao Piao, Lingqi Miao, Shimin Tao, Minggui HE, Chenxin Liu, Zhang Li, Mahongxia, Jiaxin Guo, Chen Liu, Liqun Deng, Jiansheng Wei, Xiaojun Meng, Fanyi Du, Daimeng Wei, Yanghua Xiao. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Yilun Liu 0001, Chunguang Zhao, Mengyao Piao, Lingqi Miao, Shimin Tao, Minggui He, Chenxin Liu, Hongxia Ma, Liqun Deng, Jiansheng Wei, Xiaojun Meng, Fanyi Du, Daimeng Wei, Yanghua Xiao
ACL (1)13
2026 Code LLMs Still Fall Short of Top Programmers: Evaluating Algorithmic Code Generation Through Computational Thinking
abstract
Evaluating the coding capabilities of models through algorithmic code generation is challenging, as it requires deep problem understanding and complex algorithm design. Current benchmarks suffer from a narrow focus on final execution results (such as pass@k), neglecting the crucial reasoning and problem-solving processes inherent in code generation. To address this limitation, we introduce a multi-phase algorithmic code generation benchmark, MUPA, structured around human computational thinking. MUPA dissects the evaluation into four distinct phases: example understanding, algorithm selection, solution description, and code generation. This framework facilitates a comprehensive assessment by providing insights into the model's intermediate problem-solving steps, rather than just the final code. We manually curated 197 high-quality competitive programming problems from Codeforces. Utilizing an LLM-as-a-judge paradigm with specialized prompts, our rigorous evaluation of several existing code generation LLMs reveals significant across-the-board challenges. Notably, we establish a positive correlation, indicating that proficiency in an earlier phase directly impacts performance in subsequent phases, underscoring the interdependency of these algorithmic skills. The benchmark is publicly available at https://github.com/cheniison/MUPA.
Shisong Chen, Ziyu Zhou 0019, Zhixu Li, Yanghua Xiao, Xin Lin 0001, Xiaojun Meng, Jiansheng Wei, Kuien Liu
WSDM9
2025 Sparsing Law: Towards Large Language Models with Greater Activation Sparsity
abstract
Activation sparsity denotes the existence of substantial weakly-contributed neurons within feed-forward networks of large language models (LLMs), providing wide potential benefits such as computation acceleration. However, existing works lack thorough quantitative studies on this useful property, in terms of both its measurement and influential factors. In this paper, we address three underexplored research questions: (1) How can activation sparsity be measured more accurately? (2) How is activation sparsity affected by the model architecture and training process? (3) How can we build a more sparsely activated and efficient LLM? Specifically, we develop a generalizable and performance-friendly metric, named CETT-PPL-1%, to measure activation sparsity. Based on CETT-PPL-1%, we quantitatively study the influence of various factors and observe several important phenomena, such as the convergent power-law relationship between sparsity and training data amount, the higher competence of ReLU activation than mainstream SiLU activation, the potential sparsity merit of a small width-depth ratio, and the scale insensitivity of activation sparsity. Finally, we provide implications for building sparse and effective LLMs, and demonstrate the reliability of our findings by training a 2.4B model with a sparsity ratio of 93.52%, showing 4.1$\times$ speedup compared with its dense version. The codes and checkpoints are available at https://github.com/thunlp/SparsingLaw/.
Yuqi Luo, Xu Han 0007, Yingfa Chen, Chaojun Xiao, Xiaojun Meng, Liqun Deng, Jiansheng Wei, Zhiyuan Liu 0001, Maosong Sun 0001
ICML8
2025 ReasVQA: Advancing VideoQA with Imperfect Reasoning Process
abstract
Jianxin Liang, Xiaojun Meng, Huishuai Zhang, Yueqian Wang, Jiansheng Wei, Dongyan Zhao. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Jianxin Liang, Xiaojun Meng, Huishuai Zhang, Yueqian Wang, Jiansheng Wei, Dongyan Zhao 0001
NAACL (Long Papers)5
2024 LAM-Depth: Laplace-Attention Module-Based Self-Supervised Monocular Depth Estimation
abstract
Depth estimation is extremely important in the world of driverless vehicles, which can provide key distance information for 3D scene perception and local path planning. Using convolutional neural network (CNN) to completely recover depth information from monocular images becomes a hot research trend. Supervised monocular depth estimation needs large numbers of per-pixel ground-truth depth data collected from LiDAR to train the model. This leads to high resource consumption and poor generalization ability. For the above reasons, the self-supervised learning method seems to be a promising alternative to monocular depth estimation. However, the up-sampling operations in existing encoder-decoder based architecture may lose some critical image information, which leads to boundary blurring and depth artifact in depth maps. In this paper, the Laplace-Attention module based self-supervised monocular depth estimation network (LAM-Depth) is designed to resolve this problem. Specifically, multi-scale Laplacian features are introduced into the corresponding streams in the decoder to fuse the low-level and skip-connection features. The concatenated features are then re-calibrated with a channel-wise attention unit for emphasizing the Laplacian features. Based on the above operations, the image information is preserved to the greatest extent in feature processing. The experiment results show that the proposed model LAM-Depth achieves a high ranking among the existing unsupervised methods and outperforms several supervised models trained with LiDAR data. Furthermore, we conduct experiments in real scenes to evaluate the generalization ability of LAM-Depth and obtain high-quality depth maps.
Jiansheng Wei, Shuguo Pan
IEEE Trans. Intell. Transp. Syst.1
2023 Wukong-Reader: Multi-modal Pre-training for Fine-grained Visual Document Understanding
abstract
Haoli Bai, Zhiguang Liu, Xiaojun Meng, Li Wentao, Shuang Liu, Yifeng Luo, Nian Xie, Rongfu Zheng, Liangwei Wang, Lu Hou, Jiansheng Wei, Xin Jiang, Qun Liu. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Haoli Bai, Xiaojun Meng, Yifeng Luo, Nian Xie, Rongfu Zheng, Liangwei Wang 0004, Lu Hou 0002, Jiansheng Wei, Xin Jiang 0002, Qun Liu 0001
ACL (1)11
2022 Towards Accurate Network Quantization with Equivalent Smooth Regularizer
Kirill Solodskikh, Vladimir Chikin, Ruslan Aydarkhanov, Dehua Song, Irina Zhelavskaya, Jiansheng Wei
ECCV (11)6
2022 Diffusion-Based Voice Conversion with Fast Maximum Likelihood Sampling Scheme
Vadim Popov, Ivan Vovk, Vladimir Gogoryan, Tasnima Sadekova, Mikhail Sergeevich Kudinov, Jiansheng Wei
ICLR6
2022 A Unified System for Voice Cloning and Voice Conversion through Diffusion Probabilistic Modeling
Tasnima Sadekova, Vladimir Gogoryan, Ivan Vovk, Vadim Popov, Mikhail A. Kudinov, Jiansheng Wei
INTERSPEECH6
2022 Fast Grad-TTS: Towards Efficient Diffusion-Based Speech Generation on CPU
Ivan Vovk, Tasnima Sadekova, Vladimir Gogoryan, Vadim Popov, Mikhail A. Kudinov, Jiansheng Wei
INTERSPEECH6
2022 A dynamic object filtering approach based on object detection and geometric constraint between frames
abstract
Abstract In order to eliminate the influence of moving targets in visual positioning, a dynamic object filtering approach based on object detection and inter‐frame geometric constraints is proposed to filter the dynamic objects in monocular images. The object detection algorithm is firstly used to identify and locate objects in the single‐frame image, and the object matching between frames is performed. Then, the depth estimation network and pose recovery network are trained jointly to output estimated depth along with transformation matrix of object centroids between frames respectively. Finally, the mapping centroid is obtained from the estimated depth and inter‐frame transformation matrix. The joint constraint function is performed to complete the detection and filtering of dynamic objects. The experimental results show that the proposed dynamic object filtering approach not only can filter moving objects in sequence images accurately but also allows to reserve the dynamic objects that are temporarily in the stop state. The generalization ability of this approach is also verified in a real urban road scene and meets the requirements of follow‐up research in visual positioning.
Jiansheng Wei, Shuguo Pan, Tao Zhao 0005
IET Image Process.1
2022 Triaxial Squeeze Attention Module and Mutual-Exclusion Loss Based Unsupervised Monocular Depth Estimation
Jiansheng Wei, Shuguo Pan, Tao Zhao 0005
Neural Process. Lett.1
2022 Attention Unet++ for lightweight depth estimation from sparse depth samples and a single RGB image
Tao Zhao 0005, Shuguo Pan, Chao Sheng, Yingchun Sun, Jiansheng Wei
Vis. Comput.6
2016 HiGene: A high-performance platform for genomic data analysis
abstract
Post-sequencing genomic data analysis becomes a major challenge while next-generation sequencing technologies evolve by leaps and bounds. The data-intensive and compute-intensive nature of genome analysis makes cluster computing an attractive choice for building efficient solutions. This paper presents HiGene, a high-performance genome analysis platform that exploits big data technology to revolutionize genomics data crunching power. HiGene reconstructs the genome analysis pipeline by exploiting both multi-core and multi-node parallelization using Apache Spark, and employs two key techniques to further boost the performance. First, a dynamic computing resource re-allocator is implemented, which allows flexible on-demand resource allocation for operations inside tasks. Second, an efficient skew mitigation approach is proposed, which automatically identifies and resolves data skew and computation skew through task repartitioning and resource reallocating respectively. HiGene has been evaluated with a whole human genome dataset on a 10-node Huawei 5885 cluster. Experimental results show that HiGene achieves remarkable high performance that reduces the total running time on a whole genome sequence dataset from days to nearly one hour. Furthermore, it is two times faster than state-of-the-art cluster based approaches.
Liqun Deng, Guowei Huang 0002, Yuzheng Zhuang, Jiansheng Wei, Youliang Yan
BIBM4
2014 Efficiently Representing Membershipfor Variable Large Data Sets
abstract
Cloud computing has raised new challenges for the membership representation scheme of storage systems that manage very large data sets. This paper proposes DBA, a dynamic Bloom filter array aimed at representing membership for variable large data sets in storage systems in a scalable way. DBA consists of dynamically created groups of space-efficient Bloom filters (BFs) to accommodate changes in set sizes. Within a group, BFs are homogeneous and the data layout is optimized at the bit level to enable parallel access and thus achieve high query performance. DBA can effectively control its query accuracy by partially adjusting the error rate of the constructing BFs, where each BF only represents an independent subset to help locate elements and confirm membership. Further, DBA supports element deletion by introducing a lazy update policy. We prototype and evaluate our DBA scheme as a scalable fast index in the MAD2 deduplication storage system. Experimental results reveal that DBA (with 64 BFs per group) shows significantly higher query performance than the state-of-the-art approach while scaling up to 160 BFs. DBA is also shown to excel in scalability, query accuracy, and space efficiency by theoretical analysis and experimental evaluation.
Jiansheng Wei, Hong Jiang 0001, Ke Zhou 0001, Dan Feng 0001
IEEE Trans. Parallel Distributed Syst.1
2011 DBA: A Dynamic Bloom Filter Array for Scalable Membership Representation of Variable Large Data Sets
abstract
This paper proposes a Dynamic Bloom filter Array (DBA) to represent membership for variable large data sets in storage systems in a scalable way. DBA consists of dynamically created groups of space-efficient Bloom Filters (BFs) to accommodate changes in set sizes. In each group, BFs are homogeneous and the data layout is optimized at the bit level, so that they can be accessed in parallel to achieve high query performance. DBA can effectively control its query accuracy by partially adjusting the error rate of constructing BFs, where each BF corresponds to an independent subset of the data set to facilitate element location and membership confirmation. Further, DBA supports element deletion by introducing a lazy update policy. We prototype and evaluate our DBA scheme as a scalable fast index in the MAD2 deduplication storage system. Experimental results show that DBA (with 64 BFs per group) is capable of maintaining 90% of the peek query performance while scaling up to 160 BFs. DBA is also shown to excel in performance and space efficiency by theoretical analysis and other experiments based on real-world data sets.
Jiansheng Wei, Hong Jiang 0001, Ke Zhou 0001, Dan Feng 0001
MASCOTS1
2011 Detecting Duplicates over Sliding Windows with RAM-Efficient Detached Counting Bloom Filter Arrays
abstract
Detecting duplicates over sliding windows is an important technique for monitoring and analysing data streams. Since recording the exact information of elements in a sliding window can be RAM-resource-intensive and introduce an unacceptable search complexity, several approximate membership representation schemes have been proposed to build in-memory fast indices. However, various challenges facing RAM utilization and scalability remain. This paper proposes a Detached Counting Bloom filter Array (DCBA) to flexibly and efficiently detect duplicates over sliding windows. A DCBA consists of an array of detached counting Bloom filters (DCBFs), where each DCBF is essentially a Bloom filter that is associated with a detached timer (counter) array. The DCBA scheme functions as a circular FIFO queue and keeps a filling DCBF for accommodating fresh elements and a decaying DCBF for evicting stale elements. DCBA allows the timer arrays belonging to fully filled DCBFs to be offloaded to disks to greatly improve the memory space efficiency. The fully filled DCBFs will remain stable until their elements become stale, which allows a DCBA to be efficiently replicated for the purpose of data reliability or information sharing. Further, DCBA can be cooperatively maintained by clustered nodes, which provides scalable solution for mining massive data streams. Mathematical analysis and experimental results show that a DCBA (containing 64 DCBFs) requires less than 10% of its components to be kept in RAM while maintaining more than 95% of its query performance, which significantly outperforms existing schemes in memory efficiency and scalability.
Jiansheng Wei, Hong Jiang 0001, Ke Zhou 0001, Dan Feng 0001, Hua Wang 0008
NAS1
2010 MAD2: A scalable high-throughput exact deduplication approach for network backup services
abstract
Deduplication has been widely used in disk-based secondary storage systems to improve space efficiency. However, there are two challenges facing scalable high-throughput deduplication storage. The first is the duplicate-lookup disk bottleneck due to the large size of data index that usually exceeds the available RAM space, which limits the deduplication throughput. The second is the storage node island effect resulting from duplicate data among multiple storage nodes that are difficult to eliminate. Existing approaches fail to completely eliminate the duplicates while simultaneously addressing the challenges. This paper proposes MAD2, a scalable high-throughput exact deduplication approach for network backup services. MAD2 eliminates duplicate data both at the file level and at the chunk level by employing four techniques to accelerate the deduplication process and evenly distribute data. First, MAD2 organizes fingerprints into a Hash Bucket Matrix (HBM), whose rows can be used to preserve the data locality in backups. Second, MAD2 uses Bloom Filter Array (BFA) as a quick index to quickly identify non-duplicate incoming data objects or indicate where to find a possible duplicate. Third, Dual Cache is integrated in MAD2 to effectively capture and exploit data locality. Finally, MAD2 employs a DHT-based Load-Balance technique to evenly distribute data objects among multiple storage nodes in their backup sequences to further enhance performance with a well-balanced load. We evaluate our MAD2 approach on the backend storage of B-Cloud, a research-oriented distributed system that provides network backup services. Experimental results show that MAD2 significantly outperforms the state-of-the-art approximate deduplication approaches in terms of deduplication efficiency, supporting a deduplication throughput of at least 100MB/s for each storage component.
Jiansheng Wei, Hong Jiang 0001, Ke Zhou 0001, Dan Feng 0001
MSST1