Wenbo Su

dblp:147/1636 · DBLP profile ↗
← Back
30ranked-venue papers
3as first author
28since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 23 · 1 first-author · 23 since 2021Databases, data management, data science and information retrieval · 6 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Computer networks · 2 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Think-J: Learning to Think for Generative LLM-as-a-Judge
abstract
LLM-as-a-Judge refers to the automatic modeling of preferences for responses generated by Large Language Models (LLMs), which is of significant importance for both LLM evaluation and reward modeling. Although generative LLMs have made substantial progress in various tasks, their performance as LLM-Judge still falls short of expectations. In this work, we propose Think-J, which improves generative LLM-as-a-Judge by learning how to think. We first utilized a small amount of curated data to develop the model with initial judgment thinking capabilities. Subsequently, we optimize the judgment thinking traces based on reinforcement learning (RL). We propose two methods for judgment thinking optimization, based on offline and online RL, respectively. The offline method requires training a critic model to construct positive and negative examples for learning. The online method defines rule-based reward as feedback for optimization. Experimental results showed that our approach can significantly enhance the evaluation capability of generative LLM-Judge, surpassing both generative and classifier-based LLM-Judge without requiring extra human annotations.
Hui Huang 0021, Yancheng He, Hongli Zhou 0001, Weixun Wang, Wenbo Su
AAAI8
2026 Read As Human: Compressing Context via Parallelizable Close Reading and Skimming
abstract
Jiwei Tang, Shilei Liu, Zhicheng Zhang, Qingsong Lv, Runsong Zhao, Tingwei Lu, Langming Liu, Haibin Chen, Yujin Yuan, Hai-Tao Zheng, Wenbo Su, Bo Zheng. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Jiwei Tang, Shilei Liu, Zhicheng Zhang 0008, Qingsong Lv, Runsong Zhao, Tingwei Lu, Langming Liu, Yujin Yuan, Wenbo Su
ACL (1)11
2026 SELECting over Tokens: Curating Pre-training Data at Scale via Token Classification
abstract
Xin Tong, Weidong Zhang, Jiaang Li, Haibin Chen, Shilei Liu, Langming Liu, Kangtao Lv, Yujin Yuan, Wenbo Su, Bo Zheng. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Jiaang Li 0004, Shilei Liu, Langming Liu, Kangtao Lv, Yujin Yuan, Wenbo Su, Bo Zheng 0007
ACL (1)9
2026 ShopSimulator: Evaluating and Exploring RL-Driven LLM Agent for Shopping Assistants
abstract
Pei Wang, Yanan Wu, Xiaoshuai Song, Weixun Wang, Gengru Chen, Zhongwen Li, Kezhong Yan, Qi Liu, Ken Deng, Shuaibing Zhao, Shaopan Xiong, Xuepeng Liu, Xuefeng Chen, Wanxi Deng, Wenbo Su, Bo Zheng. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Xiaoshuai Song, Weixun Wang, Gengru Chen, Kezhong Yan, Ken Deng, Shuaibing Zhao, Shaopan Xiong, Xuepeng Liu, Wanxi Deng, Wenbo Su
ACL (1)15
2026 CoMeT: Collaborative Memory Transformer for Efficient Long Context Modeling
abstract
Runsong Zhao, Shilei Liu, Jiwei Tang, Langming Liu, Haibin Chen, Weidong Zhang, Yujin Yuan, Tong Xiao, JingBo Zhu, Wenbo Su, Bo Zheng. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Runsong Zhao, Shilei Liu, Jiwei Tang, Langming Liu, Yujin Yuan, Tong Xiao 0001, Wenbo Su, Bo Zheng 0007
ACL (1)10
2026 USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models
abstract
Baolin Zheng, Guanlin Chen, Qingyang Teng, Hongqiong Zhong, Yingshui Tan, Zhendong Liu, Weixun Wang, Jiaheng Liu, Jian Yang, Huiyun Jing, Jincheng Wei, Wenbo Su, Xiaoyong Zhu, Bo Zheng, Kaifu Zhang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Baolin Zheng, Qingyang Teng, Hongqiong Zhong, Yingshui Tan, Weixun Wang, Jian Yang 0037, Huiyun Jing, Jincheng Wei, Wenbo Su, Xiaoyong Zhu, Bo Zheng 0007, Kaifu Zhang
ACL (1)12
2026 RollPacker: Taming Long-Tail Rollouts for RL Post-Training with Tail Batching
Yuheng Zhao, Dakai An, Tianyuan Wu, Lunxi Cao, Shaopan Xiong, Ju Huang, Weixun Wang, Siran Yang, Wenbo Su, Jiamang Wang, Lin Qu, Bo Zheng 0007, Wei Wang 0030
NSDI10
2026 Unlocking Scaling Law in Industrial Recommendation Systems with a Three-step Paradigm based Large User Model
abstract
Recent advancements in autoregressive Large Language Models (LLMs) have achieved remarkable progress, largely driven by their scalability—commonly formalized as the scaling law. Inspired by these successes, there has been growing interest in adapting LLMs to recommendation systems (RecSys) by reformulating recommendation tasks as generative sequence modeling problems. However, existing End-to-End Generative Recommendation (E2E-GR) methods often sacrifice the practical advantages of traditional Deep Learning-based Recommendation Models (DLRMs)—including mature feature engineering, modular architectures, and production-grade optimization practices. This trade-off introduces critical challenges that hinder the effective application of scaling laws in industrial RecSys. In this paper, we present Large User Model (LUM), a scalable and production-aware framework that bridges the gap between generative modeling and industrial recommendation requirements. LUM addresses these limitations through a principled three-step paradigm, designed to preserve the flexibility of autoregressive generation while maintaining compatibility with real-world deployment constraints. Extensive experiments show that LUM outperforms state-of-the-art DLRMs and E2E-GR approaches across multiple benchmarks. Notably, LUM exhibits strong scalability: performance improves consistently as the model scales up to 7 billion parameters. Furthermore, LUM has been successfully deployed in a large-scale industrial application, where it delivered statistically significant gains in a live A/B test, demonstrating both its effectiveness and practical viability.
Bencheng Yan, Shilei Liu, Yizhen Zhang 0005, Yujin Yuan, Langming Liu, Wenbo Su, Pengjie Wang 0002, Jian Xu 0015, Bo Zheng 0007
WSDM10
2026 NEZHA: A Zero-sacrifice and Hyperspeed Decoding Architecture for Generative Recommendations
abstract
Generative Recommendation (GR), powered by Large Language Models (LLMs), represents a promising new paradigm for industrial recommender systems. However, their practical application is severely hindered by high inference latency, making them infeasible for high-throughput, real-time services and limiting their overall business impact. While Speculative Decoding (SD) has been proposed to accelerate the autoregressive generation process, existing implementations introduce new bottlenecks: they typically require separate draft models and model-based verifiers, which require additional training and increase latency overhead. In this paper, we address these challenges with NEZHA, a novel architecture that achieves hyperspeed decoding for GR systems without sacrificing recommendation quality. Specifically, NEZHA integrates a nimble autoregressive draft head directly into the primary model, enabling efficient self-drafting. This design, combined with a specialized input prompt structure, preserves the integrity of sequence-to-sequence generation. Furthermore, to tackle the critical problem of hallucination—a major source of performance degradation—we introduce an efficient, model-free verifier based on a hash set. We demonstrate the effectiveness of NEZHA through extensive experiments on public datasets and have successfully deployed the system on Taobao since October 2025, achieving 1.2% business improvement, translating to billion-level advertising revenue and serving hundreds of millions of daily active users. The code is available at https://github.com/Applied-Machine-Learning- Lab/WWW2026_NEZHA.
Yejing Wang, Shengyu Zhou, Jinyu Lu, Ziwei Liu 0010, Langming Liu, Maolin Wang 0001, Wenlin Zhang 0001, Feng Li 0067, Wenbo Su, Pengjie Wang 0002, Jian Xu 0015, Xiangyu Zhao 0001
WWW9
2025 Chinese SimpleQA: A Chinese Factuality Evaluation for Large Language Models
abstract
Yancheng He, Shilong Li, Jiaheng Liu, Yingshui Tan, Weixun Wang, Hui Huang, Xingyuan Bu, Hangyu Guo, Chengwei Hu, Boren Zheng, Zhuoran Lin, Dekai Sun, Zhicheng Zheng, Wenbo Su, Bo Zheng. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Yancheng He, Yingshui Tan, Weixun Wang, Hui Huang 0021, Xingyuan Bu, Hangyu Guo, Chengwei Hu, Boren Zheng, Zhuoran Lin, Dekai Sun, Zhicheng Zheng, Wenbo Su, Bo Zheng 0007
ACL (1)14
2025 Can Large Language Models Detect Errors in Long Chain-of-Thought Reasoning?
abstract
Yancheng He, Shilong Li, Jiaheng Liu, Weixun Wang, Xingyuan Bu, Ge Zhang, Z.y. Peng, Zhaoxiang Zhang, Zhicheng Zheng, Wenbo Su, Bo Zheng. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Yancheng He, Weixun Wang, Xingyuan Bu, Ge Zhang 0009, Z. Y. Peng, Zhaoxiang Zhang 0001, Zhicheng Zheng, Wenbo Su, Bo Zheng 0007
ACL (1)10
2025 M2RC-EVAL: Massively Multilingual Repository-level Code Completion Evaluation
abstract
Jiaheng Liu, Ken Deng, Congnan Liu, Jian Yang, Shukai Liu, He Zhu, Peng Zhao, Linzheng Chai, Yanan Wu, JinKe JinKe, Ge Zhang, Zekun Moore Wang, Guoan Zhang, Yingshui Tan, Bangyu Xiang, Zhaoxiang Zhang, Wenbo Su, Bo Zheng. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Ken Deng, Congnan Liu, Jian Yang 0030, Linzheng Chai, Ge Zhang 0009, Zekun Moore Wang, Guoan Zhang, Yingshui Tan, Bangyu Xiang, Zhaoxiang Zhang 0001, Wenbo Su, Bo Zheng 0007
ACL (1)17
2025 Chinese SafetyQA: A Safety Short-form Factuality Benchmark for Large Language Models
abstract
Yingshui Tan, Boren Zheng, Baihui Zheng, Kerui Cao, Huiyun Jing, Jincheng Wei, Jiaheng Liu, Yancheng He, Wenbo Su, Xiaoyong Zhu, Bo Zheng, Kaifu Zhang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Yingshui Tan, Boren Zheng, Baihui Zheng, Kerui Cao, Huiyun Jing, Jincheng Wei, Yancheng He, Wenbo Su, Xiaoyong Zhu, Bo Zheng 0007, Kaifu Zhang
ACL (1)9
2025 ECKGBench: Benchmarking Large Language Models in E-commerce Leveraging Knowledge Graph
abstract
Large language models (LLMs) have demonstrated their capabilities across various natural language processing (NLP) tasks. Their potential in e-commerce is also substantial, evidenced by existing implementations in scenarios such as platform search and recommender systems. One obstinate concern associated with LLMs is the factuality issue (e.g., hallucination), which is urgent in e-commerce due to its significant impact on user experience and revenue. While some methods aim to evaluate the factuality of LLMs, issues such as lack of objectivity, high consumption, and lack of domain expertise arise. To this end, leveraging a collected knowledge graph (KG) as a reliable source, we propose ECKGBench, a question-answering dataset to assess LLMs' capacity in e-commerce. Specifically, each question is automatically generated based on one KG triple through a standardized pipeline, guaranteeing evaluation quality and reliability. We evaluate advanced LLMs using ECKGBench and provide insights into experimental results. The dataset is available online at~ https://github.com/OpenStellarTeam/ECKGBench.
Langming Liu, Yuhao Wang 0006, Yujin Yuan, Shilei Liu, Wenbo Su, Xiangyu Zhao 0001, Bo Zheng 0007
CIKM6
2025 AIR: Complex Instruction Generation via Automatic Iterative Refinement
abstract
With the development of large language models, their ability to follow simple instructions has significantly improved.However, adhering to complex instructions remains a major challenge.Current approaches to generating complex instructions are often irrelevant to the current instruction requirements or suffer from limited scalability and diversity.Moreover, methods such as back-translation, while effective for simple instruction generation, fail to leverage the rich knowledge and formatting in human written documents.In this paper, we propose a novel Automatic Iterative Refinement (AIR) framework to generate complex instructions with constraints, which not only better reflects the requirements of real scenarios but also significantly enhances LLMs' ability to follow complex instructions.The AIR framework consists of two stages: 1) Generate an initial instruction from a document; 2) Iteratively refine instructions with LLM-as-judge guidance by comparing the model's output with the document to incorporate valuable constraints.Finally, we construct the AIR-10K dataset with 10K complex instructions and demonstrate that instructions generated with our approach significantly improve the model's ability to follow complex instructions, outperforming existing methods for instruction generation 1 . Model Answer Refined D Model Answer Check Constraints Identify ConstraintsI: Write a casual review of a water-proof camera.
Yancheng He, Yu Li 0007, Hui Huang 0021, Chengwei Hu, Wenbo Su, Bo Zheng 0007
EMNLP8
2025 How to inject knowledge efficiently? Knowledge Infusion Scaling Law for Pre-training Large Language Models
abstract
Large language models (LLMs) have attracted significant attention due to their impressive general capabilities across diverse downstream tasks.However, without domain-specific optimization, they often underperform on specialized knowledge benchmarks and even produce hallucination.Recent studies show that strategically infusing domain knowledge during pretraining can substantially improve downstream performance.A critical challenge lies in balancing this infusion trade-off: injecting too little domain-specific data yields insufficient specialization, whereas excessive infusion triggers catastrophic forgetting of previously acquired knowledge.In this work, we focus on the phenomenon of memory collapse induced by overinfusion.Through systematic experiments, we make two key observations, i.e. 1) Critical collapse point: each model exhibits a threshold beyond which its knowledge retention capabilities sharply degrade.2) Scale correlation: these collapse points scale consistently with the model's size.Building on these insights, we propose a knowledge infusion scaling law that predicts the optimal amount of domain knowledge to inject into large LLMs by analyzing their smaller counterparts.Extensive experiments across different model sizes and pertaining token budgets validate both the effectiveness and generalizability of our scaling law.
Kangtao Lv, Yujin Yuan, Langming Liu, Shilei Liu, Wenbo Su, Bo Zheng 0007
EMNLP7
2025 MTU-Bench: A Multi-granularity Tool-Use Benchmark for Large Language Models
abstract
Large Language Models (LLMs) have displayed massive improvements in reason- ing and decision-making skills and can hold natural conversations with users. Recently, many tool-use benchmark datasets have been proposed. However, existing datasets have the following limitations: (1). Insufficient evaluation scenarios (e.g., only cover limited tool-use scenes). (2). Extensive evaluation costs (e.g., GPT API costs). To address these limitations, in this work, we propose a multi-granularity tool-use benchmark for large language models called MTU-Bench. For the "multi-granularity" property, our MTU-Bench covers five tool usage scenes (i.e., single-turn and single-tool, single-turn and multiple-tool, multiple-turn and single-tool, multiple-turn and multiple-tool, and out-of-distribution tasks). Besides, all evaluation metrics of our MTU-Bench are based on the prediction results and the ground truth without using any GPT or human evaluation metrics. Moreover, our MTU-Bench is collected by transforming existing high-quality datasets to simulate real-world tool usage scenarios, and we also propose an instruction dataset called MTU-Instruct data to enhance the tool-use abilities of existing LLMs. Comprehensive experimental results demonstrate the effectiveness of our MTU-Bench.
Noah Wang, Xiaoshuai Song, Z. Y. Peng, Ken Deng, Jiakai Wang, Junran Peng, Ge Zhang 0009, Hangyu Guo, Zhaoxiang Zhang 0001, Wenbo Su, Bo Zheng 0007
ICLR14
2025 ChineseEcomQA: A Scalable E-commerce Concept Evaluation Benchmark for Large Language Models
abstract
With the increasing use of Large Language Models (LLMs) in fields such as e-commerce, domain-specific concept evaluation benchmarks are crucial for assessing their domain capabilities. Existing LLMs may generate factually incorrect information within the complex e-commerce applications. Therefore, it is necessary to build an e-commerce concept benchmark. Existing benchmarks encounter two primary challenges: (1) handle the heterogeneous and diverse nature of tasks(2) distinguish between generality and specificity within the e-commerce field. To address these problems, we propose ChineseEcomQA, a scalable question-answering benchmark focused on fundamental e-commerce concepts. ChineseEcomQA is built on three core characteristics: Focus on Fundamental Concept, E-commerce Generality and E-commerce Expertise. Fundamental concepts are designed to be applicable across a diverse array of e-commerce tasks, thus addressing the challenge of heterogeneity and diversity. Additionally, by carefully balancing generality and specificity, ChineseEcomQA effectively differentiates between broad e-commerce concepts, allowing for precise validation of domain capabilities. We achieve this through a scalable benchmark construction process that combines LLM validation, Retrieval-Augmented Generation (RAG) validation, and rigorous manual annotation. Based on ChineseEcomQA, we conduct extensive evaluations on mainstream LLMs and provide some valuable insights. We hope that ChineseEcomQA could guide future domain-specific evaluations, and facilitate broader LLM adoption in e-commerce applications.
Kangtao Lv, Chengwei Hu, Yanshi Li, Yujin Yuan, Yancheng He, Xingyao Zhang 0003, Langming Liu, Shilei Liu, Wenbo Su, Bo Zheng 0007
KDD (2)10
2025 UQABench: Evaluating User Embedding for Prompting LLMs in Personalized Question Answering
abstract
Large language models (LLMs) achieve remarkable success in natural language processing (NLP). In practical scenarios like recommendations, as users increasingly seek personalized experiences, it becomes crucial to incorporate user interaction history into the context of LLMs to enhance personalization. However, from a practical utility perspective, user interactions' extensive length and noise present challenges when used directly as text prompts. A promising solution is to compress and distill interactions into compact embeddings, serving as soft prompts to assist LLMs in generating personalized responses. Although this approach brings efficiency, a critical concern emerges: Can user embeddings adequately capture valuable information and prompt LLMs? To address this concern, we propose UQABench, a benchmark designed to evaluate the effectiveness of user embeddings in prompting LLMs for personalization. We establish a fair and standardized evaluation process, encompassing pre-training, fine-tuning, and evaluation stages. To thoroughly evaluate user embeddings, we design three dimensions of tasks: sequence understanding, action prediction, and interest perception. These evaluation tasks cover the industry's demands in traditional recommendation tasks, such as improving prediction accuracy, and its aspirations for LLM-based methods, such as accurately understanding user interests and enhancing the user experience. We conduct extensive experiments on various state-of-the-art methods for modeling user embeddings. Additionally, we reveal the scaling laws of leveraging user embeddings to prompt LLMs. The benchmark is available online at https://github.com/OpenStellarTeam/UQABench.
Langming Liu, Shilei Liu, Yujin Yuan, Yizhen Zhang 0005, Bencheng Yan, Wenbo Su, Pengjie Wang 0002, Jian Xu 0015, Bo Zheng 0007
KDD (2)10
2025 Multi-task Offline Reinforcement Learning for Online Advertising in Recommender Systems
abstract
Online advertising in recommendation platforms has gained significant attention, with a predominant focus on channel recommendation and budget allocation strategies. However, current offline reinforcement learning (RL) methods face substantial challenges when applied to sparse advertising scenarios, primarily due to severe overestimation, distributional shifts, and overlooking budget constraints. To address these issues, we propose MTORL, a novel multi-task offline RL model that targets two key objectives. First, we establish a Markov Decision Process (MDP) framework specific to the nuances of advertising. Then, we develop a causal state encoder to capture dynamic user interests and temporal dependencies, facilitating offline RL through conditional sequence modeling. Causal attention mechanisms are introduced to enhance user sequence representations by identifying correlations among causal states. We employ multi-task learning to decode actions and rewards, simultaneously addressing channel recommendation and budget allocation. Notably, our framework includes an automated system for integrating these tasks into online advertising. Extensive experiments on offline and online environments demonstrate MTORL's superiority over state-of-the-art methods. The code is available online at https://github.com/Applied-Machine-Learning-Lab/MTORL.
Langming Liu, Chi Zhang 0060, Bo Li 0156, Hongzhi Yin, Xuetao Wei, Wenbo Su, Bo Zheng 0007, Xiangyu Zhao 0001
KDD (2)7
2025 Towards Visualization-of-Thought Jailbreak Attack against Large Visual Language Models
abstract
As Visual Language Models (VLMs) continue to evolve, they have demonstrated increasingly sophisticated logical reasoning capabilities and multimodal thought generation, opening doors to widespread applications. However, this advancement raises serious concerns about content security, particularly when these models process complex multimodal inputs requiring intricate reasoning. When faced with these safety challenges, the critical competition between logical reasoning and safety objectives of VLMs is often overlooked in previous works. In this paper, we introduce Visualization-of-Thought Attack (\textbf{VoTA}), a novel and automated attack framework that strategically constructs chains of images with risky visual thoughts to challenge victim models. Our attack provokes the inherent conflict between the model's logical processing and safety protocols, ultimately leading to the generation of unsafe content. Through comprehensive experiments, VoTA achieves remarkable effectiveness, improving the average attack success rate (ASR) by 26.71\% (from 63.70\% to 90.41\%) on 9 open-source and 6 commercial VLMs, compared to the state-of-the-art methods. These results expose a critical vulnerability: current VLMs struggle to maintain safety guarantees when processing insecure multimodal visualization-of-thought inputs, highlighting the urgency and necessity of enhancing safety alignment. Our code and dataset are available at https://github.com/Hongqiong12/VoTA. Content Warning: This paper contains harmful contents that may be offensive.
Hongqiong Zhong, Qingyang Teng, Baolin Zheng, Yingshui Tan, Wenbo Su, Xiaoyong Zhu, Bo Zheng 0007, Kaifu Zhang
NeurIPS8
2025 Comparative Analysis and Optimization of Magnetic Field Energy Harvesters Based on Split Three-Phase Power Line Joint Energy Harvesting
abstract
With the rapid development of smart grids, abundant online monitoring sensors are installed at critical nodes of medium and low voltage power lines. To ensure reliable operation of these monitoring sensors, the power source problem needs to be solved urgently. The magnetic field energy can be captured and used as a stable energy source for online monitoring sensors. However, the single-phase magnetic field energy harvester (MFEH) tends to achieve low energy at low currents, and the output power is significantly influenced by load fluctuations. Motivated by these challenges, a three-phase distributed magnetic field energy harvester (TDMFEH) with transforms is proposed in this article. The power multiplier changes of the TDMFEH compared with the single-phase MFEH are further analyzed. Moreover, this article proposes an output power boosting control method under load fluctuations with single-phase and three phase joint energy harvesting, to achieve higher output power with a wide range of load connected. The experimental results demonstrate that the TDMFEH can achieve a maximum power ratio of 2.13 compared to an equal-volume single-phase MFEH. Besides, the output power can be enhanced with the switching between single-phase MFEH and TDMFEH when the load fluctuates.
Chenjin Xu, Wei Wang 0148, Wenbo Su, Minqiang Hu
IEEE Trans. Ind. Informatics4
2024 MT-Bench-101: A Fine-Grained Benchmark for Evaluating Large Language Models in Multi-Turn Dialogues
abstract
Ge Bai, Jie Liu, Xingyuan Bu, Yancheng He, Jiaheng Liu, Zhanhui Zhou, Zhuoran Lin, Wenbo Su, Tiezheng Ge, Bo Zheng, Wanli Ouyang. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Ge Bai, Jie Liu 0047, Xingyuan Bu, Yancheng He, Zhanhui Zhou, Zhuoran Lin, Wenbo Su, Tiezheng Ge, Bo Zheng 0007, Wanli Ouyang
ACL (1)8
2024 DDK: Distilling Domain Knowledge for Efficient Large Language Models
abstract
Despite the advanced intelligence abilities of large language models (LLMs) in various applications, they still face significant computational and storage demands. Knowledge Distillation (KD) has emerged as an effective strategy to improve the performance of a smaller LLM (i.e., the student model) by transferring knowledge from a high-performing LLM (i.e., the teacher model). Prevailing techniques in LLM distillation typically use a black-box model API to generate high-quality pretrained and aligned datasets, or utilize white-box distillation by altering the loss function to better transfer knowledge from the teacher LLM. However, these methods ignore the knowledge differences between the student and teacher LLMs across domains. This results in excessive focus on domains with minimal performance gaps and insufficient attention to domains with large gaps, reducing overall performance. In this paper, we introduce a new LLM distillation framework called DDK, which dynamically adjusts the composition of the distillation dataset in a smooth manner according to the domain performance differences between the teacher and student models, making the distillation process more stable and effective. Extensive evaluations show that DDK significantly improves the performance of student models, outperforming both continuously pretrained baselines and existing knowledge distillation methods by a large margin.
Yuanxing Zhang, Haoran Que, Ken Deng, Zhiqi Bai, Jie Liu 0047, Ge Zhang 0009, Jiakai Wang, Congnan Liu, Jiamang Wang, Lin Qu, Wenbo Su, Bo Zheng 0007
NeurIPS15
2024 D-CPT Law: Domain-specific Continual Pre-Training Scaling Law for Large Language Models
abstract
Continual Pre-Training (CPT) on Large Language Models (LLMs) has been widely used to expand the model’s fundamental understanding of specific downstream domains (e.g., math and code). For the CPT on domain-specific LLMs, one important question is how to choose the optimal mixture ratio between the general-corpus (e.g., Dolma, Slim-pajama) and the downstream domain-corpus. Existing methods usually adopt laborious human efforts by grid-searching on a set of mixture ratios, which require high GPU training consumption costs. Besides, we cannot guarantee the selected ratio is optimal for the specific domain. To address the limitations of existing methods, inspired by the Scaling Law for performance prediction, we propose to investigate the Scaling Law of the Domain-specific Continual Pre-Training (D-CPT Law) to decide the optimal mixture ratio with acceptable training costs for LLMs of different sizes. Specifically, by fitting the D-CPT Law, we can easily predict the general and downstream performance of arbitrary mixture ratios, model sizes, and dataset sizes using small-scale training costs on limited experiments. Moreover, we also extend our standard D-CPT Law on cross-domain settings and propose the Cross-Domain D-CPT Law to predict the D-CPT law of target domains, where very small training costs (about 1\% of the normal training costs) are needed for the target domains. Comprehensive experimental results on six downstream domains demonstrate the effectiveness and generalizability of our proposed D-CPT Law and Cross-Domain D-CPT Law.
Haoran Que, Ge Zhang 0009, Xingwei Qu, Yinghao Ma, Feiyu Duan, Zhiqi Bai, Jiakai Wang, Yuanxing Zhang, Xu Tan 0003, Jie Fu 0001, Jiamang Wang, Lin Qu, Wenbo Su, Bo Zheng 0007
NeurIPS15
2023 Sensitivity Analysis of Free-Standing Columnar Magnetic Field Energy Harvester for Powering Wireless Monitoring Sensors
abstract
Magnetic field energy harvesting is the effective method to power the monitoring sensors by capturing the magnetic field around the current-carrying structures. However, conventional toroidal magnetic field energy harvester (MFEH) face challenges in meeting the requirements of high anti-saturation characteristics and convenient installation. These limitations hinder the advancement of magnetic field energy scavenging. This paper introduces the integration of the columnar MFEH and the low-power current sensor to enable current status monitoring of the power equipment. The sensitivity of free-standing columnar magnetic field energy harvester is analyzed using the method of magnetic flux block superposition at first. Accordingly, comprehensive studies are conducted on various parameters such as the number of turns, the winding configurations, magnetic core dimension and the through-hole radius of the magnetic core. Further, the operational performance of the energy harvester is investigated in scenarios involving offset and rotation. The experiment evaluations of the proposed design are presented by powering the wireless current sensor periodically. The experimental results demonstrate that the periodic wireless chip current sensor can be driven under a 300 A power line current with intervals of 29 seconds. As a result, the proposed solution proves to be highly effective in scavenging magnetic field energy around power lines and can be employed to power wireless monitoring sensors.
Chenjin Xu, Wei Wang 0148, Wenbo Su, Mingrong Duan, Minqiang Hu
IEEE Trans. Circuits Syst. I Regul. Pap.3
2023 Antisaturation and Power Decoupling Control of Multiwinding Energy Harvester Based on Magnetomotive Force Compensation
abstract
With the rapid development of smart grids, numerous monitoring sensors have been widely used for the state inspection of power lines. To ensure reliable operation of these sensors, the problem of their power source needs to be solved urgently. In particular, noninvasive inductive energy harvesters clamped over the lines are suitable as power sources for the monitoring sensors. This article presents the multiwinding energy harvester (MWEH), equipped with multiple energy output channels. To improve the antisaturation ability of the magnetic core and eliminate power cross effects between the windings simultaneously, an auxiliary winding is introduced to regulate the magnetomotive force. The method of maintaining the unsaturated working state of the magnetic core and the power decoupling control strategy between the windings are proposed. Hence, the global power fluctuation can be eliminated with a constant unsaturated state when the load and turns of any winding change. The experimental results suggest that regulating the energy harvesting factor of the auxiliary winding can desaturate the magnetic core. Furthermore, based on the linear operation of the MWEH, the problem of power linkage of the windings can also be solved, which stabilizes the output power of one or more windings.
Chenjin Xu, Wei Wang 0148, Wenbo Su, Mingrong Duan, Minqiang Hu
IEEE Trans. Ind. Informatics3
2022 GBA: A Tuning-free Approach to Switch between Synchronous and Asynchronous Training for Recommendation Models
abstract
High-concurrency asynchronous training upon parameter server (PS) architecture and high-performance synchronous training upon all-reduce (AR) architecture are the most commonly deployed distributed training modes for recommendation models. Although synchronous AR training is designed to have higher training efficiency, asynchronous PS training would be a better choice for training speed when there are stragglers (slow workers) in the shared cluster, especially under limited computing resources. An ideal way to take full advantage of these two training modes is to switch between them upon the cluster status. However, switching training modes often requires tuning hyper-parameters, which is extremely time- and resource-consuming. We find two obstacles to a tuning-free approach: the different distribution of the gradient values and the stale gradients from the stragglers. This paper proposes Global Batch gradients Aggregation (GBA) over PS, which aggregates and applies gradients with the same global batch size as the synchronous training. A token-control process is implemented to assemble the gradients and decay the gradients with severe staleness. We provide the convergence analysis to reveal that GBA has comparable convergence properties with the synchronous training, and demonstrate the robustness of GBA the recommendation models against the gradient staleness. Experiments on three industrial-scale recommendation tasks show that GBA is an effective tuning-free approach for switching. Compared to the state-of-the-art derived asynchronous training, GBA achieves up to 0.2% improvement on the AUC metric, which is significant for the recommendation models. Meanwhile, under the strained hardware resource, GBA speeds up at least 2.4x compared to synchronous training.
Wenbo Su, Yuanxing Zhang, Yufeng Cai, Kaixu Ren, Pengjie Wang 0002, Huimin Yi, Hongbo Deng, Jian Xu 0015, Lin Qu, Bo Zheng 0007
NeurIPS1
2015 Modeling and Analysis of Availability in Multi-Tenant SaaS
abstract
Software as a Service (SaaS) has become an important application development and service delivery model. Among different architectures, multi-tenant architecture (MTA) not only has advantage on maintenance, but also increases resource utilization by sharing instances. However, sharing instances brings challenges to the security of the service. As one of the three principal properties of the security, availability receives more and more attentions. Recently, there are extensive efforts on technical methods to implement a secure multi-tenant SaaS, but few works on the modeling and analysis of its availability. In this paper, we firstly present the availability issues of the multi-tenant SaaS. Two important mechanisms to implement the MTA SaaS are then introduced: network isolation and database sharing. After that, a stochastic Petri net (SPN) model is developed to analyze the availability. Specific metrics are proposed to measure the availability both from the aspects of the system and the tenant. To extend the SPN model for large scale analysis, we solve the state space explosion problem of SPN model based on the theory of Markov chain aggregation. Finally, numerical results are provided to demonstrate the effectiveness of the SPN model and the analysis is efficient.
Wenbo Su, Qu Liu, Chuang Lin 0002, Xuemin Shen
ICCCN1
2015 SLA-Aware Tenant Placement and Dynamic Resource Provision in SaaS
abstract
Software as a Service (SaaS) is an increasingly important service delivery model in cloud computing, and multitenancy makes it possible to support large scale customized tenants with only one code base. However, the complexity of multi-tenant architecture may lead to poor performance and low resource utilization. The customized demands may also lead to high operating cost. It is very important to develop an accurate model to predict the performance of the multi-tenant SaaS. To this end, a multi-tenant queueing network model is developed. Based on the model, a balanced SLA-aware tenant placement algorithm is proposed considering that customized tenants may need more resources to be placed together. The algorithm is effective in nearly 90% of the simulations comparing with other heuristic algorithms. Furthermore, the optimization problem on dynamic resource provision to minimize the operating cost is studied. As the original optimization problem is NPhard, a continuous upper bound is used to convert the original optimization problem into a convex optimization which can be solved efficiently in polynomial time. Finally, it is demonstrated that the approximate ratio of the proposed approach is no greater than 1.2 in more than 90% of the simulations.
Wenbo Su, Jie Hu 0003, Chuang Lin 0002, Xuemin Shen
ICWS1