Ziwei He

dblp:213/1850 · DBLP profile ↗
← Back
9ranked-venue papers
3as first author
8since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-authorTheory of computation · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 LongLLaDA: Unlocking Long Context Capabilities in Diffusion LLMs
abstract
Large Language Diffusion Models, or dLLMs, have emerged as a significant focus in NLP research, with substantial effort directed toward understanding their scalability and downstream task performance. However, their long-context capabilities remain unexplored, lacking systematic analysis or methods for context extension. In this work, we present the first systematic investigation comparing the long-context performance of diffusion LLMs and traditional auto-regressive LLMs. We first identify a unique characteristic of dLLMs, unlike auto-regressive LLMs, they maintain remarkably ***stable perplexity*** during direct context extrapolation. Moreover, where auto-regressive models fail outright during the Needle-In-A-Haystack task with context exceeding their pretrained length, we discover dLLMs exhibit a distinct ***local perception*** phenomenon, enabling successful retrieval from recent context segments. We explain both phenomena through the lens of Rotary Position Embedding (RoPE) scaling theory. Building on these observations, we propose LongLLaDA, a training-free method that integrates LLaDA with the NTK-based RoPE extrapolation. Our results validate that established extrapolation scaling laws remain effective for extending the context windows of dLLMs. Furthermore, we identify long-context tasks where dLLMs outperform auto-regressive LLMs and others where they fall short. Consequently, this study establishes the first length extrapolation method for diffusion LLMs while providing essential theoretical insights and empirical benchmarks critical for advancing future research on long-context diffusion LLMs.
Yuerong Song, Zhigeng Liu, Zengfeng Huang, Qipeng Guo, Ziwei He, Xipeng Qiu
AAAI6
2026 Sparse-dLLM: Accelerating Diffusion LLMs with Dynamic Cache Eviction
abstract
Diffusion Large Language Models (dLLMs) enable breakthroughs in reasoning and parallel decoding but suffer from prohibitive quadratic computational complexity and memory overhead during inference. Current caching techniques accelerate decoding by storing full-layer states, yet impose substantial memory usage that limit long-context applications. Our analysis of attention patterns in dLLMs reveals persistent cross-layer sparsity, with pivotal tokens remaining salient across decoding steps and low-relevance tokens staying unimportant, motivating selective cache eviction. We propose Sparse-dLLM, the first training-free framework integrating dynamic cache eviction with sparse attention via delayed bidirectional sparse caching. By leveraging the stability of token saliency over steps, it retains critical tokens and dynamically evicts unimportant prefix/suffix entries using an attention-guided strategy. Extensive experiments on LLaDA and Dream series demonstrate Sparse-dLLM achieves up to 10 times higher throughput than vanilla dLLMs, with comparable performance and similar peak memory costs, outperforming previous methods in efficiency and effectiveness.
Yuerong Song, Ruixiao Li, Zhigeng Liu, Zengfeng Huang, Qipeng Guo, Ziwei He, Xipeng Qiu
AAAI7
2025 WeightedKV: Attention Scores Weighted Key-Value Cache Merging for Large Language Models
abstract
Large Language Models (LLMs) use key-value (KV) cache to reduce redundant computation in autoregressive generation. However, the KV cache size increases linearly during generation, leading to excessive memory usage, especially for long texts. Most KV cache compression methods evict the unimportant KV pairs to maintain a fixed cache size, which leads to the permanent loss of tokens during generation. However, singular value decomposition shows that values do not exhibit a strong low-rank property as keys do, suggesting that information is distributed more evenly across values, in contrast to its more redundant distribution within keys. Therefore, methods that evict both keys and values risk losing crucial information and compromise context integrity, ultimately degrading the output quality. To address this problem, we propose WeightedKV, a novel, training-free approach that discards the keys of less important tokens, while merging their values into neighboring tokens via a convex combination weighted by their average attention scores. In this way, the retained keys serve as anchors that guide the generation process, while the merged values provide a rich contextual backdrop. We assess our method on four widely used language modeling datasets, demonstrating superior performance compared to all baseline methods, particularly with a lower budget ratio.
Ziwei He, Haoli Bai, Jingwen Leng
ICASSP2
2025 TreeKV: Smooth Key-Value Cache Compression with Tree Structures
abstract
Efficient key-value (KV) cache compression is critical for scaling transformer-based Large Language Models (LLMs) in long sequences and resource-limited settings. Existing methods evict tokens based on their positions or importance, but position-based strategies can miss crucial information outside predefined regions, while those relying on global importance scores resulting in strong regional biases, limiting the KV cache's overall context retention and potentially impairing the performance of LLMs on complex tasks. Our wavelet analysis reveals that as tokens approach the end of sequence, their contributions to generation gradually increase and tends to diverge more from neighboring tokens, indicating a smooth transition with increasing complexity and variability from distant to nearby context. Motivated by this observation, we propose TreeKV, an intuitive, training-free method that employs a tree structure for smooth cache compression. TreeKV maintains a fixed cache size, allowing LLMs to deliver high-quality output in long text scenarios and is applicable during both the generation and prefilling stages. TreeKV consistently surpasses all baseline models in language modeling tasks on PG19 and OpenWebText2, allowing LLMs trained with short context window to generalize to longer window with a 16x cache reduction. On the Longbench benchmark, TreeKV achieves the best performance with only 6% of the budget at optimal efficiency.
Ziwei He, Haoli Bai, Jingwen Leng
IJCAI1
2024 Towards Controlled Table-to-Text Generation with Scientific Reasoning
abstract
The sheer volume of scientific experimental results and complex technical statements, often presented in tabular formats, presents a formidable barrier to individuals acquiring preferred information. The realms of scientific reasoning and content generation that adhere to user preferences encounter distinct challenges. In this work, we present a new task for generating fluent and logical descriptions that match user preferences over scientific tabular data, aiming to automate scientific document analysis. To facilitate research in this direction, we construct a new challenging dataset CTRLSciTab consisting of table-description pairs extracted from the scientific literature, with highlighted cells and corresponding domain-specific knowledge base. We evaluated popular pre-trained language models to establish a baseline and proposed a novel architecture outperforming competing approaches. The results showed that large models struggle to produce accurate content that aligns with user preferences. As the first of its kind, our work should motivate further research in scientific domains.1
Zhixin Guo, Jianping Zhou 0004, Jiexing Qi, Mingxuan Yan, Ziwei He, Guanjie Zheng, Zhouhan Lin, Xinbing Wang, Chenghu Zhou
ICASSP5
2024 Fovea Transformer: Efficient Long-Context Modeling with Structured Fine-To-Coarse Attention
abstract
The quadratic complexity of self-attention in Transformers has hindered the processing of long text. To alleviate this problem, previous works have proposed to sparsify the attention matrix, taking advantage of the observation that crucial information about a token can be derived from its neighbors. These methods typically combine one or another form of local attention and global attention. Such combinations introduce abrupt changes in contextual granularity when going from local to global, which may be undesirable. We believe that a smoother transition could potentially enhance model’s ability to capture long-context dependencies. In this study, we introduce Fovea Transformer, a long-context focused transformer that addresses the challenges of capturing global dependencies while maintaining computational efficiency. To achieve this, we construct a multi-scale tree from the input sequence, and use representations of context tokens with a progressively coarser granularity in the tree, as their distance to the query token increases. We evaluate our model on three long-context summarization tasks1. It achieves state-of-the-art performance on two of them, and competitive results on the third with mixed improvement and setback of the evaluation metrics.
Ziwei He, Jingwen Leng
ICASSP1
2023 Cloud-Edge Collaboration-Based Distribution Network Reconfiguration for Voltage Preventive Control
abstract
The distribution network reconfiguration (DNR) can realize the voltage preventive control of the distribution network (DN) under alert state, which solves the voltage security issues and maintains the safe operation of the DN. However, since the great scale and complexity of the DN would make the traditional centralized DN reconfiguration method have a heavy computing burden, a DNR method based on cloud-edge collaborative architecture for voltage preventive control is proposed in this article, which can reduce the huge computational pressure caused by the excessive concentration of computing tasks. In order to formulate the optimal topology reconfiguration strategy according to the specifics of voltage alerts, a differential hybrid Petri-net model with event-triggered strategy is constructed based on the cloud-edge collaborative architecture to characterize the logical relations of the solution process of the reconfiguration strategy. Considering the short time scale characteristic of preventive control, corresponding to the solution process described by the constructed model, an evaluation network based on the graph convolutional neural network (GCN) is proposed for the reachability discrimination of solutions to significantly reduce the number of candidate solutions, as well as a decision network based on multilayer perceptron is proposed for the selection of the optimal solution among the reachable solutions. Numerical tests are conducted on the modified IEEE 33-bus and IEEE 118-bus distribution systems to validate the effectiveness of the proposed method in dealing with voltage alert problems.
Dong Yue 0001, Ziwei He, Chun-xia Dou
IEEE Trans. Ind. Informatics2
2022 RASAT: Integrating Relational Structures into Pretrained Seq2Seq Model for Text-to-SQL
abstract
Jiexing Qi, Jingyao Tang, Ziwei He, Xiangpeng Wan, Yu Cheng, Chenghu Zhou, Xinbing Wang, Quanshi Zhang, Zhouhan Lin. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022.
Jiexing Qi, Ziwei He, Xiangpeng Wan, Yu Cheng 0003, Chenghu Zhou, Xinbing Wang, Quanshi Zhang, Zhouhan Lin
EMNLP3
2017 A Comparative Study of Different Approaches for Tracking Communities in Evolving Social Networks
abstract
In real-world social networks, there is an increasing interest in tracking the evolution of groups of users and detecting the various changes they are liable to undergo. Several approaches have been proposed for this. In studying these approaches, we observed that most of them use a two-stage process. In the first stage, they run an algorithm to identify groups of users at each timestamp. In the second stage, a pairwise comparison based on a similarity measure is employed to track groups of users and detect changes they may undergo. While the majority of existing approaches use a two-stage process, they all run different algorithms to identify communities and rely on different similarity measures to track groups of users over time. Noting that the different approaches may perform differently depending on the dynamic social network under investigation, we decided to make a high level survey of some existing tracking approaches and then do a comparative analysis of some of them. In our analysis, we compared the algorithms in two main situations: (1) when groups of users do not overlap and (2) when the groups are overlapping. The study was done on three different testbeds extracted from the DBLP, Autonomous System (AS) and Yelp datasets.
Ziwei He, Etienne Gael Tajeuna, Shengrui Wang, Mohamed Bouguessa
DSAA1