Jinyu Wang 0002

dblp:63/6030-2 · DBLP profile ↗
← Back
10ranked-venue papers
1as first author
10since 2021 · last 2026
0000-0002-9449-3453ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 1 first-author · 5 since 2021Computer networks · 3 · 3 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 KVmix: Gradient-Based Layer Importance-Aware Mixed-Precision Quantization for KV Cache
abstract
The high memory demands of the Key-Value (KV) Cache during the inference of Large Language Models (LLMs) severely restrict their deployment in resource-constrained platforms. Quantization can effectively alleviate the memory pressure caused by KV Cache. However, existing methods either rely on static one-size-fits-all precision allocation or fail to dynamically prioritize critical KV in long-context tasks, forcing memory-accuracy-throughput tradeoffs. In this work, we propose a novel mixed-precision quantization method for KV Cache named KVmix. KVmix leverages gradient-based importance analysis to evaluate how individual Key and Value projection matrices affect the model loss, enabling layer-specific bit-width allocation for mix-precision quantization. It dynamically prioritizes higher precision for important layers while aggressively quantizing less influential ones, achieving a tunable balance between accuracy and efficiency. KVmix introduces a dynamic long-context optimization strategy that adaptively keeps full-precision KV pairs for recent pivotal tokens and compresses older ones, achieving high-quality sequence generation with low memory usage. Additionally, KVmix provides efficient low-bit quantization and CUDA kernels to optimize computational overhead. On LLMs such as Llama and Mistral, KVmix achieves near-lossless inference performance with extremely low quantization configuration (Key 2.19bit Value 2.38bit), while delivering a remarkable 4.9× memory compression and a 5.3× speedup in inference throughput.
Fei Li 0042, Song Liu 0007, Weiguo Wu, Shiqiang Nie, Jinyu Wang 0002
AAAI5
2026 EADA: Efficient adaptive data augmentation
Song Liu 0007, Weiguo Wu, Jinyu Wang 0002, Shiqiang Nie
Comput. Vis. Image Underst.5
2026 BLSA: A cache-aware balanced load scheduling approach on task graphs
Song Liu 0007, Fei Li 0042, Shiqiang Nie, Jinyu Wang 0002, Weiguo Wu
Future Gener. Comput. Syst.5
2026 FDSR: Efficient Model Training via Adaptive Tensor Quantization Based on Frequency Domain Division and Similarity Data Reuse
abstract
As deep neural networks (DNNs) continue to grow in scale and complexity, GPU memory limitations have become a significant challenge for DNN model training, especially on resource-constrained commercial GPUs. While model quantization facilitates memory-efficient training, it often necessitates a tradeoff between quantization granularity and model accuracy. And quantization imposes additional computational overhead, which adversely affects the training throughput and apportions out the performance gains it brings. In this article, we propose FDSR, an adaptive tensor quantization method that leverages frequency domain division and similarity-based data reuse to break the memory bottleneck in visual model training. FDSR leverages the frequency-domain characteristics of tensors in terms of memory consumption and model accuracy, and proposes a fine-grained tensor quantization with different quantization bit-widths. It adaptively optimizes the quantization parameters according to model accuracy during training while employing sparsification according to data frequency-domain features, minimizing memory consumption and accuracy loss. To counteract the computational cost, FDSR incorporates a novel similarity-based reuse strategy that avoids redundant quantization/dequantization computations, further enhanced by a tailored Locality-Sensitive Hashing (LSH) mechanism and optimized kernels. Experimental results demonstrate that FDSR achieves an average of 10.20× activation memory compression with only 1.10% average accuracy loss across various models on the commercial GPU. Compared to the state-of-the-art quantization methods, FDSR improves memory optimization by up to 68.6% and increases throughput by up to 25.55%, with consistent performance improvements on different GPU architectures.
Song Liu 0007, Fei Li 0042, Qin Xia, Shiqiang Nie, Jinyu Wang 0002, Weiguo Wu
ACM Trans. Archit. Code Optim.6
2026 Efficient Headroom Allocation With Two-Level Flow Control for Lossless Datacenter Networks
abstract
In datacenters, lossless network is very attractive as it can achieve ultra-low latency. In commodity Ethernet, lossless forwarding is achieved by hop-by-hop Priority-based Flow Control (PFC). To avoid buffer overflow, PFC-enabled switches need to reserve some buffer asheadroom, absorbing in-flight packets during the delay for backpressure messages to take effect. However, with the growing link speed in production networks, the buffer becomes increasingly insufficient, and the headroom can occupy a considerable fraction of buffer. As a result, the remaining buffer for absorbing normal traffic bursts is significantly squeezed, leading to frequent PFC messages that degrade the network performance. Worse yet, we find that the current static and queue-independent headroom allocation scheme is quite inefficient, resulting in significant buffer wastage. In light of this, we propose Dynamic and Shared Headroom allocation scheme (DSH), which dynamically allocates headroom to congested queues and enables sharing of allocated headroom among different queues. To achieve this, DSH first introduces port-level flow control, which performs flow control at the granularity of individual ports, guaranteeing lossless forwarding with a small fraction of per-port headroom. With this lossless guarantee, the switch is liberated for dynamic headroom adjustment. DSH dynamically allocates per-queue headroom based on the congestion status of each queue. Meanwhile, DSH preserves the queue-level flow control to protect the non-congested queues from being paused by congested queues, ensuring performance isolation on buffer sharing. Extensive experiments show that DSH can reduce the flow completion time by up to ~78.8%.
Danfeng Shan, Jinchao Ma, Yunguang Li, Boxuan Hu, Tong Zhang 0018, Yazhe Tang, Hao Li 0011, Jinyu Wang 0002, Peng Zhang 0011
IEEE Trans. Netw.8
2025 Olsync: Object-level tiering and coordination in tiered storage systems based on software-defined network
Zhike Li, Shiqiang Nie, Jinyu Wang 0002, Chi Zhang 0095, Fangxing Yu, Zhankun Zhang, Song Liu 0007, Weiguo Wu
Future Gener. Comput. Syst.4
2023 An efficient computation offloading and resource allocation algorithm in RIS empowered MEC
Xiangjun Zhang, Weiguo Wu, Song Liu 0007, Jinyu Wang 0002
Comput. Commun.4
2023 MCB: a multidevice cooperative buffer management strategy for boosting the write performance of the SSD-SMR hybrid storage
Chi Zhang 0095, Shiqiang Nie, Jinyu Wang 0002, Song Liu 0007, Weiguo Wu
J. Supercomput.3
2023 BiLSTM-based Federated Learning Computation Offloading and Resource Allocation Algorithm in MEC
abstract
Mobile edge computing (MEC) driven by 5G cellular systems has recently emerged as a promising paradigm, enabling mobile devices (MDs) with limited computing resources to offload various computation-intensive tasks (such as autopilot, online game) to edge servers to enhance the data processing capabilities of MDs. However, the uncertainty of wireless channel state and data volume of offloading tasks, as well as the data security privacy of offloading tasks, bring serious challenges to computation offloading in MEC. In this article, we consider a time-varying MEC scenario and formalize the delay and energy consumption during the computation offloading process as a joint optimization problem. Then the optimization problem is decomposed into two sub-problems: intelligent task prediction and resource allocation. Different from traditional methods, we improve the federated learning (FL) algorithm and propose a thoughtful cloud-edge-client FL task prediction mechanism based on Bidirectional Long Short-Term Memory. Each participating MD trains the model locally without uploading data to the server, and periodically aggregates the model in the edge and in the cloud. The algorithm both eliminates the need to solve complex optimization problems and ensures user privacy security. Finally, experimental results show that our proposed algorithm significantly outperforms other benchmark algorithms in energy efficiency.
Xiangjun Zhang, Weiguo Wu, Jinyu Wang 0002, Song Liu 0007
ACM Trans. Sens. Networks3
2021 DUPRFloor: Dynamic Modeling and Floorplanning for Partially Reconfigurable FPGAs
abstract
Nowadays, field-programmable gate array (FPGA) devices have been widely used in various fields. However, modules of circuits to be executed on FPGAs are placed within rectangular reconfigurable regions (RRs) with current floorplanners, leading to internal fragments, and lower utilization of resources. To address this, a dynamic description model of RRs and the corresponding floorplanner named dynamic union partial reconfiguration floorplan (DUPRFloor) are proposed in this article. The RR dynamic description modeling adds an anchor within a rectangular RR to reduce internal fragments. In this way, the modeling can represent both rectangular and nonrectangular shapes. Then, to find the optimal anchor, a clipping method is devised by constraining width and height of the candidate region. Finally, the mixed-integer linear programming (MILP) is used to optimize an objective function which considers the resources utilization and communication costs to obtain a desirable floorplanning result. The proposed method has been validated by simulation on three kinds of devices. And experimental results show that reconfigurable resources can be saved as much as 19.16% compared to rectangular modeling method. The DUPRFloor is also validated on the Microelectronics Center of North Carolina standard benchmark data sets. Results show that DUPRFloor can reduce 18.65% global wire length at most with almost the same execution time compared to state-of-the-art algorithms. Our approach is tested on a FPGA implemented software-defined radio (SDR) and reduced 29.41% wasted configurable frames, and to the overall design, 2% configurable frames are saved at most.
Jinyu Wang 0002, Yifei Kang, Weiguo Wu, Guoliang Xing, Linlin Tu
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1