EDBT 2026 Demo / reviewers in the wild / expert
Hanyu Wei
dblp:203/9734
· DBLP profile ↗
12ranked-venue papers
2as first author
10since 2021 · last 2026
0009-0009-2581-2519ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 3 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Network and information security
1 paper |
Cryptographic primitives and cryptanalysis · 100% | |
| Artificial intelligence
1 paper |
Efficient and distributed learning · 70% Language models and text generation · 30% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
GPUs and heterogeneous computing · 100% |
Topics — the 9 heaviest of 9, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Cryptographic primitives and cryptanalysis › post-quantum cryptography › lattice-based cryptography
falcon |
1.0 | 1 | 2026 | cuFalcon: An Adaptive Parallel GPU Implementation for High-Performance Falcon Acceleration · IEEE Trans. Parallel Distributed Syst. 2026 |
Cryptographic primitives and cryptanalysis › public-key cryptography › digital signatures › post-quantum signatures
lattice-based signatures |
1.0 | 1 | 2026 | cuFalcon: An Adaptive Parallel GPU Implementation for High-Performance Falcon Acceleration · IEEE Trans. Parallel Distributed Syst. 2026 |
Cryptographic primitives and cryptanalysis
post-quantum cryptography |
1.0 | 1 | 2026 | cuFalcon: An Adaptive Parallel GPU Implementation for High-Performance Falcon Acceleration · IEEE Trans. Parallel Distributed Syst. 2026 |
GPUs and heterogeneous computing › GPU computing
cryptographic acceleration |
1.0 | 1 | 2026 | cuFalcon: An Adaptive Parallel GPU Implementation for High-Performance Falcon Acceleration · IEEE Trans. Parallel Distributed Syst. 2026 |
GPUs and heterogeneous computing
GPU computing |
1.0 | 1 | 2026 | cuFalcon: An Adaptive Parallel GPU Implementation for High-Performance Falcon Acceleration · IEEE Trans. Parallel Distributed Syst. 2026 |
Machine learning › Efficient and distributed learning › model compression › quantization
KV cache quantization |
0.9 | 1 | 2025 | RotateKV: Accurate and Robust 2-Bit KV Cache Quantization for LLMs via Outlier-Aware Adaptive Rotations · IJCAI 2025 |
Natural language and speech › Language models and text generation › large language model inference
long-context inference |
0.9 | 1 | 2025 | RotateKV: Accurate and Robust 2-Bit KV Cache Quantization for LLMs via Outlier-Aware Adaptive Rotations · IJCAI 2025 |
Machine learning › Efficient and distributed learning
model compression |
0.9 | 1 | 2025 | RotateKV: Accurate and Robust 2-Bit KV Cache Quantization for LLMs via Outlier-Aware Adaptive Rotations · IJCAI 2025 |
Machine learning › Efficient and distributed learning › model compression › quantization
low-bit quantization |
0.3 | 1 | 2025 | RotateKV: Accurate and Robust 2-Bit KV Cache Quantization for LLMs via Outlier-Aware Adaptive Rotations · IJCAI 2025 |
Methods — techniques the papers use, named apart from their topics
adaptive parallelization · 2.0rotary position embedding · 0.9outlier-aware rotation · 0.9fast walsh-hadamard transform · 0.9attention-sink-aware quantization · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SDFP: Speculative Decoding with FIT-Pruned Models for Training-Free and Plug-and-Play LLM Acceleration
Hanyu Wei, Zunhai Su, Spandan Tiwari, Ashish Sirasao, Yuhan Dong |
ICIC (22) | 1 |
| 2026 | Optimizing CTRU for TLS 1.3 on Edge PlatformsabstractThe rise of quantum computing necessitates post-quantum cryptography (PQC) to secure Internet communications on edge platforms. CTRU is an NTRU-based key encapsulation mechanism (KEM) that features a small modulus and an efficient scaledE8lattice encoding, offering strong security and performance. However, two critical gaps impede its deployment for Internet of Things (IoT) edge platforms: the absence of an optimized implementation for ARMv8 devices, which dominate the mobile and embedded IoT ecosystems, and a lack of integration and evaluation in practical protocols such as Transport Layer Security (TLS) 1.3. To address these challenges, we present the first NEON-optimized CTRU implementation for ARMv8-A. Our implementation reduces CPU cycles by 3.24×, 2.62×, and 2.99× in key generation (KeyGen), encapsulation (Encaps), and decapsulation (Decaps) compared to the reference C (REF-C) implementation, while lowering energy consumption by 55.2%. Among all evaluated KEMs, it achieves the lowest energy consumption and a competitive memory footprint. We further design a batch key generation scheme that boosts KeyGen throughput by 3.73×. Finally, we integrate CTRU into TLS 1.3 to enable post-quantum (PQ) key exchange. Comprehensive benchmarks show that CTRU outperforms all evaluated KEMs and classical key exchange scheme Elliptic Curve Diffie-Hellman Secp384r1 (ECDH P384) on ARMv8-A platform. Specifically, it achieves lower TLS handshake latency and higher throughput under ideal network conditions, and its handshake latency distribution is superior to that of all alternatives under Narrow Band IoT (NB-IoT) constraints. These results demonstrate CTRU’s practical suitability for real-world deployment on IoT edge platforms. Zhuo Zhang 0027, Jieyu Zheng, Hanyu Wei, Yunlei Zhao |
IEEE Internet Things J. | 5 |
| 2026 | cuFalcon: An Adaptive Parallel GPU Implementation for High-Performance Falcon Acceleration
Hanyu Wei, Shiyu Shen 0001, Hao Yang 0062, Wangchen Dai, Yunlei Zhao |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2025 | AKVQ-VL: Attention-Aware KV Cache Adaptive 2-Bit Quantization for Vision-Language ModelsabstractVision-language models (VLMs) show remarkable performance in multimodal tasks. However, excessively long multimodal inputs lead to oversized Key-Value (KV) caches, resulting in significant memory consumption and I/O bottlenecks. Previous KV quantization methods for Large Language Models (LLMs) may alleviate these issues but overlook the attention saliency differences of multimodal tokens, resulting in suboptimal performance. In this paper, we investigate the attention-aware token saliency patterns in VLM and propose AKVQ-VL. AKVQ-VL leverages the proposed Text-Salient Attention (TSA) and Pivot-Token-Salient Attention (PSA) patterns to adaptively allocate bit budgets. Moreover, achieving extremely low-bit quantization requires effectively addressing outliers in KV tensors. AKVQ-VL utilizes the Walsh-Hadamard transform (WHT) to construct outlier-free KV caches, thereby reducing quantization difficulty. Evaluations of 2-bit quantization on 12 long-context and multimodal tasks demonstrate that AKVQ-VL maintains or even improves accuracy, outperforming LLM-oriented methods. AKVQ-VL can reduce peak memory usage by 2.13×, support up to 3.25× larger batch sizes and 2.46× throughput. Zunhai Su, Wang Shen, Linge Li, Hanyu Wei, Huangqi Yu, Kehong Yuan |
ICME | 5 |
| 2025 | RotateKV: Accurate and Robust 2-Bit KV Cache Quantization for LLMs via Outlier-Aware Adaptive RotationsabstractKey-Value (KV) cache facilitates efficient large language models (LLMs) inference by avoiding recomputation of past KVs. As the batch size and context length increase, the oversized KV caches become a significant memory bottleneck, highlighting the need for efficient compression. Existing KV quantization rely on fine-grained quantization or the retention of a significant portion of high bit-widths caches, both of which compromise compression ratio and often fail to maintain robustness at extremely low average bit-widths. In this work, we explore the potential of rotation technique for 2-bit KV quantization and propose RotateKV, which achieves accurate and robust performance through the following innovations: (i) Outlier-Aware Rotation, which utilizes channel-reordering to adapt the rotations to varying channel-wise outlier distributions without sacrificing the computational efficiency of the fast Walsh-Hadamard transform (FWHT); (ii) Pre-RoPE Grouped-Head Rotation, which mitigates the impact of rotary position embedding (RoPE) on proposed outlier-aware rotation and further smooths outliers across heads; (iii) Attention-Sink-Aware Quantization, which leverages the massive activations to precisely identify and protect attention sinks. RotateKV achieves less than 0.3 perplexity (PPL) degradation with 2-bit quantization on WikiText-2 using LLaMA-2-13B, maintains strong CoT reasoning and long-context capabilities, with less than 1.7% degradation on GSM8K, outperforming existing methods even at lower average bit-widths. RotateKV also showcases a 3.97× reduction in peak memory usage, supports 5.75× larger batch sizes, and achieves a 2.32× speedup in decoding stage. Zunhai Su, Hanyu Wei, Wang Shen, Linge Li, Huangqi Yu, Kehong Yuan |
IJCAI | 2 |
| 2025 | OSKR/OKAI: Systematic Optimization of Key Encapsulation Mechanisms from Module Lattice
Shiyu Shen 0001, Zhichuang Liang, Jieyu Zheng, Hanyu Wei, Yang Wang 0050, Zhenfeng Zhang, Yunlei Zhao |
J. Comput. Sci. Technol. | 6 |
| 2024 | Continual MARL-assisted Traffic-Adaptive GC-MAC Protocol for Task-Driven Directional MANETsabstractFaced with the abundance of mobile ad-hoc network (MANET) applications, there emerges a strong incentive to provision MANET in millimeter wave. However, the deafness of directional antennas and the decentralized structure of MANET make it difficult to achieve consistent medium access control (MAC) among nodes through random competition. Therefore, graph coloring-based MAC (GC-MAC) scheme is proposed to implement time division multiplexed scheduling, but it allocates equal slots to links, disregarding the unbalanced and piecewise stationary traffic distribution in task-driven MANETs. Here, we propose a continual multi-agent reinforcement learning (RL)-assisted traffic-adaptive scheme to enhance the agility of slot allocation. Specifically, we add a contention period to frames of GC-MAC, during which nodes analyze stochastic characteristics of traffic to derive link traffic distribution for each task, and adjust slot assignment to reach cooperation through decentralized multi-agent deep Q-network (MA-DQN). Besides, considering the traffic variation due to task switch, continual RL is incorporated to accommodate changes of the environment more sensitively. Finally, simulation results prove the proposed scheme achieves faster convergence speed, lower delay and higher throughput. Chan Wang, Rongpeng Li, Hanyu Wei, Minjian Zhao |
VTC Fall | 4 |
| 2024 | Swinv2-Imagen: hierarchical vision transformer diffusion models for text-to-image generationabstractAbstract Recently, diffusion models have been proven to perform remarkably well in text-to-image synthesis tasks in a number of studies, immediately presenting new study opportunities for image generation. Google’s Imagen follows this research trend and outperforms DALLE2 as the best model for text-to-image generation. However, Imagen merely uses a T5 language model for text processing, which cannot ensure learning the semantic information of the text. Furthermore, the Efficient UNet leveraged by Imagen is not the best choice in image processing. To address these issues, we propose the Swinv2-Imagen, a novel text-to-image diffusion model based on a Hierarchical Visual Transformer and a Scene Graph incorporating a semantic layout. In the proposed model, the feature vectors of entities and relationships are extracted and involved in the diffusion model, effectively improving the quality of generated images. On top of that, we also introduce a Swin-Transformer-based UNet architecture, called Swinv2-Unet, which can address the problems stemming from the CNN convolution operations. Extensive experiments are conducted to evaluate the performance of the proposed model by using three real-world datasets, i.e. MSCOCO, CUB and MM-CelebA-HQ. The experimental results show that the proposed Swinv2-Imagen model outperforms several popular state-of-the-art methods. Ruijun Li, Weihua Li 0007, Yi Yang 0036, Hanyu Wei, Jianhua Jiang, Quan Bai 0001 |
Neural Comput. Appl. | 4 |
| 2023 | Matrix Factorization and Deep Autoencoder based Clustering Scheme for Large-scale UAV NetworksabstractIn recent years, unmanned aerial vehicle (UAV) ad-hoc networks have achieved rapid development due to their autonomy and high reliability. Typically, clustering is widely adopted to reduce the degradation of network performance in large-scale UAV networks. However, due to the limited channel resources and complex communication environment, it becomes challenging to obtain complete information needed for clustering from remote nodes. In addition, the high dimension and non-linear relationships of data also make large-scale clustering diffcult. Therefore, in this paper, we propose a noval distributed clustering scheme. First, a matrix factorization (MF) based time series algorithm is proposed to predict and complement the incomplete information. Second, we adopt deep autoencoder to wisely incorporate the non-linear relationship of information needed for clustering. Finally we extract features as input to k-means to obtain clustering results. Simulation results demonstrate the effectiveness of the proposed clustering scheme. Jiaolan Fang, Chan Wang, Rongpeng Li, Hanyu Wei, Minjian Zhao |
VTC2023-Spring | 4 |
| 2022 | Mean-field MARL-based Priority-Aware CSMA/CA Strategy in Large-Scale MANETsabstractMobile ad-hoc network (MANET) has attracted ex-tensive attention in many applications with nodes operating over bandwidth-constrained wireless links in a self-organized manner. Typically, to avoid potential collisions, CSMA/CA method is widely adopted in MANETs and affects the provisioning quality of service (QoS). However, due to the potentially severe collisions in a large-scale MANET, current CSMA/CA methods fail to support QoS effectively, especially for scenarios with different priority services. Here, we propose a priority-aware CSMA/CA strategy to use the higher-priority packets to collect the queueing information from adjacent nodes in a piggyback manner and model the channel access process by a partial observation Markov decision process (POMDP). Furthermore mean-field multi-agent reinforcement learning (MARL) is adopted to allow each individual node in the MANET to dynamically select the appropriate contention window (CW) and gradually compute the global optimal policy. Finally, extensive simulation results verify significantly lower delay and packet loss rate for the higher- priority services and comparable provisioning quality for lower- priority services, which reflects the superiority over baselines. Hanyu Wei, Chan Wang, Rongpeng Li, Minjian Zhao |
GLOBECOM | 1 |
| 2020 | Predicting Internet of Things Data Traffic Through LSTM and Autoregressive Spectrum AnalysisabstractThe rapid increase of Internet of Things (IoT) applications and services has led to massive amounts of heterogeneous data. Hence, we need to re-think how IoT data influences the network. In this paper, we study the characteristics of IoT data traffic in the context of smart cities. Aiming at analyzing the influence of IoT data traffic on the access and core network, we generate various IoT data traffic according to the characteristics of different IoT applications. Based on the analysis of the inherent features of the aggregated IoT data traffic, we propose a Long Short-Term Memory (LSTM) model combined with autoregressive spectrum analysis to predict the IoT data traffic. In this model, the autoregressive spectrum analysis is used to estimate the minimum length of the historical data needed for predicting the traffic in the future, which alleviates LSTM’s performance deterioration with the increase of sequence length. A sliding window enables predicting the long¬term tendency of IoT data traffic while keeping the inherent features of the data traffic. The evaluation results show that the proposed model converges quickly and can predict the variations of IoT traffic more accurately than other methods and the general LSTM model. Bailin Wang, Xiang Su 0001, Jukka Riekki, Hanyu Wei |
NOMS | 7 |
| 2017 | Gamma-modulated Wavelet model for Internet of Things trafficabstractPromoted by sensor, big data and mobile computing technologies, the number of Internet of Things (IoT) applications and services is increasing rapidly. The massive amounts of heterogeneous data produced by a large variety of IoT devices require us to re-think its influence on the network. In this paper, we study the characteristics of IoT data traffic in the context of smart city. We generate data traffic according to the characteristics of different IoT applications. We propose a Gamma modulated wavelet method for statistical characterization of both IoT data and the aggregated traffic, aiming at analyzing the influence of IoT data traffic on the access and core network. By using Gamma function to modulate the coefficients of the wavelet, both the long range and short range dependency of the IoT data traffic can be described through fewer parameters. The Gamma modulation also reduces the independency of the coefficients and improves the accuracy of the Wavelet model. Xiang Su 0001, Jukka Riekki, Huber Flores, Hanyu Wei |
ICC | 7 |