Linge Li

dblp:211/0561 · DBLP profile ↗
← Back
7ranked-venue papers
1as first author
6since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2025 AKVQ-VL: Attention-Aware KV Cache Adaptive 2-Bit Quantization for Vision-Language Models
abstract
Vision-language models (VLMs) show remarkable performance in multimodal tasks. However, excessively long multimodal inputs lead to oversized Key-Value (KV) caches, resulting in significant memory consumption and I/O bottlenecks. Previous KV quantization methods for Large Language Models (LLMs) may alleviate these issues but overlook the attention saliency differences of multimodal tokens, resulting in suboptimal performance. In this paper, we investigate the attention-aware token saliency patterns in VLM and propose AKVQ-VL. AKVQ-VL leverages the proposed Text-Salient Attention (TSA) and Pivot-Token-Salient Attention (PSA) patterns to adaptively allocate bit budgets. Moreover, achieving extremely low-bit quantization requires effectively addressing outliers in KV tensors. AKVQ-VL utilizes the Walsh-Hadamard transform (WHT) to construct outlier-free KV caches, thereby reducing quantization difficulty. Evaluations of 2-bit quantization on 12 long-context and multimodal tasks demonstrate that AKVQ-VL maintains or even improves accuracy, outperforming LLM-oriented methods. AKVQ-VL can reduce peak memory usage by 2.13×, support up to 3.25× larger batch sizes and 2.46× throughput.
Zunhai Su, Wang Shen, Linge Li, Hanyu Wei, Huangqi Yu, Kehong Yuan
ICME3
2025 RotateKV: Accurate and Robust 2-Bit KV Cache Quantization for LLMs via Outlier-Aware Adaptive Rotations
abstract
Key-Value (KV) cache facilitates efficient large language models (LLMs) inference by avoiding recomputation of past KVs. As the batch size and context length increase, the oversized KV caches become a significant memory bottleneck, highlighting the need for efficient compression. Existing KV quantization rely on fine-grained quantization or the retention of a significant portion of high bit-widths caches, both of which compromise compression ratio and often fail to maintain robustness at extremely low average bit-widths. In this work, we explore the potential of rotation technique for 2-bit KV quantization and propose RotateKV, which achieves accurate and robust performance through the following innovations: (i) Outlier-Aware Rotation, which utilizes channel-reordering to adapt the rotations to varying channel-wise outlier distributions without sacrificing the computational efficiency of the fast Walsh-Hadamard transform (FWHT); (ii) Pre-RoPE Grouped-Head Rotation, which mitigates the impact of rotary position embedding (RoPE) on proposed outlier-aware rotation and further smooths outliers across heads; (iii) Attention-Sink-Aware Quantization, which leverages the massive activations to precisely identify and protect attention sinks. RotateKV achieves less than 0.3 perplexity (PPL) degradation with 2-bit quantization on WikiText-2 using LLaMA-2-13B, maintains strong CoT reasoning and long-context capabilities, with less than 1.7% degradation on GSM8K, outperforming existing methods even at lower average bit-widths. RotateKV also showcases a 3.97× reduction in peak memory usage, supports 5.75× larger batch sizes, and achieves a 2.32× speedup in decoding stage.
Zunhai Su, Hanyu Wei, Wang Shen, Linge Li, Huangqi Yu, Kehong Yuan
IJCAI5
2025 SAHLUT: Efficient Image Enhancement using Spatial-Aware High-Light Compensation Look-up Tables
abstract
Abstract Recently, the look‐up table (LUT)‐based method has achieved remarkable success in image enhancement tasks with its high efficiency and lightweight nature. However, when considering edge scenarios with limited computational resources, most existing methods fail to meet practical requirements due to their costly floating‐point operations on convolution layers, which limit their general use. Moreover, most LUT‐based methods may not perform well in handling high‐light regions. To address these issues, we propose SAHLUT, an efficient and practical image enhancement method by using spatial‐aware high‐light compensation look‐up tables (LUTs), which comprise two parts. Firstly, we propose a spatial‐aware weight predictor to reduce the computational burden. A lightweight network is trained to predict spatial‐aware weight values, and then we transfer the values to the LUTs. Additionally, to correct overexposure in high‐light regions, we propose a high‐light compensation 3D LUT. Our proposed method allows us to directly retrieve the values from the LUTs to achieve efficient image enhancement at test time. Extensive experimental results demonstrate that SAHLUT exhibits competitive performance compared to other LUT‐based methods both quantitatively and qualitatively in a more efficient manner. For instance, SAHLUT significantly reduces computational resources (at least 18 times in GFLOPs compared to other LUT‐based methods), while excelling in high‐light region handling.
Xin Chen 0104, Linge Li, Linhong Mu, Jingwei Guan
Comput. Graph. Forum2
2024 Foggy image restoration using deep sub-pixel reconstruction network
abstract
Abstract Light undergoes attenuation due to scattering and refraction when propagating through aerosols. In foggy conditions, Aerosol particles in the troposphere exhibit high mobility introducing intricate non‐linear noise into images. Foggy image restoration represents an ill‐posed problem, where traditional physical models and image enhancement techniques often prove inadequate in delivering effective solutions. This paper introduces a novel deep sub‐pixel reconstruction algorithm for foggy image restoration, pioneering the application of sub‐pixel reconstruction modules to this domain. This model employs convolutional layers to extract low‐level features and dense‐connected layers for high‐level feature extraction. Furthermore, a specialized sub‐pixel reconstruction module tailored for the task of foggy image restoration is designed, with the purpose of reconstructing dehazed images from latent vectors. During training, a generative adversarial training framework is adopted, incorporating a purpose‐designed discriminator. Additionally, a fusion loss is implemented to facilitate model refinement. Quantitative and qualitative evaluation experiments conducted on synthetic and real‐world image datasets demonstrate the effectiveness of the proposed method in preserving finer details. The Structural Similarity Index (SSIM) is observed to improve by 2.5%, attesting to enhanced perceptual quality for grayscale foggy images.
Linge Li, Feiyu Shi, Yihua Cai, Ping Fang, Chao Mu, Ningquan Weng
IET Image Process.1
2023 Towards Practical Consistent Video Depth Estimation
abstract
Monocular depth estimation algorithms aim to explore the possible links between 2D and 3D data, but challenges remain for existing methods to predict consistent depth from a casual video. Relying on camera poses and the optical flow in the time-consuming test-time training phases makes these methods fail in many scenarios and cannot be used for practical applications. In this work, we present a data-driven post-processing method to overcome these challenges and achieve online processing. Based on a deep recurrent network, our method takes the adjacent original and optimized depth map as inputs to learn temporal consistency from the dataset and achieves higher depth accuracy. Our approach can be applied to multiple single-frame depth estimation models and used for various real-world scenes in real-time. In addition, to tackle the lack of a temporally consistent video depth training dataset of dynamic scenes, we propose an approach to generate the training video sequences dataset from a single image based on inferring motion field. To the best of our knowledge, this is the first data-driven plug-and-play method to improve the temporal consistency of depth estimation for casual videos. Extensive experiments on three datasets and three depth estimation models show that our method outperforms the state-of-the-art methods.
Pengzhi Li, Yikang Ding, Linge Li, Jingwei Guan, Zhiheng Li 0001
ICMR3
2021 How Video Super-Resolution and Frame Interpolation Mutually Benefit
abstract
Video super-resolution (VSR) and video frame interpolation (VFI) are inter-dependent for enhancing videos of low resolution and low frame rate. However, most studies treat VSR and temporal VFI as independent tasks. In this work, we design a spatial-temporal super-resolution network based on exploring the interaction between VSR and VFI. The main idea is to improve the middle frame of VFI by the super-resolution (SR) frames and feature maps from VSR. In the meantime, VFI also provides extra information for VSR and thus, through interacting, the SR of consecutive frames of the original video can also be improved by the feedback from the generated middle frame. Drawing on this, our approach leverages a simple interaction of VSR and VFI and achieves state-of-the-art performance on various datasets. Due to such a simple strategy, our approach is universally applicable to any existing VSR or VFI networks for effectively improving their video enhancement performance.
Chengcheng Zhou, Zongqing Lu 0001, Linge Li, Qiangyu Yan, Jing-Hao Xue
ACM Multimedia3
2017 Efficient algorithms for HEVC bitrate transcoding
Linge Li, Guoming Zhi, Zuping Zhang 0001, Hao Zhang 0032
Multim. Tools Appl.2