Akide Liu

dblp:323/5463 · DBLP profile ↗
← Back
7ranked-venue papers
2as first author
7since 2021 · last 2026
0000-0001-9870-8303ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 2 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Efficient and distributed learning · 55% Generative modeling · 22% Deep learning architectures and training · 11%
Computer graphics and multimedia
2 papers
Rendering · 48% Visual content generation and editing · 28% Image and video coding · 24%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Hardware accelerators and domain-specific architectures · 100%

Topics — the 18 heaviest of 19, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning › KV cache management
KV cache compression
1.822026
OmniSparse: Training-Aware Fine-Grained Sparse Attention for Long-Video MLLMs · AAAI 2026
MiniCache: KV Cache Compression in Depth Dimension for Large Language Models · NeurIPS 2024
Machine learning › Efficient and distributed learning
model compression
1.622025
FPSAttention: Training-Aware FP8 and Sparsity Co-Design for Fast Video Diffusion · NeurIPS 2025
MiniCache: KV Cache Compression in Depth Dimension for Large Language Models · NeurIPS 2024
Machine learning › Efficient and distributed learning
inference acceleration
1.012026
OmniSparse: Training-Aware Fine-Grained Sparse Attention for Long-Video MLLMs · AAAI 2026
Machine learning › Deep learning architectures and training › attention mechanism
sparse attention
1.012026
OmniSparse: Training-Aware Fine-Grained Sparse Attention for Long-Video MLLMs · AAAI 2026
Visual content generation and editing › image editing › text-guided image editing
instruction-based image editing
1.012026
MCIE: Multimodal LLM-Driven Complex Instruction Image Editing with Spatial Guidance · AAAI 2026
Machine learning › Generative modeling
diffusion model
0.912025
FPSAttention: Training-Aware FP8 and Sparsity Co-Design for Fast Video Diffusion · NeurIPS 2025
Machine learning › Efficient and distributed learning › model quantization
FP8 quantization
0.912025
FPSAttention: Training-Aware FP8 and Sparsity Co-Design for Fast Video Diffusion · NeurIPS 2025
Machine learning › Efficient and distributed learning › model compression
quantization
0.912025
FPSAttention: Training-Aware FP8 and Sparsity Co-Design for Fast Video Diffusion · NeurIPS 2025
Machine learning › Generative modeling
video generation
0.912025
FPSAttention: Training-Aware FP8 and Sparsity Co-Design for Fast Video Diffusion · NeurIPS 2025
Rendering › gaussian splatting
3d gaussian splatting
0.912025
ZPressor: Bottleneck-Aware Compression for Scalable Feed-Forward 3DGS · NeurIPS 2025
Image and video coding › image compression
multi-view image compression
0.912025
ZPressor: Bottleneck-Aware Compression for Scalable Feed-Forward 3DGS · NeurIPS 2025
Rendering
novel view synthesis
0.912025
ZPressor: Bottleneck-Aware Compression for Scalable Feed-Forward 3DGS · NeurIPS 2025
Hardware accelerators and domain-specific architectures › machine learning accelerator › transformer accelerator
attention acceleration
0.912025
FPSAttention: Training-Aware FP8 and Sparsity Co-Design for Fast Video Diffusion · NeurIPS 2025
Hardware accelerators and domain-specific architectures
machine learning accelerator
0.912025
FPSAttention: Training-Aware FP8 and Sparsity Co-Design for Fast Video Diffusion · NeurIPS 2025
Natural language and speech › Language models and text generation
large language model inference
0.812024
MiniCache: KV Cache Compression in Depth Dimension for Large Language Models · NeurIPS 2024
Machine learning › Generative modeling
motion generation
0.812024
Motion Mamba: Efficient and Long Sequence Motion Generation · ECCV (1) 2024
Computer vision › 3D vision
3d scene understanding
0.312025
ZPressor: Bottleneck-Aware Compression for Scalable Feed-Forward 3DGS · NeurIPS 2025
Machine learning › Deep learning architectures and training › sequence modeling
efficient sequence modeling
0.212024
Motion Mamba: Efficient and Long Sequence Motion Generation · ECCV (1) 2024

Methods — techniques the papers use, named apart from their topics

cross-attention · 2.7sparsity · 2.5latent compression · 1.7information bottleneck · 1.7flashattention · 1.73d bi-directional attention · 1.7query selection · 1.0multimodal large language model · 1.0dynamic token budget allocation · 1.0diffusion model · 1.0KV selection · 1.0KV cache slimming · 1.0
YearPublicationVenuePosition
2026 MCIE: Multimodal LLM-Driven Complex Instruction Image Editing with Spatial Guidance
abstract
Recent advances in instruction-based image editing have shown remarkable progress. However, existing methods remain limited to relatively simple editing operations, hindering real-world applications that require complex and compositional instructions. In this work, we address these limitations from the perspectives of architectural design, data, and evaluation protocols. Specifically, we identify two key challenges in current models: insufficient instruction compliance and background inconsistency. To this end, we propose MCIE-E1, a Multimodal Large Language Model–Driven Complex Instruction Image Editing method that integrates two key modules: a spatial-aware cross-attention module and a background-consistent cross-attention module. The former enhances instruction-following capability by explicitly aligning semantic instructions with spatial regions through spatial guidance during the denoising process, while the latter preserves features in unedited regions to maintain background consistency. To enable effective training, we construct a dedicated data pipeline to mitigate the scarcity of complex instruction-based image editing datasets, combining fine-grained automatic filtering via a powerful MLLM with rigorous human validation. Finally, to comprehensively evaluate complex instruction-based image editing, we introduce CIE-Bench, a new benchmark with two new evaluation metrics. Experimental results on CIE-Bench demonstrate that MCIE-E1 consistently outperforms previous state-of-theart methods in both quantitative and qualitative assessments, achieving a 23.96% improvement in instruction compliance.
Xuehai Bai, Xiaoling Gu, Akide Liu, Hangjie Yuan, YiFan Zhang, Jack Ma
AAAI3
2026 OmniSparse: Training-Aware Fine-Grained Sparse Attention for Long-Video MLLMs
abstract
Existing sparse attention methods primarily target inference-time acceleration by selecting critical tokens under predefined sparsity patterns. However, they often fail to bridge the training–inference gap and lack the capacity for fine-grained token selection across multiple dimensions—such as queries, key-values (KV), and heads—leading to suboptimal performance and acceleration gains. In this paper, we introduce OmniSparse, a training-aware fine-grained sparse attention of long-video MLLMs, which is applied in both training and inference with dynamic token budget allocation. Specifically, OmniSparse contains three adaptive and complementary mechanisms: (1) query selection as lazy-active classification, aiming to retain active queries that capture broader semantic similarity, while discarding most of lazy ones that focus on limited local context and exhibit high functional redundancy with their neighbors, (2) KV selection with head-level dynamic budget allocation, where a shared budget is determined based on the flattest head and applied uniformly across all heads to ensure attention recall after selection, and (3) KV cache slimming to alleviate head-level redundancy, which selectively fetches visual KV cache according to the head-level decoding query pattern. Experimental results demonstrate that OmniSparse can achieve comparable performance with full attention, achieving 2.7x speedup during prefill and 2.4x memory reduction for decoding.
Feng Chen 0047, Yefei He, Shaoxuan He, Yuanyu He, Jing Liu 0048, Lequan Lin, Akide Liu, Zhenbang Sun, Bohan Zhuang, Qi Wu 0001
AAAI7
2025 FPSAttention: Training-Aware FP8 and Sparsity Co-Design for Fast Video Diffusion
abstract
Diffusion generative models have become the standard for producing high-quality, coherent video content, yet their slow inference speeds and high computational demands hinder practical deployment. Although both quantization and sparsity can independently accelerate inference while maintaining generation quality, naively combining these techniques in existing training-free approaches leads to significant performance degradation, as they fail to achieve proper joint optimization. We introduce FPSAttention, a novel training-aware co-design of FP8 quantization and Sparsity for video generation, with a focus on the 3D bi-directional attention mechanism. Our approach features three key innovations: 1) A unified 3D tile-wise granularity that simultaneously supports both quantization and sparsity. 2) A denoising step-aware strategy that adapts to the noise schedule, addressing the strong correlation between quantization/sparsity errors and denoising steps. 3) A native, hardware-friendly kernel that leverages FlashAttention and is implemented with optimized Hopper architecture features, enabling highly efficient execution. Trained on Wan2.1's 1.3B and 14B models and evaluated on the vBench benchmark, FPSAttention achieves a 7.09$\times$ kernel speedup for attention operations and a 4.96$\times$ end-to-end speedup for video generation compared to the BF16 baseline at 720p resolution—without sacrificing generation quality.
Akide Liu, Zeyu Zhang 0006, Zhexin Li, Xuehai Bai, Yuanjie Xing, Yizeng Han, Jiasheng Tang, Jichao Wu, Mingyang Yang, Yuanyu He, Fan Wang 0019, Gholamreza Haffari, Bohan Zhuang
NeurIPS1
2025 ZPressor: Bottleneck-Aware Compression for Scalable Feed-Forward 3DGS
abstract
Feed-forward 3D Gaussian Splatting (3DGS) models have recently emerged as a promising solution for novel view synthesis, enabling one-pass inference without the need for per-scene 3DGS optimization. However, their scalability is fundamentally constrained by the limited capacity of their encoders, leading to degraded performance or excessive memory consumption as the number of input views increases. In this work, we analyze feed-forward 3DGS frameworks through the lens of the Information Bottleneck principle and introduce ZPressor, a lightweight architecture-agnostic module that enables efficient compression of multi-view inputs into a compact latent state $Z$ that retains essential scene information while discarding redundancy. Concretely, ZPressor enables existing feed-forward 3DGS models to scale to over 100 input views at 480P resolution on an 80GB GPU, by partitioning the views into anchor and support sets and using cross attention to compress the information from the support views into anchor views, forming the compressed latent state $Z$. We show that integrating ZPressor into several state-of-the-art feed-forward 3DGS models consistently improves performance under moderate input views and enhances robustness under dense view settings on two large-scale benchmarks DL3DV-10K and RealEstate10K.
Weijie Wang 0014, Donny Y. Chen, Zeyu Zhang 0006, Duochao Shi, Akide Liu, Bohan Zhuang
NeurIPS5
2025 CIT: Rethinking class-incremental semantic segmentation with a Class Independent Transformation
abstract
Class-incremental semantic segmentation (CSS) requires that a model learn to segment new classes without forgetting how to segment previous ones: this is typically achieved by distilling the current knowledge and incorporating the latest data. However, bypassing iterative distillation by directly transferring outputs of initial classes to the current learning task is not supported in existing class-specific CSS methods. Via Softmax, they enforce dependency between classes and adjust the output distribution at each learning step, resulting in a large probability distribution gap between initial and current tasks. We introduce a simple, yet effective Class Independent Transformation (CIT) that converts the outputs of existing semantic segmentation models into class-independent forms with negligible cost or performance loss. By utilizing class-independent predictions facilitated by CIT, we establish an accumulative distillation framework, ensuring equitable incorporation of all class information. We conduct extensive experiments on various segmentation architectures, including DeepLabV3, Mask2Former, and SegViTv2. Results from these experiments show minimal task forgetting across different datasets, with less than 5% for ADE20K in the most challenging 11 task configurations and less than 1% across all configurations for the PASCAL VOC 2012 dataset. • Softmax interdependency causes incremental forgetting in continual learning. • We introduce a class-independent transformation (CIT) to reduce forgetting. • CIT reformulates segmentation as class-agnostic, enhancing CSS training pipelines. • Our method significantly reduces forgetting on ADE20K compared to CSS baselines. • CIT achieves near-zero forgetting ( ≤ 1%) in Pascal-VOC 2012 settings.
Jinchao Ge, Bowen Zhang 0009, Akide Liu, Vu Minh Hieu Phan, Qi Chen 0014, Yangyang Shu, Yang Zhao 0019
Pattern Recognit.3
2024 Motion Mamba: Efficient and Long Sequence Motion Generation
Zeyu Zhang 0006, Akide Liu, Ian D. Reid 0001, Richard I. Hartley, Bohan Zhuang, Hao Tang 0005
ECCV (1)2
2024 MiniCache: KV Cache Compression in Depth Dimension for Large Language Models
abstract
A critical approach for efficiently deploying computationally demanding large language models (LLMs) is Key-Value (KV) caching. The KV cache stores key-value states of previously generated tokens, significantly reducing the need for repetitive computations and thereby lowering latency in autoregressive generation. However, the size of the KV cache grows linearly with sequence length, posing challenges for applications requiring long context input and extensive sequence generation. In this paper, we present a simple yet effective approach, called MiniCache, to compress the KV cache across layers from a novel depth perspective, significantly reducing the memory footprint for LLM inference. Our approach is based on the observation that KV cache states exhibit high similarity between the adjacent layers in the middle-to-deep portion of LLMs. To facilitate merging, we propose disentangling the states into the magnitude and direction components, interpolating the directions of the state vectors while preserving their lengths unchanged. Furthermore, we introduce a token retention strategy to keep highly distinct state pairs unmerged, thus preserving the information with minimal additional storage overhead. Our MiniCache is training-free and general, complementing existing KV cache compression strategies, such as quantization and sparsity. We conduct a comprehensive evaluation of MiniCache utilizing various models including LLaMA-2, LLaMA-3, Phi-3, Mistral, and Mixtral across multiple benchmarks, demonstrating its exceptional performance in achieving superior compression ratios and high throughput. On the ShareGPT dataset, LLaMA-2-7B with cross-layer merging achieves a compression ratio of $1.53\times$. Additionally, since MiniCache is orthogonal to existing quantization techniques, it can achieve a compression ratio of up to $5.02\times$ when combined with the 4-bit quantization technique, enhancing inference throughput by approximately $5\times$ and reducing the memory footprint by $41\%$ compared to the FP16 full cache baseline, all while maintaining near-lossless performance. Project is available at https://minicache.vmv.re .
Akide Liu, Jing Liu 0048, Zizheng Pan, Yefei He, Gholamreza Haffari, Bohan Zhuang
NeurIPS1