VLDB 2026 Research / reviewers in the wild / expert
Shaoxuan He
dblp:384/6643
· DBLP profile ↗
5ranked-venue papers
0as first author
5since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Efficient and distributed learning · 31% Generative modeling · 27% Language models and text generation · 9% | |
| Computer graphics and multimedia
1 paper |
Visual content generation and editing · 100% | |
| Databases, data mining, and information retrieval
1 paper |
Recommender systems · 100% |
Topics — the 16 heaviest of 19, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Efficient and distributed learning
inference acceleration |
1.9 | 2 | 2026 | OmniSparse: Training-Aware Fine-Grained Sparse Attention for Long-Video MLLMs · AAAI 2026 ZipAR: Parallel Autoregressive Image Generation through Spatial Locality · ICML 2025 |
Machine learning › Generative modeling
autoregressive model |
1.7 | 2 | 2025 | ZipAR: Parallel Autoregressive Image Generation through Spatial Locality · ICML 2025 Neighboring Autoregressive Modeling for Efficient Visual Generation · ICCV 2025 |
Natural language and speech › Language models and text generation › decoding › decoding strategy
parallel decoding |
1.1 | 2 | 2025 | ZipAR: Parallel Autoregressive Image Generation through Spatial Locality · ICML 2025 Neighboring Autoregressive Modeling for Efficient Visual Generation · ICCV 2025 |
Machine learning › Efficient and distributed learning › KV cache management
KV cache compression |
1.0 | 1 | 2026 | OmniSparse: Training-Aware Fine-Grained Sparse Attention for Long-Video MLLMs · AAAI 2026 |
Machine learning › Deep learning architectures and training › attention mechanism
sparse attention |
1.0 | 1 | 2026 | OmniSparse: Training-Aware Fine-Grained Sparse Attention for Long-Video MLLMs · AAAI 2026 |
Machine learning › Generative modeling › autoregressive model
autoregressive image generation |
0.9 | 1 | 2025 | Neighboring Autoregressive Modeling for Efficient Visual Generation · ICCV 2025 |
Machine learning › Generative modeling
image generation |
0.9 | 1 | 2025 | ZipAR: Parallel Autoregressive Image Generation through Spatial Locality · ICML 2025 |
Knowledge, reasoning and agents › Knowledge representation and reasoning
spatial reasoning |
0.9 | 1 | 2025 | SpatialCLIP: Learning 3D-aware Image Representations from Spatially Discriminative Language · CVPR 2025 |
Computer vision › 3D vision
spatial understanding |
0.9 | 1 | 2025 | SpatialCLIP: Learning 3D-aware Image Representations from Spatially Discriminative Language · CVPR 2025 |
Machine learning › Efficient and distributed learning › inference acceleration
speculative decoding |
0.9 | 1 | 2025 | ZipAR: Parallel Autoregressive Image Generation through Spatial Locality · ICML 2025 |
Computer vision › Vision and language
vision-language model |
0.9 | 1 | 2025 | SpatialCLIP: Learning 3D-aware Image Representations from Spatially Discriminative Language · CVPR 2025 |
Visual content generation and editing › image generation
autoregressive image generation |
0.9 | 1 | 2025 | Neighboring Autoregressive Modeling for Efficient Visual Generation · ICCV 2025 |
Visual content generation and editing
image and video generation |
0.9 | 1 | 2025 | Neighboring Autoregressive Modeling for Efficient Visual Generation · ICCV 2025 |
Recommender systems
sequential recommendation |
0.8 | 1 | 2024 | Semantic Codebook Learning for Dynamic Recommendation Models · ACM Multimedia 2024 |
Machine learning › Efficient and distributed learning
inference efficiency |
0.3 | 1 | 2025 | Neighboring Autoregressive Modeling for Efficient Visual Generation · ICCV 2025 |
Machine learning › Representation and self-supervised learning › representation learning
disentangled representation learning |
0.2 | 1 | 2024 | Semantic Codebook Learning for Dynamic Recommendation Models · ACM Multimedia 2024 |
Methods — techniques the papers use, named apart from their topics
outpainting · 1.7autoregressive modeling · 1.7query selection · 1.0dynamic token budget allocation · 1.0KV selection · 1.0KV cache slimming · 1.0rejection sampling · 0.9hard negative captions · 0.9contrastive language-image pretraining · 0.93d-inspired vit · 0.9semantic metacode · 0.8semantic codebook · 0.8dual parameter model · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | OmniSparse: Training-Aware Fine-Grained Sparse Attention for Long-Video MLLMsabstractExisting sparse attention methods primarily target inference-time acceleration by selecting critical tokens under predefined sparsity patterns. However, they often fail to bridge the training–inference gap and lack the capacity for fine-grained token selection across multiple dimensions—such as queries, key-values (KV), and heads—leading to suboptimal performance and acceleration gains. In this paper, we introduce OmniSparse, a training-aware fine-grained sparse attention of long-video MLLMs, which is applied in both training and inference with dynamic token budget allocation. Specifically, OmniSparse contains three adaptive and complementary mechanisms: (1) query selection as lazy-active classification, aiming to retain active queries that capture broader semantic similarity, while discarding most of lazy ones that focus on limited local context and exhibit high functional redundancy with their neighbors, (2) KV selection with head-level dynamic budget allocation, where a shared budget is determined based on the flattest head and applied uniformly across all heads to ensure attention recall after selection, and (3) KV cache slimming to alleviate head-level redundancy, which selectively fetches visual KV cache according to the head-level decoding query pattern. Experimental results demonstrate that OmniSparse can achieve comparable performance with full attention, achieving 2.7x speedup during prefill and 2.4x memory reduction for decoding. Feng Chen 0047, Yefei He, Shaoxuan He, Yuanyu He, Jing Liu 0048, Lequan Lin, Akide Liu, Zhenbang Sun, Bohan Zhuang, Qi Wu 0001 |
AAAI | 3 |
| 2025 | SpatialCLIP: Learning 3D-aware Image Representations from Spatially Discriminative LanguageabstractContrastive Language-Image Pre-training (CLIP) learns robust visual models through language supervision, making it a crucial visual encoding technique for various applications. However, CLIP struggles with comprehending spatial concepts in images, potentially restricting the spatial intelligence of CLIP-based AI systems. In this work, we propose SpatialCLIP, an enhanced version of CLIP with better spatial understanding capabilities. To capture the intricate 3D spatial relationships in images, we improve both "visual model" and "language supervision" of CLIP. Specifically, we design 3D-inspired ViT to replace the standard ViT in CLIP. By lifting 2D image tokens into 3D space and incorporating design insights from point cloud networks, our visual model gains greater potential for spatial perception. Meanwhile, captions with accurate and detailed spatial information are very rare. To explore better language supervision for spatial understanding, we re-caption images and perturb their spatial phrases as negative descriptions, which compels the visual model to seek spatial cues to distinguish these hard negative captions. With the enhanced visual model, we introduce SpatialLLaVA, following the same LLaVA-1.5 training protocol, to investigate the importance of visual representations for MLLM’s spatial intelligence. Furthermore, we create SpatialBench, a benchmark specifically designed to evaluate CLIP and MLLM in spatial reasoning. Spatial-CLIP and SpatialLLaVA achieve substantial performance improvements, demonstrating stronger capabilities in spatial perception and reasoning, while maintaining comparable results on general-purpose benchmarks. Zehan Wang 0001, Sashuai Zhou, Shaoxuan He, Haifeng Huang 0001, Lihe Yang, Xize Cheng, Shengpeng Ji, Tao Jin 0004, Hengshuang Zhao, Zhou Zhao 0001 |
CVPR | 3 |
| 2025 | Neighboring Autoregressive Modeling for Efficient Visual GenerationabstractVisual autoregressive models typically adhere to a raster-order ``next-token prediction" paradigm, which overlooks the spatial and temporal locality inherent in visual content. Specifically, visual tokens exhibit significantly stronger correlations with their spatially or temporally adjacent tokens compared to those that are distant. In this paper, we propose Neighboring Autoregressive Modeling (NAR), a novel paradigm that formulates autoregressive visual generation as a progressive outpainting procedure, following a near-to-far ``next-neighbor prediction" mechanism. Starting from an initial token, the remaining tokens are decoded in ascending order of their Manhattan distance from the initial token in the spatial-temporal space, progressively expanding the boundary of the decoded region. To enable parallel prediction of multiple adjacent tokens in the spatial-temporal space, we introduce a set of dimension-oriented decoding heads, each predicting the next token along a mutually orthogonal dimension. During inference, all tokens adjacent to the decoded tokens are processed in parallel, substantially reducing the model forward steps for generation. Experiments on ImageNet$256\times 256$ and UCF101 demonstrate that NAR achieves 2.4$\times$ and 8.6$\times$ higher throughput respectively, while obtaining superior FID/FVD scores for both image and video generation tasks compared to the PAR-4X approach. When evaluating on text-to-image generation benchmark GenEval, NAR with 0.8B parameters outperforms Chameleon-7B while using merely 0.4 of the training data. Code is available at https://github.com/ThisisBillhe/NAR. Yefei He, Yuanyu He, Shaoxuan He, Feng Chen 0047, Kaipeng Zhang, Bohan Zhuang |
ICCV | 3 |
| 2025 | ZipAR: Parallel Autoregressive Image Generation through Spatial LocalityabstractIn this paper, we propose ZipAR, a training-free, plug-and-play parallel decoding framework for accelerating autoregressive (AR) visual generation. The motivation stems from the observation that images exhibit local structures, and spatially distant regions tend to have minimal interdependence. Given a partially decoded set of visual tokens, in addition to the original next-token prediction scheme in the row dimension, the tokens corresponding to spatially adjacent regions in the column dimension can be decoded in parallel. To ensure alignment with the contextual requirements of each token, we employ an adaptive local window assignment scheme with rejection sampling analogous to speculative decoding. By decoding multiple tokens in a single forward pass, the number of forward passes required to generate an image is significantly reduced, resulting in a substantial improvement in generation efficiency. Experiments demonstrate that ZipAR can reduce the number of model forward passes by up to 91% on the Emu3-Gen model without requiring any additional retraining. Yefei He, Feng Chen 0047, Yuanyu He, Shaoxuan He, Kaipeng Zhang, Bohan Zhuang |
ICML | 4 |
| 2024 | Semantic Codebook Learning for Dynamic Recommendation ModelsabstractDynamic sequential recommendation (DSR) can generate model parameters based on user behavior to improve the personalization of sequential recommendation under various user preferences. However, it faces the challenges of large parameter search space and sparse and noisy user-item interactions, which reduces the applicability of the generated model parameters. The Semantic Codebook Learning for Dynamic Recommendation Models (SOLID) framework presents a significant advancement in DSR by effectively tackling these challenges. By transforming item sequences into semantic sequences and employing a dual parameter model, SOLID compresses the parameter generation search space and leverages homogeneity within the recommendation system. The introduction of the semantic metacode and semantic codebook, which stores disentangled item representations, ensures robust and accurate parameter generation. Extensive experiments demonstrates that SOLID consistently outperforms existing DSR, delivering more accurate, stable, and robust recommendations. Zheqi Lv, Shaoxuan He, Tianyu Zhan, Shengyu Zhang 0001, Wenqiao Zhang, Jingyuan Chen 0003, Zhou Zhao 0001, Fei Wu 0001 |
ACM Multimedia | 2 |