VLDB 2026 Research / reviewers in the wild / expert
Renshan Zhang
dblp:251/4783
· DBLP profile ↗
3ranked-venue papers
1as first author
3since 2021 · last 2026
0000-0003-1833-9996ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Efficient and distributed learning · 36% Robot manipulation · 30% Representation and self-supervised learning · 16% |
Topics — the 8 heaviest of 8, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Robotics › Robot manipulation › embodied foundation models
vision-language-action model |
1.9 | 2 | 2026 | SemanticVLA: Semantic-Aligned Sparsification and Enhancement for Efficient Robotic Manipulation · AAAI 2026 CogVLA: Cognition-Aligned Vision-Language-Action Models via Instruction-Driven Routing & Sparsification · NeurIPS 2025 |
Machine learning › Representation and self-supervised learning
semantic alignment |
1.0 | 1 | 2026 | SemanticVLA: Semantic-Aligned Sparsification and Enhancement for Efficient Robotic Manipulation · AAAI 2026 |
Machine learning › Efficient and distributed learning › model compression › sparsity
model sparsification |
0.9 | 1 | 2025 | CogVLA: Cognition-Aligned Vision-Language-Action Models via Instruction-Driven Routing & Sparsification · NeurIPS 2025 |
Computer vision › Vision and language › vision-language model
multimodal large language model |
0.9 | 1 | 2025 | FALCON: Resolving Visual Redundancy and Fragmentation in High-Resolution Multimodal Large Language Models via Visual Registers · ICCV 2025 |
Machine learning › Efficient and distributed learning › model compression › token compression
visual token reduction |
0.9 | 1 | 2025 | FALCON: Resolving Visual Redundancy and Fragmentation in High-Resolution Multimodal Large Language Models via Visual Registers · ICCV 2025 |
Machine learning › Efficient and distributed learning › token reduction
visual token sparsification |
0.3 | 1 | 2026 | SemanticVLA: Semantic-Aligned Sparsification and Enhancement for Efficient Robotic Manipulation · AAAI 2026 |
Robotics › Motion planning and robot control
robot control |
0.3 | 1 | 2025 | CogVLA: Cognition-Aligned Vision-Language-Action Models via Instruction-Driven Routing & Sparsification · NeurIPS 2025 |
Machine learning › Efficient and distributed learning › model compression
token compression |
0.3 | 1 | 2025 | FALCON: Resolving Visual Redundancy and Fragmentation in High-Resolution Multimodal Large Language Models via Visual Registers · ICCV 2025 |
Methods — techniques the papers use, named apart from their topics
visual token pruning · 1.0mixture of experts · 1.0feature fusion · 1.0visual register · 0.9token pruning · 0.9register-based representation compacting · 0.9register interactive attention · 0.9instruction-driven routing · 0.9coupled attention · 0.9FiLM conditioning · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SemanticVLA: Semantic-Aligned Sparsification and Enhancement for Efficient Robotic ManipulationabstractVision-Language-Action (VLA) models have advanced in robotic manipulation, yet practical deployment remains hindered by two key limitations: **1) perceptual redundancy**, where irrelevant visual inputs are processed inefficiently, and **2) superficial instruction-vision alignment**, which hampers semantic grounding of actions. In this paper, we propose **SemanticVLA**, a novel VLA framework that performs Semantic-Aligned Sparsification and Enhancement for Efficient Robotic Manipulation. Specifically: **1)** To sparsify redundant perception while preserving semantic alignment, **Semantic-guided Dual Visual Pruner (SD-Pruner)** performs: Instruction-driven Pruner (ID-Pruner) extracts global action cues and local semantic anchors in SigLIP; Spatial-aggregation Pruner (SA-Pruner) compacts geometry-rich features into task-adaptive tokens in DINOv2. **2)** To exploit sparsified features and integrate semantics with spatial geometry, **Semantic-complementary Hierarchical Fuser (SH-Fuser)** fuses dense patches and sparse tokens across SigLIP and DINOv2 for coherent representation. **3)** To enhance the transformation from perception to action, **Semantic-conditioned Action Coupler (SA-Coupler)** replaces the conventional observation-to-DoF approach, yielding more efficient and interpretable behavior modeling for manipulation tasks. Extensive experiments on simulation and real-world tasks show that SemanticVLA sets a new SOTA in both performance and efficiency. SemanticVLA surpasses OpenVLA on LIBERO benchmark by **21.1%** in success rate, while reducing training cost and inference latency by **3.0×** and **2.7×**. Renshan Zhang, Rui Shao 0001, Zhijian Fang, Kaiwen Zhou 0001, Zhuotao Tian, Liqiang Nie |
AAAI | 2 |
| 2025 | FALCON: Resolving Visual Redundancy and Fragmentation in High-Resolution Multimodal Large Language Models via Visual RegistersabstractThe incorporation of high-resolution visual input equips multimodal large language models (MLLMs) with enhanced visual perception capabilities for real-world tasks. However, most existing high-resolution MLLMs rely on a cropping-based approach to process images, which leads to fragmented visual encoding and a sharp increase in redundant tokens. To tackle these issues, we propose the FALCON model. FALCON introduces a novel visual register technique to simultaneously: 1) Eliminate redundant tokens at the stage of visual encoding. To directly address the visual redundancy present in the output of vision encoder, we propose a Register-based Representation Compacting (ReCompact) mechanism. This mechanism introduces a set of learnable visual registers designed to adaptively aggregate essential information while discarding redundancy. It enables the encoder to produce a more compact visual representation with a minimal number of output tokens, thus eliminating the need for an additional compression module. 2) Ensure continuity in visual encoding. To address the potential encoding errors caused by fragmented visual inputs, we develop a Register Interactive Attention (ReAtten) module. This module facilitates effective and efficient information exchange across sub-images by enabling interactions between visual registers. It ensures the continuity of visual semantics throughout the encoding. We conduct comprehensive experiments with FALCON on high-resolution benchmarks across a wide range of scenarios. FALCON demonstrates superior performance with a remarkable 9-fold reduction in visual tokens. Renshan Zhang, Rui Shao 0001, Gongwei Chen, Miao Zhang 0022, Kaiwen Zhou 0001, Weili Guan, Liqiang Nie |
ICCV | 1 |
| 2025 | CogVLA: Cognition-Aligned Vision-Language-Action Models via Instruction-Driven Routing & SparsificationabstractRecent Vision-Language-Action (VLA) models built on pre-trained Vision-Language Models (VLMs) require extensive post-training, resulting in high computational overhead that limits scalability and deployment. Existing sparsification strategies—such as Mixture-of-Depths, layer skipping, and early exit—fall short by neglecting the semantic coupling across vision-language-action modalities, and focusing narrowly on intra-LLM computation while overlooking end-to-end coherence from perception to control. To address these challenges, we propose **CogVLA**, a Cognition-Aligned Vision-Language-Action framework that leverages instruction-driven routing and sparsification to improve both efficiency and performance. CogVLA draws inspiration from human multimodal coordination and introduces a 3-stage progressive architecture. 1) **Encoder-FiLM based Aggregation Routing (EFA-Routing)** injects instruction information into the vision encoder to selectively aggregate and compress dual-stream visual tokens, forming a instruction-aware latent representation. 2) Building upon this compact visual encoding, **LLM-FiLM based Pruning Routing (LFP-Routing)** introduces action intent into the language model by pruning instruction-irrelevant visually grounded tokens, thereby achieving token-level sparsity. 3) To ensure that compressed perception inputs can still support accurate and coherent action generation, we introduce **V‑L‑A Coupled Attention (CAtten)**, which combines causal vision-language attention with bidirectional action parallel decoding.
Extensive experiments on the LIBERO benchmark and real-world robotic tasks demonstrate that CogVLA achieves state-of-the-art performance with success rates of 97.4\% and 70.0\%, respectively, while reducing training costs by 2.5$\times$ and decreasing inference latency by 2.8$\times$ compared to OpenVLA. Renshan Zhang, Rui Shao 0001, Liqiang Nie |
NeurIPS | 2 |