EDBT 2026 Demo / reviewers in the wild / expert
Jianke Zhang
dblp:46/6970
· DBLP profile ↗
13ranked-venue papers
3as first author
9since 2021 · last 2026
0000-0003-0424-5067ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 2 first-author · 8 since 2021Systems, architecture and hardware · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Motion planning and robot control · 26% Robot manipulation · 19% Reinforcement learning · 16% | |
| Software engineering, system software, and programming languages
1 paper |
Program synthesis and code generation · 100% |
Topics — the 17 heaviest of 18, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Robotics › Motion planning and robot control
robot learning |
1.7 | 2 | 2025 | Improving Vision-Language-Action Model with Online Reinforcement Learning · ICRA 2025 Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations · ICML 2025 |
Robotics › Robot manipulation › embodied foundation models
vision-language-action model |
1.7 | 2 | 2025 | Improving Vision-Language-Action Model with Online Reinforcement Learning · ICRA 2025 UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent · ICML 2025 |
Machine learning › Generative modeling
diffusion model |
1.0 | 2 | 2025 | Prediction with Action: Visual Policy Learning via Joint Denoising Process · NeurIPS 2024 Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations · ICML 2025 |
Computer vision › Vision and language › vision-language model › multimodal large language model
chart understanding |
1.0 | 1 | 2026 | RealChart2Code: Bridging the Gap in Real-World Chart-to-Code Generation via Multi-Task Evaluation · ACL (1) 2026 |
Computer vision › Vision and language
multimodal evaluation |
1.0 | 1 | 2026 | RealChart2Code: Bridging the Gap in Real-World Chart-to-Code Generation via Multi-Task Evaluation · ACL (1) 2026 |
Program synthesis and code generation › code generation with language models
chart-to-code generation |
1.0 | 1 | 2026 | RealChart2Code: Bridging the Gap in Real-World Chart-to-Code Generation via Multi-Task Evaluation · ACL (1) 2026 |
Machine learning › Reinforcement learning
embodied control |
0.9 | 1 | 2025 | UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent · ICML 2025 |
Robotics › Motion planning and robot control › robot learning › robot policy learning
generalist robot policy |
0.9 | 1 | 2025 | Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations · ICML 2025 |
Robotics › Motion planning and robot control › robot dynamics
inverse dynamics |
0.9 | 1 | 2025 | Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations · ICML 2025 |
Computer vision › 3D vision
spatial understanding |
0.9 | 1 | 2025 | UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent · ICML 2025 |
Machine learning › Representation and self-supervised learning › representation learning
visual representation learning |
0.9 | 1 | 2025 | Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations · ICML 2025 |
Robotics › Robot manipulation
diffusion policy |
0.8 | 1 | 2024 | Prediction with Action: Visual Policy Learning via Joint Denoising Process · NeurIPS 2024 |
Machine learning › Reinforcement learning › deep reinforcement learning
visual policy learning |
0.8 | 1 | 2024 | Prediction with Action: Visual Policy Learning via Joint Denoising Process · NeurIPS 2024 |
Machine learning › Representation and self-supervised learning › pre-training
multimodal pretraining |
0.3 | 1 | 2025 | UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent · ICML 2025 |
Machine learning › Reinforcement learning › online decision making
online reinforcement learning |
0.3 | 1 | 2025 | Improving Vision-Language-Action Model with Online Reinforcement Learning · ICRA 2025 |
Machine learning › Generative modeling › diffusion model
video diffusion model |
0.3 | 1 | 2025 | Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations · ICML 2025 |
Machine learning › Reinforcement learning
imitation learning |
0.2 | 1 | 2024 | Prediction with Action: Visual Policy Learning via Joint Denoising Process · NeurIPS 2024 |
Methods — techniques the papers use, named apart from their topics
multi-task evaluation · 2.0vision-language model pretraining · 0.9vision-language model · 0.9video diffusion model · 0.9supervised fine-tuning · 0.9reinforcement learning · 0.9inverse dynamics learning · 0.9future prediction objective · 0.9fine-tuning on robot data · 0.9diffusion transformer · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RealChart2Code: Bridging the Gap in Real-World Chart-to-Code Generation via Multi-Task EvaluationabstractJiajun Zhang, Yuying Li, Zhixun Li, Xingyu Guo, Jingzhuo Wu, Leqi Zheng, Yiran Yang, Jianke Zhang, Qingbin Li, Shannan Yan, Changguo Jia, Junfei Wu, Zilei Wang, Qiang Liu, Liang Wang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Zhixun Li, Xingyu Guo, Jingzhuo Wu, Leqi Zheng, Jianke Zhang, Qingbin Li, Shannan Yan, Changguo Jia, Junfei Wu, Zilei Wang |
ACL (1) | 8 |
| 2025 | Video Prediction Policy: A Generalist Robot Policy with Predictive Visual RepresentationsabstractVisual representations play a crucial role in developing generalist robotic policies. Previous vision encoders, typically pre-trained with single-image reconstruction or two-image contrastive learning, tend to capture static information, often neglecting the dynamic aspects vital for embodied tasks. Recently, video diffusion models (VDMs) demonstrate the ability to predict future frames and showcase a strong understanding of physical world. We hypothesize that VDMs inherently produce visual representations that encompass both current static information and predicted future dynamics, thereby providing valuable guidance for robot action learning. Based on this hypothesis, we propose the Video Prediction Policy (VPP), which learns implicit inverse dynamics model conditioned on predicted future representations inside VDMs. To predict more precise future, we fine-tune pre-trained video foundation model on robot datasets along with internet human manipulation data. In experiments, VPP achieves a 18.6% relative improvement on the Calvin ABC-D generalization benchmark compared to the previous state-of-the-art, and demonstrates a 31.6% increase in success rates for complex real-world dexterous manipulation tasks. For your convenience, videos can be found at https://video-prediction-policy.github.io/ Yanjiang Guo, Pengchao Wang, Yen-Jen Wang, Jianke Zhang, Koushil Sreenath, Chaochao Lu, Jianyu Chen 0002 |
ICML | 6 |
| 2025 | UP-VLA: A Unified Understanding and Prediction Model for Embodied AgentabstractRecent advancements in Vision-Language-Action (VLA) models have leveraged pre-trained Vision-Language Models (VLMs) to improve the generalization capabilities. VLMs, typically pre-trained on vision-language understanding tasks, provide rich semantic knowledge and reasoning abilities. However, prior research has shown that VLMs often focus on high-level semantic content and neglect low-level features, limiting their ability to capture detailed spatial information and understand physical dynamics. These aspects, which are crucial for embodied control tasks, remain underexplored in existing pre-training paradigms. In this paper, we investigate the training paradigm for VLAs, and introduce UP-VLA, a Unified VLA model training with both multi-modal Understanding and future Prediction objectives, enhancing both high-level semantic comprehension and low-level spatial understanding. Experimental results show that UP-VLA achieves a 33% improvement on the Calvin ABC-D benchmark compared to the previous state-of-the-art method. Additionally, UP-VLA demonstrates improved success rates in real-world manipulation tasks, particularly those requiring precise spatial information. Jianke Zhang, Yanjiang Guo, Jianyu Chen 0002 |
ICML | 1 |
| 2025 | Improving Vision-Language-Action Model with Online Reinforcement LearningabstractRecent studies have successfully integrated large vision-language models (VLMs) into low-level robotic control by supervised fine-tuning (SFT) with expert robotic datasets, resulting in what we term vision-language-action (VLA) models. Although the VLA models are powerful, how to improve these large models during interaction with environments remains an open question. In this paper, we explore how to further improve these VLA models via Reinforcement Learning (RL), a commonly used fine-tuning technique for large models. However, we find that directly applying online RL to large VLA models presents significant challenges, including training instability that severely impacts the performance of large models, and computing burdens that exceed the capabilities of most local machines. To address these challenges, we propose iRe-VLA framework, which iterates between Reinforcement Learning and Supervised Learning to effectively improve VLA models, leveraging the exploratory benefits of RL while maintaining the stability of supervised learning. Experiments in two simulated benchmarks and a real-world manipulation suite validate the effectiveness of our method. Yanjiang Guo, Jianke Zhang, Yen-Jen Wang, Jianyu Chen 0002 |
ICRA | 2 |
| 2025 | Cross-modal adapter for vision-language retrieval
Haojun Jiang, Jianke Zhang, Rui Huang 0012, Chunjiang Ge, Zanlin Ni, Shiji Song, Gao Huang 0001 |
Pattern Recognit. | 2 |
| 2024 | Prediction with Action: Visual Policy Learning via Joint Denoising ProcessabstractDiffusion models have demonstrated remarkable capabilities in image generation tasks, including image editing and video creation, representing a good understanding of the physical world. On the other line, diffusion models have also shown promise in robotic control tasks by denoising actions, known as diffusion policy. Although the diffusion generative model and diffusion policy exhibit distinct capabilities—image prediction and robotic action, respectively—they technically follow similar denoising process. In robotic tasks, the ability to predict future images and generate actions is highly correlated since they share the same underlying dynamics of the physical world. Building on this insight, we introduce \textbf{PAD}, a novel visual policy learning framework that unifies image \textbf{P}rediction and robot \textbf{A}ction within a joint \textbf{D}enoising process. Specifically, PAD utilizes Diffusion Transformers (DiT) to seamlessly integrate images and robot states, enabling the simultaneous prediction of future images and robot actions. Additionally, PAD supports co-training on both robotic demonstrations and large-scale video datasets and can be easily extended to other robotic modalities, such as depth images.
PAD outperforms previous methods, achieving a significant 38.9\% relative improvement on the full Metaworld benchmark, by utilizing a single text-conditioned visual policy within a data-efficient imitation learning setting. Furthermore, PAD demonstrates superior generalization to unseen tasks in real-world robot manipulation settings with 28.0\% success rate increase compared to the strongest baseline.
Videos of PAD can be found at https://sites.google.com/view/pad-paper Yanjiang Guo, Jianke Zhang, Yen-Jen Wang, Chaochao Lu, Jianyu Chen 0002 |
NeurIPS | 3 |
| 2024 | Sufficient conditions for interval-valued optimal control problems in admissible orders
Lifeng Li, Jianke Zhang |
Soft Comput. | 2 |
| 2022 | Optimality conditions for fuzzy optimization problems under granular convexity conceptabstractIn this paper, under the condition of granular differentiation, we consider the fuzzy optimization problems with the general fuzzy function as the objective function. Firstly, we introduce the concept of granular convexity, and propose the properties of the granular convex fuzzy functions. Secondly, we present the Karush-Kuhn-Tucker type optimality conditions of the fuzzy relative optimal solution of more general fuzzy programming problems and some test examples. Finally, the relationships between a class of variational inequalities and the fuzzy optimization problems are established. Jianke Zhang, Lifeng Li, Xiaojue Ma |
Fuzzy Sets Syst. | 1 |
| 2022 | Two classes of granular solutions and related optimality conditions for interval type-2 fuzzy optimization
Jianke Zhang, Zeshui Xu, Feng Feng 0003, Ronald R. Yager |
Inf. Sci. | 1 |
| 2016 | Virtual machine consolidated placement based on multi-objective biogeography-based optimization
Rui Li 0073, Xiuqi Li, Nazaraf Shah, Jianke Zhang, Feng Tian 0002, Kuo-Ming Chao |
Future Gener. Comput. Syst. | 5 |
| 2015 | On fuzzy generalized convex mappings and optimality conditions for fuzzy weakly univex mappings
Lifeng Li, Jianke Zhang |
Fuzzy Sets Syst. | 3 |
| 2014 | Biogeography-based optimization with improved migration operator and self-adaptive clear duplicate operator
Quanxi Feng, Jianke Zhang, Longquan Yong |
Appl. Intell. | 3 |
| 2010 | Attribute reduction in fuzzy concept lattices based on the T implication
Lifeng Li, Jianke Zhang |
Knowl. Based Syst. | 2 |