VLDB 2026 Research / reviewers in the wild / expert
Sizhe Yang
dblp:351/1712
· DBLP profile ↗
6ranked-venue papers
3as first author
6since 2021 · last 2025
0009-0007-4462-5191ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Reinforcement learning · 40% Motion planning and robot control · 26% Robot manipulation · 13% |
Topics — the 8 heaviest of 8, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning › deep reinforcement learning
visual reinforcement learning |
1.3 | 2 | 2023 | RL-ViGen: A Reinforcement Learning Benchmark for Visual Generalization · NeurIPS 2023 MoVie: Visual Model-Based Policy Adaptation for View Generalization · NeurIPS 2023 |
Robotics › Motion planning and robot control › robot dynamics
inverse dynamics |
0.9 | 1 | 2025 | Predictive Inverse Dynamics Models are Scalable Learners for Robotic Manipulation · ICLR 2025 |
Robotics › Motion planning and robot control
robot learning |
0.9 | 1 | 2025 | Predictive Inverse Dynamics Models are Scalable Learners for Robotic Manipulation · ICLR 2025 |
Robotics › Robot manipulation › embodied foundation models
vision-language-action model |
0.9 | 1 | 2025 | Predictive Inverse Dynamics Models are Scalable Learners for Robotic Manipulation · ICLR 2025 |
Machine learning › Learning theory
generalization |
0.7 | 1 | 2023 | RL-ViGen: A Reinforcement Learning Benchmark for Visual Generalization · NeurIPS 2023 |
Machine learning › Reinforcement learning
model-based reinforcement learning |
0.7 | 1 | 2023 | MoVie: Visual Model-Based Policy Adaptation for View Generalization · NeurIPS 2023 |
Machine learning › Reinforcement learning
policy adaptation |
0.7 | 1 | 2023 | MoVie: Visual Model-Based Policy Adaptation for View Generalization · NeurIPS 2023 |
Computer vision › 3D vision
view generalization |
0.7 | 1 | 2023 | MoVie: Visual Model-Based Policy Adaptation for View Generalization · NeurIPS 2023 |
Methods — techniques the papers use, named apart from their topics
world model · 0.9transformer · 0.9behavior cloning · 0.9test-time adaptation · 0.7benchmark · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | FCoDT-Net: A Novel Framework for High-Precision Medical Image Segmentation Using Contextual Distillation TransformerabstractCurrent methods in medical image semantic segmentation often rely on simple skip connections within U-shaped network structures. These approaches fail to bridge the semantic gap between the encoder and decoder module and do not fully exploit the rich contextual information among. The unused information leads to suboptimal segmentation results. In this paper, we propose the Feature Context Distillation Transformer Network (FCoDT-Net), a deep learning model designed to address these limitations by leveraging the rich contextual information within the skip connections. FCoDT-Net introduces a novel Context Distillation Transformer (CoDT) within its decoder module, which effectively utilizes abundant contextual information from skip connections. CoDT leverages attention matrix calculations between feature maps at similar scales to exploit rich context information. The narrowed semantic gap then can generate features which can be more efficiently utilized by the decoder. Further, we integrate Convolutional Network Next(ConvNeXt) block into our encoder module to enhance feature extraction capabilities. Experimental simulations show that our FCoDT-Net excels in both 3D and 2D applications. In 3D, FCoDT-Net achieves exceptional results on the lung dataset, with a Dice score of 0.9250 for airway segmentation. In 2D, FCoDT-Net demonstrates state-of-the-art performance on both the Gland Segmentation (Glas) and Multi-Organ Nuclei Segmentation and Classification (MoNuSeg) datasets. Yutao Qin, Sizhe Yang, Bang Hu |
ICASSP | 2 |
| 2025 | Predictive Inverse Dynamics Models are Scalable Learners for Robotic ManipulationabstractCurrent efforts to learn scalable policies in robotic manipulation primarily fall into two categories: one focuses on "action," which involves behavior cloning from extensive collections of robotic data, while the other emphasizes "vision," enhancing model generalization by pre-training representations or generative models, also referred to as world models, using large-scale visual datasets. This paper presents an end-to-end paradigm that predicts actions using inverse dynamics models conditioned on the robot's forecasted visual states, named Predictive Inverse Dynamics Models (PIDM). By closing the loop between vision and action, the end-to-end PIDM can be a better scalable action learner. In practice, we use Transformers to process both visual states and actions, naming the model Seer. It is initially pre-trained on large-scale robotic datasets, such as DROID, and can be adapted to real-world scenarios with a little fine-tuning data. Thanks to large-scale, end-to-end training and the continuous synergy between vision and action at each execution step, Seer significantly outperforms state-of-the-art methods across both simulation and real-world experiments. It achieves improvements of 13% on the LIBERO-LONG benchmark, 22% on CALVIN ABC-D, and 43% in real-world tasks. Notably, it demonstrates superior generalization for novel objects, lighting conditions, and environments under high-intensity disturbances. Code and models will be publicly available. Sizhe Yang, Dahua Lin, Hao Dong 0003, Jiangmiao Pang |
ICLR | 2 |
| 2025 | LWMUNet: Enhancing Semantic Segmentation with Learnable Linked Sliding-Window Serialization
Sizhe Yang, Yutao Qin |
KSEM (2) | 1 |
| 2025 | BMUNet: When Pixel-Wise Precision Meets Global Context DependencyabstractCapturing both local details and long-range dependencies is essential for accurate and comprehensive results in image segmentation. Local features, such as edges and textures, are crucial for distinguishing fine-grained structures, while global dependencies provide the necessary context for understanding overall image information. Traditional CNN-based models struggle to model global information effectively, while Transformer-based models face challenges in capturing pixel-wise information due to their quadratic complexity. To address these limitations, we introduce Bifocal Mamba UNet (BMUNet), a novel segmentation network that effectively integrates both local and global feature extraction. The core component of BMUNet is Bifocal Mamba, a novel Mamba-based framework featuring a dual-branch structure: the Overlapping Sliding-Window mechanism based branch, Sliding-window Pixel VMamba (SwPiVM), enhances pixel-level feature extraction while preserving spatial locality. Activation-and-Pooling Global Tree Mamba (ApGlTM) branch, with activation and pool sampling, is designed for global context understanding. Experimental results show BMUNet outperforms all baselines on GlaS, MoNuSeg, and ISIC-2018 datasets in metrics of Dice score and HD95. Also, BMUNet achieves highest IoU scores (68.83, and 84.63) on MoNuSeg and ISIC-2018 datasets, except for one method on GlaS. Sizhe Yang, Yutao Qin |
ICMR | 1 |
| 2023 | MoVie: Visual Model-Based Policy Adaptation for View GeneralizationabstractVisual Reinforcement Learning (RL) agents trained on limited views face significant challenges in generalizing their learned abilities to unseen views. This inherent difficulty is known as the problem of $\textit{view generalization}$. In this work, we systematically categorize this fundamental problem into four distinct and highly challenging scenarios that closely resemble real-world situations. Subsequently, we propose a straightforward yet effective approach to enable successful adaptation of visual $\textbf{Mo}$del-based policies for $\textbf{Vie}$w generalization ($\textbf{MoVie}$) during test time, without any need for explicit reward signals and any modification during training time. Our method demonstrates substantial advancements across all four scenarios encompassing a total of $\textbf{18}$ tasks sourced from DMControl, xArm, and Adroit, with a relative improvement of $\mathbf{33}$%, $\mathbf{86}$%, and $\mathbf{152}$% respectively. The superior results highlight the immense potential of our approach for real-world robotics applications. Code and videos are available at https://yangsizhe.github.io/MoVie/. Sizhe Yang, Yanjie Ze, Huazhe Xu |
NeurIPS | 1 |
| 2023 | RL-ViGen: A Reinforcement Learning Benchmark for Visual GeneralizationabstractVisual Reinforcement Learning (Visual RL), coupled with high-dimensional observations, has consistently confronted the long-standing challenge of out-of-distribution generalization. Despite the focus on algorithms aimed at resolving visual generalization problems, we argue that the devil is in the existing benchmarks as they are restricted to isolated tasks and generalization categories, undermining a comprehensive evaluation of agents' visual generalization capabilities. To bridge this gap, we introduce RL-ViGen: a novel Reinforcement Learning Benchmark for Visual Generalization, which contains diverse tasks and a wide spectrum of generalization types, thereby facilitating the derivation of more reliable conclusions. Furthermore, RL-ViGen incorporates the latest generalization visual RL algorithms into a unified framework, under which the experiment results indicate that no single existing algorithm has prevailed universally across tasks. Our aspiration is that Rl-ViGen will serve as a catalyst in this area, and lay a foundation for the future creation of universal visual generalization RL agents suitable for real-world scenarios. Access to our code and implemented algorithms is provided at https://gemcollector.github.io/RL-ViGen/. Zhecheng Yuan, Sizhe Yang, Pu Hua, Can Chang, Kaizhe Hu, Huazhe Xu |
NeurIPS | 2 |