Yunhao Luo 0001

dblp:305/4684 · DBLP profile ↗
← Back
7ranked-venue papers
3as first author
7since 2021 · last 2025
0009-0005-9321-3674ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 3 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Motion planning and robot control · 32% Reinforcement learning · 23% Segmentation and scene understanding · 17%

Topics — the 17 heaviest of 19, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Segmentation and scene understanding
semantic segmentation
1.422024
Transformer-CNN Cohort: Semi-supervised Semantic Segmentation by the Best of Both Students · ICRA 2024
A Good Student is Cooperative and Reliable: CNN-Transformer Collaborative Learning for Semantic Segmentation · ICCV 2023
Machine learning › Generative modeling
diffusion model
0.912025
Generative Trajectory Stitching through Diffusion Composition · NeurIPS 2025
Machine learning › Reinforcement learning
goal-conditioned reinforcement learning
0.912025
Grounding Video Models to Actions through Goal Conditioned Exploration · ICLR 2025
Robotics › Motion planning and robot control
robot learning
0.912025
Grounding Video Models to Actions through Goal Conditioned Exploration · ICLR 2025
Machine learning › Reinforcement learning › exploration › intrinsic motivation
self-supervised exploration
0.912025
Grounding Video Models to Actions through Goal Conditioned Exploration · ICLR 2025
Machine learning › Reinforcement learning › offline reinforcement learning
trajectory stitching
0.912025
Generative Trajectory Stitching through Diffusion Composition · NeurIPS 2025
Robotics › Motion planning and robot control › robot learning › robot policy learning
video-conditioned policy learning
0.912025
Grounding Video Models to Actions through Goal Conditioned Exploration · ICLR 2025
Robotics › Motion planning and robot control › motion planning
learning-based motion planning
0.812024
Potential Based Diffusion Motion Planning · ICML 2024
Robotics › Motion planning and robot control
motion planning
0.812024
Potential Based Diffusion Motion Planning · ICML 2024
Robotics › Motion planning and robot control › path planning
potential field planning
0.812024
Potential Based Diffusion Motion Planning · ICML 2024
Machine learning › Learning paradigms
semi-supervised learning
0.812024
Transformer-CNN Cohort: Semi-supervised Semantic Segmentation by the Best of Both Students · ICRA 2024
Machine learning › Efficient and distributed learning › model compression
knowledge distillation
0.712023
A Good Student is Cooperative and Reliable: CNN-Transformer Collaborative Learning for Semantic Segmentation · ICCV 2023
Computer vision › Segmentation and scene understanding › semantic segmentation › geometry-aware semantic segmentation
panoramic semantic segmentation
0.712023
Look at the Neighbor: Distortion-aware Unsupervised Domain Adaptation for Panoramic Semantic Segmentation · ICCV 2023
Machine learning › Transfer learning and domain adaptation › domain adaptation
unsupervised domain adaptation
0.712023
Look at the Neighbor: Distortion-aware Unsupervised Domain Adaptation for Panoramic Semantic Segmentation · ICCV 2023
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › planning
long-horizon planning
0.312025
Generative Trajectory Stitching through Diffusion Composition · NeurIPS 2025
Machine learning › Reinforcement learning
offline reinforcement learning
0.312025
Generative Trajectory Stitching through Diffusion Composition · NeurIPS 2025
Robotics › Robot navigation and mapping
visual navigation
0.312025
Grounding Video Models to Actions through Goal Conditioned Exploration · ICLR 2025

Methods — techniques the papers use, named apart from their topics

diffusion model · 1.6video models · 0.9trajectory-level action generation · 0.9compositional generation · 0.9bidirectional diffusion · 0.9behavior cloning · 0.9knowledge distillation · 0.8energy-based model · 0.8convolutional neural network · 0.8consistency regularization · 0.8
YearPublicationVenuePosition
2025 Grounding Video Models to Actions through Goal Conditioned Exploration
abstract
Large video models, pretrained on massive quantities of amount of Internet video, provide a rich source of physical knowledge about the dynamics and motions of objects and tasks. However, video models are not grounded in the embodiment of an agent, and do not describe how to actuate the world to reach the visual states depicted in a video. To tackle this problem, current methods use a separate vision-based inverse dynamic model trained on embodiment-specific data to map image states to actions. Gathering data to train such a model is often expensive and challenging, and this model is limited to visual settings similar to the ones in which data is available. In this paper, we investigate how to directly ground video models to continuous actions through self-exploration in the embodied environment -- using generated video states as visual goals for exploration. We propose a framework that uses trajectory level action generation in combination with video guidance to enable an agent to solve complex tasks without any external supervision, e.g., rewards, action labels, or segmentation masks. We validate the proposed approach on 8 tasks in Libero, 6 tasks in MetaWorld, 4 tasks in Calvin, and 12 tasks in iThor Visual Navigation. We show how our approach is on par with or even surpasses multiple behavior cloning baselines trained on expert demonstrations while without requiring any action annotations.
Yunhao Luo 0001, Yilun Du
ICLR1
2025 Generative Trajectory Stitching through Diffusion Composition
abstract
Effective trajectory stitching for long-horizon planning is a significant challenge in robotic decision-making. While diffusion models have shown promise in planning, they are limited to solving tasks similar to those seen in their training data. We propose CompDiffuser, a novel generative approach that can solve new tasks by learning to compositionally stitch together shorter trajectory chunks from previously seen tasks. Our key insight is modeling the trajectory distribution by subdividing it into overlapping chunks and learning their conditional relationships through a single bidirectional diffusion model. This allows information to propagate between segments during generation, ensuring physically consistent connections. We conduct experiments on benchmark tasks of various difficulties, covering different environment sizes, agent state dimension, trajectory types, training data quality, and show that CompDiffuser significantly outperforms existing methods.
Yunhao Luo 0001, Utkarsh A. Mishra, Yilun Du, Danfei Xu
NeurIPS1
2025 Distilling efficient Vision Transformers from CNNs for semantic segmentation
Xu Zheng 0002, Yunhao Luo 0001, Peng Yuan Zhou, Lin Wang 0025
Pattern Recognit.2
2024 Potential Based Diffusion Motion Planning
abstract
Effective motion planning in high dimensional spaces is a long-standing open problem in robotics. One class of traditional motion planning algorithms corresponds to potential-based motion planning. An advantage of potential based motion planning is composability – different motion constraints can easily combined by adding corresponding potentials. However, constructing motion paths from potentials requires solving a global optimization across configuration space potential landscape, which is often prone to local minima. We propose a new approach towards learning potential based motion planning, where we train a neural network to capture and learn an easily optimizable potentials over motion planning trajectories. We illustrate the effectiveness of such approach, significantly outperforming both classical and recent learned motion planning approaches and avoiding issues with local minima. We further illustrate its inherent composability, enabling us to generalize to a multitude of different motion constraints. Project website at https://energy-based-model.github.io/potential-motion-plan.
Yunhao Luo 0001, Chen Sun 0002, Josh Tenenbaum, Yilun Du
ICML1
2024 Transformer-CNN Cohort: Semi-supervised Semantic Segmentation by the Best of Both Students
abstract
The popular methods for semi-supervised semantic segmentation mostly adopt a unitary network model using convolutional neural networks (CNNs) and enforce consistency of the model’s predictions over perturbations applied to the inputs or model. However, such a learning paradigm suffers from two critical limitations: a) learning the discriminative features for the unlabeled data; b) learning both global and local information from the whole image. In this paper, we propose a novel Semi-supervised Learning (SSL) approach, called Transformer-CNN Cohort (TCC), that consists of two students with one based on the vision transformer (ViT) and the other based on the CNN. Our method subtly incorporates the multi-level consistency regularization on the predictions and the heterogeneous feature spaces via pseudo-labeling for the unlabeled data. First, as the inputs of the ViT student are image patches, the feature maps extracted encode crucial class-wise statistics. To this end, we propose class-aware feature consistency distillation (CFCD) that first leverages the outputs of each student as the pseudo labels and generates class-aware feature (CF) maps for knowledge transfer between the two students. Second, as the ViT student has more uniform representations for all layers, we propose consistency-aware cross distillation (CCD) to transfer knowledge between the pixel-wise predictions from the cohort. We validate the TCC framework on Cityscapes and Pascal VOC 2012 datasets, which outperforms existing SSL methods by a large margin. Project page: https://vlislab22.github.io/TCC/.
Xu Zheng 0002, Yunhao Luo 0001, Chong Fu 0001, Kangcheng Liu, Lin Wang 0025
ICRA2
2023 Look at the Neighbor: Distortion-aware Unsupervised Domain Adaptation for Panoramic Semantic Segmentation
abstract
Endeavors have been recently made to transfer knowledge from the labeled pinhole image domain to the unlabeled panoramic image domain via Unsupervised Domain Adaptation (UDA). The aim is to tackle the domain gaps caused by the style disparities and distortion problem from the non-uniformly distributed pixels of equirectangular projection (ERP). Previous works typically focus on transferring knowledge based on geometric priors with specially designed multi-branch network architectures. As a result, considerable computational costs are induced, and meanwhile, their generalization abilities are profoundly hindered by the variation of distortion among pixels. In this paper, we find that the pixels’ neighborhood regions of the ERP indeed introduce less distortion. Intuitively, we propose a novel UDA framework that can effectively address the distortion problems for panoramic semantic segmentation. In comparison, our method is simpler, easier to implement, and more computationally efficient. Specifically, we propose distortion-aware attention (DA) capturing the neighboring pixel distribution without using any geometric constraints. Moreover, we propose a class-wise feature aggregation (CFA) module to iteratively update the feature representations with a memory bank. As such, the feature similarity between two domains can be consistently optimized. Extensive experiments show that our method achieves new state-of-the-art performance while remarkably reducing 80% parameters.
Xu Zheng 0002, Tianbo Pan, Yunhao Luo 0001, Lin Wang 0025
ICCV3
2023 A Good Student is Cooperative and Reliable: CNN-Transformer Collaborative Learning for Semantic Segmentation
abstract
In this paper, we strive to answer the question ‘how to collaboratively learn convolutional neural network (CNN)-based and vision transformer (ViT)-based models by selecting and exchanging the reliable knowledge between them for semantic segmentation?’ Accordingly, we propose an online knowledge distillation (KD) framework that can simultaneously learn compact yet effective CNN-based and ViT-based models with two key technical breakthroughs to take full advantage of CNNs and ViT while compensating their limitations. Firstly, we propose heterogeneous feature distillation (HFD) to improve students’ consistency in low-layer feature space by mimicking heterogeneous features between CNNs and ViT. Secondly, to facilitate the two students to learn reliable knowledge from each other, we propose bidirectional selective distillation (BSD) that can dynamically transfer selective knowledge. This is achieved by 1) region-wise BSD determining the directions of knowledge transferred between the corresponding regions in the feature space and 2) pixel-wise BSD discerning which of the prediction knowledge to be transferred in the logit space. Extensive experiments on three benchmark datasets demonstrate that our proposed framework outperforms the state-of-the-art online distillation methods by a large margin, and shows its efficacy in learning collaboratively between ViT-based and CNN-based models.
Jinjing Zhu, Yunhao Luo 0001, Xu Zheng 0002, Hao Wang 0005, Lin Wang 0025
ICCV2