Jinghuan Shang

dblp:218/7364 · DBLP profile ↗
← Back
10ranked-venue papers
5as first author
9since 2021 · last 2025
0000-0001-7301-5981ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 5 first-author · 9 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
8 papers
Reinforcement learning · 28% Motion planning and robot control · 16% Representation and self-supervised learning · 14%
Databases, data mining, and information retrieval
1 paper
Data mining · 61% Web and social media mining · 39%

Topics — the 22 heaviest of 26, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning
imitation learning
1.022024
Crossway Diffusion: Improving Diffusion-based Visuomotor Policy via Self-supervised Learning · ICRA 2024
StARformer: Transformer With State-Action-Reward Representations for Robot Learning · IEEE Trans. Pattern Anal. Mach. Intell. 2023
Robotics › Motion planning and robot control › robot learning
robot policy learning
0.912025
LLaRA: Supercharging Robot Learning Data for Vision-Language Policy · ICLR 2025
Robotics › Robot manipulation › embodied foundation models
vision-language-action model
0.912025
LLaRA: Supercharging Robot Learning Data for Vision-Language Policy · ICLR 2025
Computer vision › Vision and language
vision-language model
0.912025
LLaRA: Supercharging Robot Learning Data for Vision-Language Policy · ICLR 2025
Machine learning › Learning paradigms › continual learning
catastrophic forgetting
0.812024
Continual Learning with Global Alignment · NeurIPS 2024
Machine learning › Learning paradigms
continual learning
0.812024
Continual Learning with Global Alignment · NeurIPS 2024
Robotics › Robot manipulation
diffusion policy
0.812024
Crossway Diffusion: Improving Diffusion-based Visuomotor Policy via Self-supervised Learning · ICRA 2024
Robotics › Motion planning and robot control › robot learning › visuomotor learning
visuomotor policy learning
0.812024
Crossway Diffusion: Improving Diffusion-based Visuomotor Policy via Self-supervised Learning · ICRA 2024
Robotics › Motion planning and robot control
robot learning
0.712023
StARformer: Transformer With State-Action-Reward Representations for Robot Learning · IEEE Trans. Pattern Anal. Mach. Intell. 2023
Machine learning › Deep learning architectures and training
sequence modeling
0.712023
StARformer: Transformer With State-Action-Reward Representations for Robot Learning · IEEE Trans. Pattern Anal. Mach. Intell. 2023
Machine learning › Representation and self-supervised learning
contrastive learning
0.612022
Does Self-supervised Learning Really Improve Reinforcement Learning from Pixels? · NeurIPS 2022
Machine learning › Reinforcement learning › deep reinforcement learning › visual reinforcement learning
reinforcement learning from pixels
0.612022
Does Self-supervised Learning Really Improve Reinforcement Learning from Pixels? · NeurIPS 2022
Machine learning › Reinforcement learning
self-supervised reinforcement learning
0.612022
Does Self-supervised Learning Really Improve Reinforcement Learning from Pixels? · NeurIPS 2022
Machine learning › Deep learning architectures and training
transformer
0.612022
StARformer: Transformer with State-Action-Reward Representations for Visual Reinforcement Learning · ECCV (39) 2022
Machine learning › Reinforcement learning › deep reinforcement learning
visual reinforcement learning
0.612022
StARformer: Transformer with State-Action-Reward Representations for Visual Reinforcement Learning · ECCV (39) 2022
Machine learning › Representation and self-supervised learning › representation learning › visual representation learning
visual token representation
0.612022
Learning Viewpoint-Agnostic Visual Representations by Recovering Tokens in 3D Space · NeurIPS 2022
Data mining › structured data mining › graph mining
community detection
0.312018
Dynamic Detection of Communities and Their Evolutions in Temporal Social Networks · AAAI 2018
Web and social media mining › online community analysis
community evolution
0.312018
Dynamic Detection of Communities and Their Evolutions in Temporal Social Networks · AAAI 2018
Data mining › structured data mining › graph mining › community detection
dynamic community detection
0.312018
Dynamic Detection of Communities and Their Evolutions in Temporal Social Networks · AAAI 2018
Machine learning › Reinforcement learning › imitation learning › offline imitation learning
behavior cloning
0.312025
LLaRA: Supercharging Robot Learning Data for Vision-Language Policy · ICLR 2025
Machine learning › Reinforcement learning
demonstration dataset
0.312025
LLaRA: Supercharging Robot Learning Data for Vision-Language Policy · ICLR 2025
Web and social media mining
social network analysis
0.112018
Dynamic Detection of Communities and Their Evolutions in Temporal Social Networks · AAAI 2018

Methods — techniques the papers use, named apart from their topics

transformer · 1.8reinforcement learning · 1.2visual instruction tuning · 0.9self-supervised auxiliary tasks · 0.9token representation composition · 0.8self-supervised learning · 0.8reconstruction objective · 0.8experience replay · 0.8diffusion model · 0.8intrinsic reward · 0.7temporal affiliation strength · 0.3normal distribution · 0.3
YearPublicationVenuePosition
2025 LLaRA: Supercharging Robot Learning Data for Vision-Language Policy
abstract
Vision Language Models (VLMs) have recently been leveraged to generate robotic actions, forming Vision-Language-Action (VLA) models. However, directly adapting a pretrained VLM for robotic control remains challenging, particularly when constrained by a limited number of robot demonstrations. In this work, we introduce LLaRA: Large Language and Robotics Assistant, a framework that formulates robot action policy as visuo-textual conversations and enables an efficient transfer of a pretrained VLM into a powerful VLA, motivated by the success of visual instruction tuning in Computer Vision. First, we present an automated pipeline to generate conversation-style instruction tuning data for robots from existing behavior cloning datasets, aligning robotic actions with image pixel coordinates. Further, we enhance this dataset in a self-supervised manner by defining six auxiliary tasks, without requiring any additional action annotations. We show that a VLM finetuned with a limited amount of such datasets can produce meaningful action decisions for robotic control. Through experiments across multiple simulated and real-world tasks, we demonstrate that LLaRA achieves state-of-the-art performance while preserving the generalization capabilities of large language models. The code, datasets, and pretrained models are available at https://github.com/LostXine/LLaRA.
Xiang Li 0109, Cristina Mata, Jongwoo Park 0003, Kumara Kahatapitiya, Yoo Sung Jang, Jinghuan Shang, Kanchana Ranasinghe, Ryan D. Burgert, Mu Cai, Yong Jae Lee, Michael S. Ryoo
ICLR6
2024 Crossway Diffusion: Improving Diffusion-based Visuomotor Policy via Self-supervised Learning
abstract
Diffusion models have been adopted for behavioral cloning in a sequence modeling fashion, benefiting from their exceptional capabilities in modeling complex data distributions. The standard diffusion-based policy iteratively denoises action sequences from random noise conditioned on the input states and the model is typically trained with a singular diffusion loss. This paper explores the potential enhancements in such models when the denoising process is informed by a better visual representation. We study the scenario where the model is jointly optimized using the standard diffusion loss alongside an auxiliary objective based on self-supervised learning. After experimenting with various objectives, we introduce Crossway Diffusion, a simple yet effective way to enhance diffusion-based visuomotor policy learning via a state decoder and an auxiliary reconstruction objective. During training, the state decoder reconstructs raw image pixels and other states from the intermediate representations of the model. Experiments demonstrate the effectiveness of our method in various simulated and real-world tasks, confirming its consistent advantages over the standard diffusion-based policy and other baselines.
Xiang Li 0109, Varun Belagali, Jinghuan Shang, Michael S. Ryoo
ICRA3
2024 Continual Learning with Global Alignment
abstract
Continual learning aims to sequentially learn new tasks without forgetting previous tasks' knowledge (catastrophic forgetting). One factor that can cause forgetting is the interference between the gradients on losses from different tasks. When the gradients on the current task's loss are in opposing directions to those on previous tasks' losses, updating the model for the current task may cause performance degradation on previous tasks. In this paper, we first identify causes of the above interference, and hypothesize that correlations between data representations are a key factor of interference. We then propose a method for promoting appropriate correlations between arbitrary tasks' data representations (i.e., global alignment) in individual task learning. Specifically, we learn the data representation as a task-specific composition of pre-trained token representations shared across all tasks. Then the correlations between different tasks' data representations are grounded by correlations between pre-trained token representations. We explore different ways to learn such compositions. Without experience replay, our model achieves SOTA performance in continual learning tasks. It also achieves advanced class-incremental performance through task-incremental training.
Xueying Bai, Jinghuan Shang, Niranjan Balasubramanian
NeurIPS2
2023 Active Vision Reinforcement Learning under Limited Visual Observability
abstract
In this work, we investigate Active Vision Reinforcement Learning (ActiveVision-RL), where an embodied agent simultaneously learns action policy for the task while also controlling its visual observations in partially observable environments. We denote the former as motor policy and the latter as sensory policy. For example, humans solve real world tasks by hand manipulation (motor policy) together with eye movements (sensory policy). ActiveVision-RL poses challenges on coordinating two policies given their mutual influence. We propose SUGARL, Sensorimotor Understanding Guided Active Reinforcement Learning, a framework that models motor and sensory policies separately, but jointly learns them using with an intrinsic sensorimotor reward. This learnable reward is assigned by sensorimotor reward module, incentivizes the sensory policy to select observations that are optimal to infer its own motor action, inspired by the sensorimotor stage of humans. Through a series of experiments, we show the effectiveness of our method across a range of observability conditions and its adaptability to existed RL algorithms. The sensory policies learned through our method are observed to exhibit effective active vision strategies.
Jinghuan Shang, Michael S. Ryoo
NeurIPS1
2023 StARformer: Transformer With State-Action-Reward Representations for Robot Learning
abstract
Reinforcement Learning (RL) can be considered as a sequence modeling task, where an agent employs a sequence of past state-action-reward experiences to predict a sequence of future actions. In this work, we propose State-Action-Reward Transformer (StARformer), a Transformer architecture for robot learning with image inputs, which explicitly models short-term state-action-reward representations (StAR-representations), essentially introducing a Markovian-like inductive bias to improve long-term modeling. StARformer first extracts StAR-representations using self-attending patches of image states, action, and reward tokens within a short temporal window. These StAR-representations are combined with pure image state representations, extracted as convolutional features, to perform self-attention over the whole sequence. Our experimental results show that StARformer outperforms the state-of-the-art Transformer-based method on image-based Atari and DeepMind Control Suite benchmarks, under both offline-RL and imitation learning settings. We find that models can benefit from our combination of patch-wise and convolutional image embeddings. StARformer is also more compliant with longer sequences of inputs than the baseline method. Finally, we demonstrate how StARformer can be successfully applied to a real-world robot imitation learning setting via a human-following task.
Jinghuan Shang, Xiang Li 0109, Kumara Kahatapitiya, Yu-Cheol Lee, Michael S. Ryoo
IEEE Trans. Pattern Anal. Mach. Intell.1
2022 StARformer: Transformer with State-Action-Reward Representations for Visual Reinforcement Learning
Jinghuan Shang, Kumara Kahatapitiya, Xiang Li 0109, Michael S. Ryoo
ECCV (39)1
2022 Does Self-supervised Learning Really Improve Reinforcement Learning from Pixels?
abstract
We investigate whether self-supervised learning (SSL) can improve online reinforcement learning (RL) from pixels. We extend the contrastive reinforcement learning framework (e.g., CURL) that jointly optimizes SSL and RL losses and conduct an extensive amount of experiments with various self-supervised losses. Our observations suggest that the existing SSL framework for RL fails to bring meaningful improvement over the baselines only taking advantage of image augmentation when the same amount of data and augmentation is used. We further perform evolutionary searches to find the optimal combination of multiple self-supervised losses for RL, but find that even such a loss combination fails to meaningfully outperform the methods that only utilize carefully designed image augmentations. After evaluating these approaches together in multiple different environments including a real-world robot environment, we confirm that no single self-supervised loss or image augmentation method can dominate all environments and that the current framework for joint optimization of SSL and RL is limited. Finally, we conduct the ablation study on multiple factors and demonstrate the properties of representations learned with different approaches.
Xiang Li 0109, Jinghuan Shang, Srijan Das, Michael S. Ryoo
NeurIPS2
2022 Learning Viewpoint-Agnostic Visual Representations by Recovering Tokens in 3D Space
abstract
Humans are remarkably flexible in understanding viewpoint changes due to visual cortex supporting the perception of 3D structure. In contrast, most of the computer vision models that learn visual representation from a pool of 2D images often fail to generalize over novel camera viewpoints. Recently, the vision architectures have shifted towards convolution-free architectures, visual Transformers, which operate on tokens derived from image patches. However, these Transformers do not perform explicit operations to learn viewpoint-agnostic representation for visual understanding. To this end, we propose a 3D Token Representation Layer (3DTRL) that estimates the 3D positional information of the visual tokens and leverages it for learning viewpoint-agnostic representations. The key elements of 3DTRL include a pseudo-depth estimator and a learned camera matrix to impose geometric transformations on the tokens, trained in an unsupervised fashion. These enable 3DTRL to recover the 3D positional information of the tokens from 2D patches. In practice, 3DTRL is easily plugged-in into a Transformer. Our experiments demonstrate the effectiveness of 3DTRL in many vision tasks including image classification, multi-view video alignment, and action recognition. The models with 3DTRL outperform their backbone Transformers in all the tasks with minimal added computation. Our code is available at https://github.com/elicassion/3DTRL.
Jinghuan Shang, Srijan Das, Michael S. Ryoo
NeurIPS1
2021 Self-Supervised Disentangled Representation Learning for Third-Person Imitation Learning
abstract
Humans learn to imitate by observing others. However, robot imitation learning generally requires expert demonstrations in the first-person view (FPV). Collecting such FPV videos for every robot could be very expensive.Third-person imitation learning (TPIL) is the concept of learning action policies by observing other agents in a third-person view (TPV), similar to what humans do. This ultimately allows utilizing human and robot demonstration videos in TPV from many different data sources, for the policy learning. In this paper, we present a TPIL approach for robot tasks with egomotion. Although many robot tasks with ground/aerial mobility often involve actions with camera egomotion, study on TPIL for such tasks has been limited. Here, FPV and TPV observations are visually very different; FPV shows egomotion while the agent appearance is only observable in TPV. To enable better state learning for TPIL, we propose our disentangled representation learning method. We use a dual auto-encoder structure plus representation permutation loss and time-contrastive loss to ensure the state and viewpoint representations are well disentangled. Our experiments show the effectiveness of our approach.
Jinghuan Shang, Michael S. Ryoo
IROS1
2018 Dynamic Detection of Communities and Their Evolutions in Temporal Social Networks
abstract
In this paper, we propose a novel community detection model, which explores the dynamic community evolutions in temporal social networks by modeling temporal affiliation strength between users and communities. Instead of transforming dynamic networks into static networks, our model utilizes normal distribution to estimate the change of affiliation strength more concisely and comprehensively. Extensive quantitative and qualitative evaluation on large social network datasets shows that our model achieves improvements in terms of prediction accuracy and reveals distinctive insight about evolutions of temporal social networks.
Yaowei Huang, Jinghuan Shang, Bill Y. Lin, Luoyi Fu, Xinbing Wang
AAAI2