VLDB 2026 Research / reviewers in the wild / expert
Jinghuan Shang
dblp:218/7364
· DBLP profile ↗
10ranked-venue papers
5as first author
9since 2021 · last 2025
0000-0001-7301-5981ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 5 first-author · 9 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
8 papers |
Reinforcement learning · 28% Motion planning and robot control · 16% Representation and self-supervised learning · 14% | |
| Databases, data mining, and information retrieval
1 paper |
Data mining · 61% Web and social media mining · 39% |
Topics — the 22 heaviest of 26, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning
imitation learning |
1.0 | 2 | 2024 | Crossway Diffusion: Improving Diffusion-based Visuomotor Policy via Self-supervised Learning · ICRA 2024 StARformer: Transformer With State-Action-Reward Representations for Robot Learning · IEEE Trans. Pattern Anal. Mach. Intell. 2023 |
Robotics › Motion planning and robot control › robot learning
robot policy learning |
0.9 | 1 | 2025 | LLaRA: Supercharging Robot Learning Data for Vision-Language Policy · ICLR 2025 |
Robotics › Robot manipulation › embodied foundation models
vision-language-action model |
0.9 | 1 | 2025 | LLaRA: Supercharging Robot Learning Data for Vision-Language Policy · ICLR 2025 |
Computer vision › Vision and language
vision-language model |
0.9 | 1 | 2025 | LLaRA: Supercharging Robot Learning Data for Vision-Language Policy · ICLR 2025 |
Machine learning › Learning paradigms › continual learning
catastrophic forgetting |
0.8 | 1 | 2024 | Continual Learning with Global Alignment · NeurIPS 2024 |
Machine learning › Learning paradigms
continual learning |
0.8 | 1 | 2024 | Continual Learning with Global Alignment · NeurIPS 2024 |
Robotics › Robot manipulation
diffusion policy |
0.8 | 1 | 2024 | Crossway Diffusion: Improving Diffusion-based Visuomotor Policy via Self-supervised Learning · ICRA 2024 |
Robotics › Motion planning and robot control › robot learning › visuomotor learning
visuomotor policy learning |
0.8 | 1 | 2024 | Crossway Diffusion: Improving Diffusion-based Visuomotor Policy via Self-supervised Learning · ICRA 2024 |
Robotics › Motion planning and robot control
robot learning |
0.7 | 1 | 2023 | StARformer: Transformer With State-Action-Reward Representations for Robot Learning · IEEE Trans. Pattern Anal. Mach. Intell. 2023 |
Machine learning › Deep learning architectures and training
sequence modeling |
0.7 | 1 | 2023 | StARformer: Transformer With State-Action-Reward Representations for Robot Learning · IEEE Trans. Pattern Anal. Mach. Intell. 2023 |
Machine learning › Representation and self-supervised learning
contrastive learning |
0.6 | 1 | 2022 | Does Self-supervised Learning Really Improve Reinforcement Learning from Pixels? · NeurIPS 2022 |
Machine learning › Reinforcement learning › deep reinforcement learning › visual reinforcement learning
reinforcement learning from pixels |
0.6 | 1 | 2022 | Does Self-supervised Learning Really Improve Reinforcement Learning from Pixels? · NeurIPS 2022 |
Machine learning › Reinforcement learning
self-supervised reinforcement learning |
0.6 | 1 | 2022 | Does Self-supervised Learning Really Improve Reinforcement Learning from Pixels? · NeurIPS 2022 |
Machine learning › Deep learning architectures and training
transformer |
0.6 | 1 | 2022 | StARformer: Transformer with State-Action-Reward Representations for Visual Reinforcement Learning · ECCV (39) 2022 |
Machine learning › Reinforcement learning › deep reinforcement learning
visual reinforcement learning |
0.6 | 1 | 2022 | StARformer: Transformer with State-Action-Reward Representations for Visual Reinforcement Learning · ECCV (39) 2022 |
Machine learning › Representation and self-supervised learning › representation learning › visual representation learning
visual token representation |
0.6 | 1 | 2022 | Learning Viewpoint-Agnostic Visual Representations by Recovering Tokens in 3D Space · NeurIPS 2022 |
Data mining › structured data mining › graph mining
community detection |
0.3 | 1 | 2018 | Dynamic Detection of Communities and Their Evolutions in Temporal Social Networks · AAAI 2018 |
Web and social media mining › online community analysis
community evolution |
0.3 | 1 | 2018 | Dynamic Detection of Communities and Their Evolutions in Temporal Social Networks · AAAI 2018 |
Data mining › structured data mining › graph mining › community detection
dynamic community detection |
0.3 | 1 | 2018 | Dynamic Detection of Communities and Their Evolutions in Temporal Social Networks · AAAI 2018 |
Machine learning › Reinforcement learning › imitation learning › offline imitation learning
behavior cloning |
0.3 | 1 | 2025 | LLaRA: Supercharging Robot Learning Data for Vision-Language Policy · ICLR 2025 |
Machine learning › Reinforcement learning
demonstration dataset |
0.3 | 1 | 2025 | LLaRA: Supercharging Robot Learning Data for Vision-Language Policy · ICLR 2025 |
Web and social media mining
social network analysis |
0.1 | 1 | 2018 | Dynamic Detection of Communities and Their Evolutions in Temporal Social Networks · AAAI 2018 |
Methods — techniques the papers use, named apart from their topics
transformer · 1.8reinforcement learning · 1.2visual instruction tuning · 0.9self-supervised auxiliary tasks · 0.9token representation composition · 0.8self-supervised learning · 0.8reconstruction objective · 0.8experience replay · 0.8diffusion model · 0.8intrinsic reward · 0.7temporal affiliation strength · 0.3normal distribution · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | LLaRA: Supercharging Robot Learning Data for Vision-Language PolicyabstractVision Language Models (VLMs) have recently been leveraged to generate robotic actions, forming Vision-Language-Action (VLA) models. However, directly adapting a pretrained VLM for robotic control remains challenging, particularly when constrained by a limited number of robot demonstrations. In this work, we introduce LLaRA: Large Language and Robotics Assistant, a framework that formulates robot action policy as visuo-textual conversations and enables an efficient transfer of a pretrained VLM into a powerful VLA, motivated by the success of visual instruction tuning in Computer Vision. First, we present an automated pipeline to generate conversation-style instruction tuning data for robots from existing behavior cloning datasets, aligning robotic actions with image pixel coordinates. Further, we enhance this dataset in a self-supervised manner by defining six auxiliary tasks, without requiring any additional action annotations. We show that a VLM finetuned with a limited amount of such datasets can produce meaningful action decisions for robotic control. Through experiments across multiple simulated and real-world tasks, we demonstrate that LLaRA achieves state-of-the-art performance while preserving the generalization capabilities of large language models. The code, datasets, and pretrained models are available at https://github.com/LostXine/LLaRA. Xiang Li 0109, Cristina Mata, Jongwoo Park 0003, Kumara Kahatapitiya, Yoo Sung Jang, Jinghuan Shang, Kanchana Ranasinghe, Ryan D. Burgert, Mu Cai, Yong Jae Lee, Michael S. Ryoo |
ICLR | 6 |
| 2024 | Crossway Diffusion: Improving Diffusion-based Visuomotor Policy via Self-supervised LearningabstractDiffusion models have been adopted for behavioral cloning in a sequence modeling fashion, benefiting from their exceptional capabilities in modeling complex data distributions. The standard diffusion-based policy iteratively denoises action sequences from random noise conditioned on the input states and the model is typically trained with a singular diffusion loss. This paper explores the potential enhancements in such models when the denoising process is informed by a better visual representation. We study the scenario where the model is jointly optimized using the standard diffusion loss alongside an auxiliary objective based on self-supervised learning. After experimenting with various objectives, we introduce Crossway Diffusion, a simple yet effective way to enhance diffusion-based visuomotor policy learning via a state decoder and an auxiliary reconstruction objective. During training, the state decoder reconstructs raw image pixels and other states from the intermediate representations of the model. Experiments demonstrate the effectiveness of our method in various simulated and real-world tasks, confirming its consistent advantages over the standard diffusion-based policy and other baselines. Xiang Li 0109, Varun Belagali, Jinghuan Shang, Michael S. Ryoo |
ICRA | 3 |
| 2024 | Continual Learning with Global AlignmentabstractContinual learning aims to sequentially learn new tasks without forgetting previous tasks' knowledge (catastrophic forgetting). One factor that can cause forgetting is the interference between the gradients on losses from different tasks. When the gradients on the current task's loss are in opposing directions to those on previous tasks' losses, updating the model for the current task may cause performance degradation on previous tasks. In this paper, we first identify causes of the above interference, and hypothesize that correlations between data representations are a key factor of interference. We then propose a method for promoting appropriate correlations between arbitrary tasks' data representations (i.e., global alignment) in individual task learning. Specifically, we learn the data representation as a task-specific composition of pre-trained token representations shared across all tasks. Then the correlations between different tasks' data representations are grounded by correlations between pre-trained token representations. We explore different ways to learn such compositions. Without experience replay, our model achieves SOTA performance in continual learning tasks. It also achieves advanced class-incremental performance through task-incremental training. Xueying Bai, Jinghuan Shang, Niranjan Balasubramanian |
NeurIPS | 2 |
| 2023 | Active Vision Reinforcement Learning under Limited Visual ObservabilityabstractIn this work, we investigate Active Vision Reinforcement Learning (ActiveVision-RL), where an embodied agent simultaneously learns action policy for the task while also controlling its visual observations in partially observable environments. We denote the former as motor policy and the latter as sensory policy. For example, humans solve real world tasks by hand manipulation (motor policy) together with eye movements (sensory policy). ActiveVision-RL poses challenges on coordinating two policies given their mutual influence. We propose SUGARL, Sensorimotor Understanding Guided Active Reinforcement Learning, a framework that models motor and sensory policies separately, but jointly learns them using with an intrinsic sensorimotor reward. This learnable reward is assigned by sensorimotor reward module, incentivizes the sensory policy to select observations that are optimal to infer its own motor action, inspired by the sensorimotor stage of humans. Through a series of experiments, we show the effectiveness of our method across a range of observability conditions and its adaptability to existed RL algorithms. The sensory policies learned through our method are observed to exhibit effective active vision strategies. Jinghuan Shang, Michael S. Ryoo |
NeurIPS | 1 |
| 2023 | StARformer: Transformer With State-Action-Reward Representations for Robot LearningabstractReinforcement Learning (RL) can be considered as a sequence modeling task, where an agent employs a sequence of past state-action-reward experiences to predict a sequence of future actions. In this work, we propose State-Action-Reward Transformer (StARformer), a Transformer architecture for robot learning with image inputs, which explicitly models short-term state-action-reward representations (StAR-representations), essentially introducing a Markovian-like inductive bias to improve long-term modeling. StARformer first extracts StAR-representations using self-attending patches of image states, action, and reward tokens within a short temporal window. These StAR-representations are combined with pure image state representations, extracted as convolutional features, to perform self-attention over the whole sequence. Our experimental results show that StARformer outperforms the state-of-the-art Transformer-based method on image-based Atari and DeepMind Control Suite benchmarks, under both offline-RL and imitation learning settings. We find that models can benefit from our combination of patch-wise and convolutional image embeddings. StARformer is also more compliant with longer sequences of inputs than the baseline method. Finally, we demonstrate how StARformer can be successfully applied to a real-world robot imitation learning setting via a human-following task. Jinghuan Shang, Xiang Li 0109, Kumara Kahatapitiya, Yu-Cheol Lee, Michael S. Ryoo |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | StARformer: Transformer with State-Action-Reward Representations for Visual Reinforcement Learning
Jinghuan Shang, Kumara Kahatapitiya, Xiang Li 0109, Michael S. Ryoo |
ECCV (39) | 1 |
| 2022 | Does Self-supervised Learning Really Improve Reinforcement Learning from Pixels?abstractWe investigate whether self-supervised learning (SSL) can improve online reinforcement learning (RL) from pixels. We extend the contrastive reinforcement learning framework (e.g., CURL) that jointly optimizes SSL and RL losses and conduct an extensive amount of experiments with various self-supervised losses. Our observations suggest that the existing SSL framework for RL fails to bring meaningful improvement over the baselines only taking advantage of image augmentation when the same amount of data and augmentation is used. We further perform evolutionary searches to find the optimal combination of multiple self-supervised losses for RL, but find that even such a loss combination fails to meaningfully outperform the methods that only utilize carefully designed image augmentations. After evaluating these approaches together in multiple different environments including a real-world robot environment, we confirm that no single self-supervised loss or image augmentation method can dominate all environments and that the current framework for joint optimization of SSL and RL is limited. Finally, we conduct the ablation study on multiple factors and demonstrate the properties of representations learned with different approaches. Xiang Li 0109, Jinghuan Shang, Srijan Das, Michael S. Ryoo |
NeurIPS | 2 |
| 2022 | Learning Viewpoint-Agnostic Visual Representations by Recovering Tokens in 3D SpaceabstractHumans are remarkably flexible in understanding viewpoint changes due to visual cortex supporting the perception of 3D structure. In contrast, most of the computer vision models that learn visual representation from a pool of 2D images often fail to generalize over novel camera viewpoints. Recently, the vision architectures have shifted towards convolution-free architectures, visual Transformers, which operate on tokens derived from image patches. However, these Transformers do not perform explicit operations to learn viewpoint-agnostic representation for visual understanding. To this end, we propose a 3D Token Representation Layer (3DTRL) that estimates the 3D positional information of the visual tokens and leverages it for learning viewpoint-agnostic representations. The key elements of 3DTRL include a pseudo-depth estimator and a learned camera matrix to impose geometric transformations on the tokens, trained in an unsupervised fashion. These enable 3DTRL to recover the 3D positional information of the tokens from 2D patches. In practice, 3DTRL is easily plugged-in into a Transformer. Our experiments demonstrate the effectiveness of 3DTRL in many vision tasks including image classification, multi-view video alignment, and action recognition. The models with 3DTRL outperform their backbone Transformers in all the tasks with minimal added computation. Our code is available at https://github.com/elicassion/3DTRL. Jinghuan Shang, Srijan Das, Michael S. Ryoo |
NeurIPS | 1 |
| 2021 | Self-Supervised Disentangled Representation Learning for Third-Person Imitation LearningabstractHumans learn to imitate by observing others. However, robot imitation learning generally requires expert demonstrations in the first-person view (FPV). Collecting such FPV videos for every robot could be very expensive.Third-person imitation learning (TPIL) is the concept of learning action policies by observing other agents in a third-person view (TPV), similar to what humans do. This ultimately allows utilizing human and robot demonstration videos in TPV from many different data sources, for the policy learning. In this paper, we present a TPIL approach for robot tasks with egomotion. Although many robot tasks with ground/aerial mobility often involve actions with camera egomotion, study on TPIL for such tasks has been limited. Here, FPV and TPV observations are visually very different; FPV shows egomotion while the agent appearance is only observable in TPV. To enable better state learning for TPIL, we propose our disentangled representation learning method. We use a dual auto-encoder structure plus representation permutation loss and time-contrastive loss to ensure the state and viewpoint representations are well disentangled. Our experiments show the effectiveness of our approach. Jinghuan Shang, Michael S. Ryoo |
IROS | 1 |
| 2018 | Dynamic Detection of Communities and Their Evolutions in Temporal Social NetworksabstractIn this paper, we propose a novel community detection model, which explores the dynamic community evolutions in temporal social networks by modeling temporal affiliation strength between users and communities. Instead of transforming dynamic networks into static networks, our model utilizes normal distribution to estimate the change of affiliation strength more concisely and comprehensively. Extensive quantitative and qualitative evaluation on large social network datasets shows that our model achieves improvements in terms of prediction accuracy and reveals distinctive insight about evolutions of temporal social networks. Yaowei Huang, Jinghuan Shang, Bill Y. Lin, Luoyi Fu, Xinbing Wang |
AAAI | 2 |