Shuaiyi Huang

dblp:205/3109 · DBLP profile ↗
← Back
11ranked-venue papers
5as first author
6since 2021 · last 2025
0000-0003-0555-2077ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 4 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-author · 3 since 2021Systems, architecture and hardware · 2 · 2 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
7 papers
Video understanding and tracking · 25% Reinforcement learning · 24% Robot manipulation · 18%
Computer graphics and multimedia
2 papers
Image and video coding · 50% Geometric modeling and processing · 30% Image and video processing · 20%

Topics — the 22 heaviest of 23, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › 3D vision › correspondence estimation
semantic correspondence
1.022022
Learning Semantic Correspondence with Sparse Annotations · ECCV (14) 2022
Dynamic Context Correspondence Network for Semantic Alignment · ICCV 2019
Computer vision › Video understanding and tracking › action recognition
few-shot action recognition
0.912025
Trokens: Semantic-Aware Relational Trajectory Tokens for Few-Shot Action Recognition · ICCV 2025
Computer vision › Video understanding and tracking › motion analysis
motion modeling
0.912025
Trokens: Semantic-Aware Relational Trajectory Tokens for Few-Shot Action Recognition · ICCV 2025
Machine learning › Reinforcement learning › reinforcement learning from human feedback
preference-based reinforcement learning
0.912025
TREND: Tri-Teaching for Robust Preference-based Reinforcement Learning with Demonstrations · ICRA 2025
Machine learning › Reinforcement learning
reward learning
0.912025
TREND: Tri-Teaching for Robust Preference-based Reinforcement Learning with Demonstrations · ICRA 2025
Machine learning › Reinforcement learning › reward learning
reward modeling
0.912025
TREND: Tri-Teaching for Robust Preference-based Reinforcement Learning with Demonstrations · ICRA 2025
Computer vision › Video understanding and tracking
trajectory learning
0.912025
Trokens: Semantic-Aware Relational Trajectory Tokens for Few-Shot Action Recognition · ICCV 2025
Robotics › Robot manipulation › embodied foundation models
vision-language-action model
0.912025
TraceVLA: Visual Trace Prompting Enhances Spatial-Temporal Awareness for Generalist Robotic Policies · ICLR 2025
Geometric modeling and processing
implicit neural representation
0.712023
Towards Scalable Neural Representation for Diverse Videos · CVPR 2023
Image and video coding
neural video representation
0.712023
Towards Scalable Neural Representation for Diverse Videos · CVPR 2023
Machine learning › Efficient and distributed learning › data-efficient learning › label-efficient learning
sparse annotation learning
0.612022
Learning Semantic Correspondence with Sparse Annotations · ECCV (14) 2022
Image and video processing › image restoration
image dehazing
0.412020
Dehazing Evaluation: Real-World Benchmark Datasets, Criteria, and Baselines · IEEE Trans. Image Process. 2020
Image and video coding
image quality assessment
0.412020
Dehazing Evaluation: Real-World Benchmark Datasets, Criteria, and Baselines · IEEE Trans. Image Process. 2020
Machine learning › Deep learning architectures and training
attention mechanism
0.412019
Dynamic Context Correspondence Network for Semantic Alignment · ICCV 2019
Machine learning › Deep learning architectures and training › feature fusion
dynamic fusion
0.412019
Dynamic Context Correspondence Network for Semantic Alignment · ICCV 2019
Machine learning › Deep learning architectures and training › attention mechanism
structured attention
0.312017
Structured Attentions for Visual Question Answering · ICCV 2017
Machine learning › Deep learning architectures and training › attention mechanism
visual attention
0.312017
Structured Attentions for Visual Question Answering · ICCV 2017
Computer vision › Vision and language
visual question answering
0.312017
Structured Attentions for Visual Question Answering · ICCV 2017
Robotics › Motion planning and robot control › robot learning › robot policy learning
generalist robot policy
0.312025
TraceVLA: Visual Trace Prompting Enhances Spatial-Temporal Awareness for Generalist Robotic Policies · ICLR 2025
Robotics › Robot manipulation
learning from demonstration
0.312025
TREND: Tri-Teaching for Robust Preference-based Reinforcement Learning with Demonstrations · ICRA 2025
Robotics › Motion planning and robot control › robot learning
robot skill learning
0.312025
TREND: Tri-Teaching for Robust Preference-based Reinforcement Learning with Demonstrations · ICRA 2025
Computer vision › Video understanding and tracking
action recognition
0.212023
Towards Scalable Neural Representation for Diverse Videos · CVPR 2023

Methods — techniques the papers use, named apart from their topics

temporal reasoning · 1.3task-oriented flow · 1.3implicit neural representation · 1.3visual trace prompting · 0.9tri-teaching strategy · 0.9transformer · 0.9point tracking · 0.9histogram of oriented displacements · 0.9fine-tuning · 0.9few-shot demonstration · 0.9visibility index · 0.4realness index · 0.4full-reference image quality assessment · 0.4
YearPublicationVenuePosition
2025 Trokens: Semantic-Aware Relational Trajectory Tokens for Few-Shot Action Recognition
abstract
Video understanding requires effective modeling of both motion and appearance information, particularly for few-shot action recognition. While recent advances in point tracking have been shown to improve few-shot action recognition, two fundamental challenges persist: selecting informative points to track and effectively modeling their motion patterns. We present Trokens, a novel approach that transforms trajectory points into semantic-aware relational tokens for action recognition. First, we introduce a semantic-aware sampling strategy to adaptively distribute tracking points based on object scale and semantic relevance. Second, we develop a motion modeling framework that captures both intra-trajectory dynamics through the Histogram of Oriented Displacements (HoD) and inter-trajectory relationships to model complex action patterns. Our approach effectively combines these trajectory tokens with semantic features to enhance appearance features with motion information, achieving state-of-the-art performance across six diverse few-shot action recognition benchmarks: Something-Something-V2 (both full and small splits), Kinetics, UCF101, HMDB51, and FineGym. For project page see https://trokens-iccv25.github.io
Pulkit Kumar, Shuaiyi Huang, Matthew Walmer, Sai Saketh Rambhatla, Abhinav Shrivastava
ICCV2
2025 TraceVLA: Visual Trace Prompting Enhances Spatial-Temporal Awareness for Generalist Robotic Policies
abstract
Although large vision-language-action (VLA) models pretrained on extensive robot datasets offer promising generalist policies for robotic learning, they still struggle with spatial-temporal dynamics in interactive robotics, making them less effective in handling complex tasks, such as manipulation. In this work, we introduce visual trace prompting, a simple yet effective approach to facilitate VLA models’ spatial-temporal awareness for action prediction by encoding state-action trajectories visually. We develop a new TraceVLA model by finetuning OpenVLA on our own collected dataset of 150K robot manipulation trajectories using visual trace prompting. Evaluations of TraceVLA across 137 configurations in SimplerEnv and 4 tasks on a physical WidowX robot demonstrate state-of-the-art performance, outperforming OpenVLA by 10% on SimplerEnv and 3.5x on real-robot tasks and exhibiting robust generalization across diverse embodiments and scenarios. To further validate the effectiveness and generality of our method, we present a compact VLA model based on 4B Phi-3-Vision, pretrained on the Open-X-Embodiment and finetuned on our dataset, rivals the 7B OpenVLA baseline while significantly improving inference efficiency.
Ruijie Zheng, Yongyuan Liang, Shuaiyi Huang, Jianfeng Gao 0001, Hal Daumé III, Andrey Kolobov, Furong Huang
ICLR3
2025 TREND: Tri-Teaching for Robust Preference-based Reinforcement Learning with Demonstrations
abstract
Preference feedback collected by human or VLM annotators is often noisy, presenting a significant challenge for preference-based reinforcement learning that relies on accurate preference labels. To address this challenge, we propose TREND, a novel framework that integrates few-shot expert demonstrations with a tri-teaching strategy for effective noise mitigation. Our method trains three reward models simultaneously, where each model views its small-loss preference pairs as useful knowledge and teaches such useful pairs to its peer network for updating the parameters. Remarkably, our approach requires as few as one to three expert demonstrations to achieve high performance. We evaluate TREND on various robotic manipulation tasks, achieving up to 90% success rates even with noise levels as high as 40%, highlighting its effective robustness in handling noisy preference feedback.
Shuaiyi Huang, Mara Levy, Daniel Ekpo, Ruijie Zheng, Abhinav Shrivastava
ICRA1
2024 ARDuP: Active Region Video Diffusion for Universal Policies
abstract
Sequential decision-making can be formulated as a text-conditioned video generation problem, where a video planner, guided by a text-defined goal, generates future frames visualizing planned actions, from which control actions are subsequently derived. In this work, we introduce Active Region Video Diffusion for Universal Policies (ARDuP), a novel framework for video-based policy learning that emphasizes the generation of active regions, i.e. potential interaction areas, enhancing the conditional policy’s focus on interactive areas critical for task execution. This innovative framework integrates active region conditioning with latent diffusion models for video planning and employs latent representations for direct action decoding during inverse dynamic modeling. By utilizing motion cues in videos for automatic active region discovery, our method eliminates the need for manual annotations of active regions. We validate ARDuP’s efficacy via extensive experiments on simulator CLIPort and the real-world dataset BridgeData v2, achieving notable improvements in success rates and generating convincingly realistic video plans.
Shuaiyi Huang, Mara Levy, Zhenyu Jiang 0002, Anima Anandkumar, Yuke Zhu, Linxi Fan, De-An Huang, Abhinav Shrivastava
IROS1
2023 Towards Scalable Neural Representation for Diverse Videos
abstract
Implicit neural representations (INR) have gained increasing attention in representing 3D scenes and images, and have been recently applied to encode videos (e.g., NeRV [1], E-NeRV [2]). While achieving promising results, existing INR-based methods are limited to encoding a handful of short videos (e.g., seven 5-second videos in the UVG dataset) with redundant visual content, leading to a model design that fits individual video frames independently and is not efficiently scalable to a large number of diverse videos. This paper focuses on developing neural representations for a more practical setup - encoding long and/or a large number of videos with diverse visual content. We first show that instead of dividing videos into small subsets and encoding them with separate models, encoding long and diverse videos jointly with a unified model achieves better compression results. Based on this observation, we propose D-NeRV, a novel neural representation framework designed to encode diverse videos by (i) decoupling clip-specific visual content from motion information, (ii) introducing temporal reasoning into the implicit neural network, and (iii) employing the task-oriented flow as intermediate output to reduce spatial redundancies. Our new model largely surpasses NeRV and traditional video compression techniques on UCF101 and UVG datasets on the video compression task. Moreover, when used as an efficient data-loader, D-NeRV achieves 3%-10% higher accuracy than NeRV on action recognition tasks on the UCF101 dataset under the same compression ratios.
Bo He 0004, Xitong Yang, Hanyu Wang 0002, Zuxuan Wu, Hao Chen 0066, Shuaiyi Huang, Yixuan Ren, Ser-Nam Lim, Abhinav Shrivastava
CVPR6
2022 Learning Semantic Correspondence with Sparse Annotations
Shuaiyi Huang, Luyu Yang, Bo He 0004, Songyang Zhang 0001, Xuming He 0001, Abhinav Shrivastava
ECCV (14)1
2020 Confidence-Aware Adversarial Learning for Self-supervised Semantic Matching
Shuaiyi Huang, Qiuyue Wang, Xuming He 0001
PRCV (1)1
2020 Dehazing Evaluation: Real-World Benchmark Datasets, Criteria, and Baselines
abstract
On benchmark images, modern dehazing methods are able to achieve very comparable results whose differences are too subtle for people to qualitatively judge. Thus, it is imperative to adopt quantitative evaluation on a vast number of hazy images. However, existing quantitative evaluation schemes are not convincing due to a lack of appropriate datasets and poor correlations between metrics and human perceptions. In this work, we attempt to address these issues, and we make two contributions. First, we establish two benchmark datasets, i.e., the BEnchmark Dataset for Dehazing Evaluation (BeDDE) and the EXtension of the BeDDE (exBeDDE), which had been lacking for a long period of time. The BeDDE is used to evaluate dehazing methods via full reference image quality assessment (FR-IQA) metrics. It provides hazy images, clear references, haze level labels, and manually labeled masks that indicate the regions of interest (ROIs) in image pairs. The exBeDDE is used to assess the performance of dehazing evaluation metrics. It provides extra dehazed images and subjective scores from people. To the best of our knowledge, the BeDDE is the first dehazing dataset whose image pairs were collected in natural outdoor scenes without any simulation. Second, we provide a new insight that dehazing involves two separate aspects, i.e., visibility restoration and realness restoration, which should be evaluated independently; thus, to characterize them, we establish two criteria, i.e., the visibility index (VI) and the realness index (RI), respectively. The effectiveness of the criteria is verified through extensive experiments. Furthermore, 14 representative dehazing methods are evaluated as baselines using our criteria on BeDDE. Our datasets and relevant code are available at https://github.com/xiaofeng94/BeDDE-for-defogging.
Shiyu Zhao 0001, Lin Zhang 0014, Shuaiyi Huang, Ying Shen 0005, Shengjie Zhao 0001
IEEE Trans. Image Process.3
2019 Dynamic Context Correspondence Network for Semantic Alignment
abstract
Establishing semantic correspondence is a core problem in computer vision and remains challenging due to large intra-class variations and lack of annotated data. In this paper, we aim to incorporate global semantic context in a flexible manner to overcome the limitations of prior work that relies on local semantic representations. To this end, we first propose a context-aware semantic representation that incorporates spatial layout for robust matching against local ambiguities. We then develop a novel dynamic fusion strategy based on attention mechanism to weave the advantages of both local and context features by integrating semantic cues from multiple scales. We instantiate our strategy by designing an end-to-end learnable deep network, named as Dynamic Context Correspondence Network (DCCNet). To train the network, we adopt a multi-auxiliary task loss to improve the efficiency of our weakly-supervised learning procedure. Our approach achieves superior or competitive performance over previous methods on several challenging datasets, including PF-Pascal, PF-Willow, and TSS, demonstrating its effectiveness and generality.
Shuaiyi Huang, Qiuyue Wang, Songyang Zhang 0001, Shipeng Yan, Xuming He 0001
ICCV1
2019 Evaluation of Defogging: A Real-World Benchmark Dataset, A New Criterion and Baselines
abstract
Modern defogging methods are able to achieve very comparable results whose differences are too subtle for people to qualitatively judge. On the other hand, existing quantitative evaluation methods are also not convincing due to a lack of proper datasets. In this work, we attempt to address these issues and establish a long-term lacking benchmark dataset, namely BeDDE (BEnchmark Dataset for Defogging Evaluation), for evaluating the performance of defogging algorithms. To our knowledge, BeDDE is the first real-world dataset comprising foggy images with their registered clear counterparts. Using BeDDE, we set up a new criterion for evaluating defogging methods where VSI, a full reference image quality assessment metric, is calculated and averaged on registered ROIs of all image pairs. The evaluation results of the proposed criterion correlate well with human judgements. 10 state-of-the-art defogging methods are evaluated as baselines on BeDDE. BeDDE is available online.
Shiyu Zhao 0001, Lin Zhang 0014, Shuaiyi Huang, Ying Shen 0005, Shengjie Zhao 0001, Yukai Yang
ICME3
2017 Structured Attentions for Visual Question Answering
abstract
Visual attention, which assigns weights to image regions according to their relevance to a question, is considered as an indispensable part by most Visual Question Answering models. Although the questions may involve complex rela- tions among multiple regions, few attention models can ef- fectively encode such cross-region relations. In this paper, we demonstrate the importance of encoding such relations by showing the limited effective receptive field of ResNet on two datasets, and propose to model the visual attention as a multivariate distribution over a grid-structured Con- ditional Random Field on image regions. We demonstrate how to convert the iterative inference algorithms, Mean Field and Loopy Belief Propagation, as recurrent layers of an end-to-end neural network. We empirically evalu- ated our model on 3 datasets, in which it surpasses the best baseline model of the newly released CLEVR dataset [13] by 9.5%, and the best published model on the VQA dataset [3] by 1.25%. Source code is available at https://github.com/zhuchen03/vqa-sva.
Yanpeng Zhao, Shuaiyi Huang, Kewei Tu
ICCV3