VLDB 2026 Research / reviewers in the wild / expert
Shuaiyi Huang
dblp:205/3109
· DBLP profile ↗
11ranked-venue papers
5as first author
6since 2021 · last 2025
0000-0003-0555-2077ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 4 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-author · 3 since 2021Systems, architecture and hardware · 2 · 2 first-author · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
7 papers |
Video understanding and tracking · 25% Reinforcement learning · 24% Robot manipulation · 18% | |
| Computer graphics and multimedia
2 papers |
Image and video coding · 50% Geometric modeling and processing · 30% Image and video processing · 20% |
Topics — the 22 heaviest of 23, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › 3D vision › correspondence estimation
semantic correspondence |
1.0 | 2 | 2022 | Learning Semantic Correspondence with Sparse Annotations · ECCV (14) 2022 Dynamic Context Correspondence Network for Semantic Alignment · ICCV 2019 |
Computer vision › Video understanding and tracking › action recognition
few-shot action recognition |
0.9 | 1 | 2025 | Trokens: Semantic-Aware Relational Trajectory Tokens for Few-Shot Action Recognition · ICCV 2025 |
Computer vision › Video understanding and tracking › motion analysis
motion modeling |
0.9 | 1 | 2025 | Trokens: Semantic-Aware Relational Trajectory Tokens for Few-Shot Action Recognition · ICCV 2025 |
Machine learning › Reinforcement learning › reinforcement learning from human feedback
preference-based reinforcement learning |
0.9 | 1 | 2025 | TREND: Tri-Teaching for Robust Preference-based Reinforcement Learning with Demonstrations · ICRA 2025 |
Machine learning › Reinforcement learning
reward learning |
0.9 | 1 | 2025 | TREND: Tri-Teaching for Robust Preference-based Reinforcement Learning with Demonstrations · ICRA 2025 |
Machine learning › Reinforcement learning › reward learning
reward modeling |
0.9 | 1 | 2025 | TREND: Tri-Teaching for Robust Preference-based Reinforcement Learning with Demonstrations · ICRA 2025 |
Computer vision › Video understanding and tracking
trajectory learning |
0.9 | 1 | 2025 | Trokens: Semantic-Aware Relational Trajectory Tokens for Few-Shot Action Recognition · ICCV 2025 |
Robotics › Robot manipulation › embodied foundation models
vision-language-action model |
0.9 | 1 | 2025 | TraceVLA: Visual Trace Prompting Enhances Spatial-Temporal Awareness for Generalist Robotic Policies · ICLR 2025 |
Geometric modeling and processing
implicit neural representation |
0.7 | 1 | 2023 | Towards Scalable Neural Representation for Diverse Videos · CVPR 2023 |
Image and video coding
neural video representation |
0.7 | 1 | 2023 | Towards Scalable Neural Representation for Diverse Videos · CVPR 2023 |
Machine learning › Efficient and distributed learning › data-efficient learning › label-efficient learning
sparse annotation learning |
0.6 | 1 | 2022 | Learning Semantic Correspondence with Sparse Annotations · ECCV (14) 2022 |
Image and video processing › image restoration
image dehazing |
0.4 | 1 | 2020 | Dehazing Evaluation: Real-World Benchmark Datasets, Criteria, and Baselines · IEEE Trans. Image Process. 2020 |
Image and video coding
image quality assessment |
0.4 | 1 | 2020 | Dehazing Evaluation: Real-World Benchmark Datasets, Criteria, and Baselines · IEEE Trans. Image Process. 2020 |
Machine learning › Deep learning architectures and training
attention mechanism |
0.4 | 1 | 2019 | Dynamic Context Correspondence Network for Semantic Alignment · ICCV 2019 |
Machine learning › Deep learning architectures and training › feature fusion
dynamic fusion |
0.4 | 1 | 2019 | Dynamic Context Correspondence Network for Semantic Alignment · ICCV 2019 |
Machine learning › Deep learning architectures and training › attention mechanism
structured attention |
0.3 | 1 | 2017 | Structured Attentions for Visual Question Answering · ICCV 2017 |
Machine learning › Deep learning architectures and training › attention mechanism
visual attention |
0.3 | 1 | 2017 | Structured Attentions for Visual Question Answering · ICCV 2017 |
Computer vision › Vision and language
visual question answering |
0.3 | 1 | 2017 | Structured Attentions for Visual Question Answering · ICCV 2017 |
Robotics › Motion planning and robot control › robot learning › robot policy learning
generalist robot policy |
0.3 | 1 | 2025 | TraceVLA: Visual Trace Prompting Enhances Spatial-Temporal Awareness for Generalist Robotic Policies · ICLR 2025 |
Robotics › Robot manipulation
learning from demonstration |
0.3 | 1 | 2025 | TREND: Tri-Teaching for Robust Preference-based Reinforcement Learning with Demonstrations · ICRA 2025 |
Robotics › Motion planning and robot control › robot learning
robot skill learning |
0.3 | 1 | 2025 | TREND: Tri-Teaching for Robust Preference-based Reinforcement Learning with Demonstrations · ICRA 2025 |
Computer vision › Video understanding and tracking
action recognition |
0.2 | 1 | 2023 | Towards Scalable Neural Representation for Diverse Videos · CVPR 2023 |
Methods — techniques the papers use, named apart from their topics
temporal reasoning · 1.3task-oriented flow · 1.3implicit neural representation · 1.3visual trace prompting · 0.9tri-teaching strategy · 0.9transformer · 0.9point tracking · 0.9histogram of oriented displacements · 0.9fine-tuning · 0.9few-shot demonstration · 0.9visibility index · 0.4realness index · 0.4full-reference image quality assessment · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Trokens: Semantic-Aware Relational Trajectory Tokens for Few-Shot Action RecognitionabstractVideo understanding requires effective modeling of both motion and appearance information, particularly for few-shot action recognition. While recent advances in point tracking have been shown to improve few-shot action recognition, two fundamental challenges persist: selecting informative points to track and effectively modeling their motion patterns. We present Trokens, a novel approach that transforms trajectory points into semantic-aware relational tokens for action recognition. First, we introduce a semantic-aware sampling strategy to adaptively distribute tracking points based on object scale and semantic relevance. Second, we develop a motion modeling framework that captures both intra-trajectory dynamics through the Histogram of Oriented Displacements (HoD) and inter-trajectory relationships to model complex action patterns. Our approach effectively combines these trajectory tokens with semantic features to enhance appearance features with motion information, achieving state-of-the-art performance across six diverse few-shot action recognition benchmarks: Something-Something-V2 (both full and small splits), Kinetics, UCF101, HMDB51, and FineGym. For project page see https://trokens-iccv25.github.io Pulkit Kumar, Shuaiyi Huang, Matthew Walmer, Sai Saketh Rambhatla, Abhinav Shrivastava |
ICCV | 2 |
| 2025 | TraceVLA: Visual Trace Prompting Enhances Spatial-Temporal Awareness for Generalist Robotic PoliciesabstractAlthough large vision-language-action (VLA) models pretrained on extensive robot datasets offer promising generalist policies for robotic learning, they still struggle with spatial-temporal dynamics in interactive robotics, making them less effective in handling complex tasks, such as manipulation. In this work, we introduce visual trace prompting, a simple yet effective approach to facilitate VLA models’ spatial-temporal awareness for action prediction by encoding state-action trajectories visually. We develop a new TraceVLA model by finetuning
OpenVLA on our own collected dataset of 150K robot manipulation trajectories using visual trace prompting. Evaluations of TraceVLA across 137 configurations in SimplerEnv and 4 tasks on a physical WidowX robot demonstrate state-of-the-art performance, outperforming OpenVLA by 10% on SimplerEnv and 3.5x on real-robot tasks and exhibiting robust generalization across diverse embodiments and scenarios. To further validate the effectiveness and generality of our method, we present a compact VLA model based on 4B Phi-3-Vision, pretrained on the Open-X-Embodiment and finetuned on our dataset, rivals the 7B OpenVLA baseline while significantly improving inference efficiency. Ruijie Zheng, Yongyuan Liang, Shuaiyi Huang, Jianfeng Gao 0001, Hal Daumé III, Andrey Kolobov, Furong Huang |
ICLR | 3 |
| 2025 | TREND: Tri-Teaching for Robust Preference-based Reinforcement Learning with DemonstrationsabstractPreference feedback collected by human or VLM annotators is often noisy, presenting a significant challenge for preference-based reinforcement learning that relies on accurate preference labels. To address this challenge, we propose TREND, a novel framework that integrates few-shot expert demonstrations with a tri-teaching strategy for effective noise mitigation. Our method trains three reward models simultaneously, where each model views its small-loss preference pairs as useful knowledge and teaches such useful pairs to its peer network for updating the parameters. Remarkably, our approach requires as few as one to three expert demonstrations to achieve high performance. We evaluate TREND on various robotic manipulation tasks, achieving up to 90% success rates even with noise levels as high as 40%, highlighting its effective robustness in handling noisy preference feedback. Shuaiyi Huang, Mara Levy, Daniel Ekpo, Ruijie Zheng, Abhinav Shrivastava |
ICRA | 1 |
| 2024 | ARDuP: Active Region Video Diffusion for Universal PoliciesabstractSequential decision-making can be formulated as a text-conditioned video generation problem, where a video planner, guided by a text-defined goal, generates future frames visualizing planned actions, from which control actions are subsequently derived. In this work, we introduce Active Region Video Diffusion for Universal Policies (ARDuP), a novel framework for video-based policy learning that emphasizes the generation of active regions, i.e. potential interaction areas, enhancing the conditional policy’s focus on interactive areas critical for task execution. This innovative framework integrates active region conditioning with latent diffusion models for video planning and employs latent representations for direct action decoding during inverse dynamic modeling. By utilizing motion cues in videos for automatic active region discovery, our method eliminates the need for manual annotations of active regions. We validate ARDuP’s efficacy via extensive experiments on simulator CLIPort and the real-world dataset BridgeData v2, achieving notable improvements in success rates and generating convincingly realistic video plans. Shuaiyi Huang, Mara Levy, Zhenyu Jiang 0002, Anima Anandkumar, Yuke Zhu, Linxi Fan, De-An Huang, Abhinav Shrivastava |
IROS | 1 |
| 2023 | Towards Scalable Neural Representation for Diverse VideosabstractImplicit neural representations (INR) have gained increasing attention in representing 3D scenes and images, and have been recently applied to encode videos (e.g., NeRV [1], E-NeRV [2]). While achieving promising results, existing INR-based methods are limited to encoding a handful of short videos (e.g., seven 5-second videos in the UVG dataset) with redundant visual content, leading to a model design that fits individual video frames independently and is not efficiently scalable to a large number of diverse videos. This paper focuses on developing neural representations for a more practical setup - encoding long and/or a large number of videos with diverse visual content. We first show that instead of dividing videos into small subsets and encoding them with separate models, encoding long and diverse videos jointly with a unified model achieves better compression results. Based on this observation, we propose D-NeRV, a novel neural representation framework designed to encode diverse videos by (i) decoupling clip-specific visual content from motion information, (ii) introducing temporal reasoning into the implicit neural network, and (iii) employing the task-oriented flow as intermediate output to reduce spatial redundancies. Our new model largely surpasses NeRV and traditional video compression techniques on UCF101 and UVG datasets on the video compression task. Moreover, when used as an efficient data-loader, D-NeRV achieves 3%-10% higher accuracy than NeRV on action recognition tasks on the UCF101 dataset under the same compression ratios. Bo He 0004, Xitong Yang, Hanyu Wang 0002, Zuxuan Wu, Hao Chen 0066, Shuaiyi Huang, Yixuan Ren, Ser-Nam Lim, Abhinav Shrivastava |
CVPR | 6 |
| 2022 | Learning Semantic Correspondence with Sparse Annotations
Shuaiyi Huang, Luyu Yang, Bo He 0004, Songyang Zhang 0001, Xuming He 0001, Abhinav Shrivastava |
ECCV (14) | 1 |
| 2020 | Confidence-Aware Adversarial Learning for Self-supervised Semantic Matching
Shuaiyi Huang, Qiuyue Wang, Xuming He 0001 |
PRCV (1) | 1 |
| 2020 | Dehazing Evaluation: Real-World Benchmark Datasets, Criteria, and BaselinesabstractOn benchmark images, modern dehazing methods are able to achieve very comparable results whose differences are too subtle for people to qualitatively judge. Thus, it is imperative to adopt quantitative evaluation on a vast number of hazy images. However, existing quantitative evaluation schemes are not convincing due to a lack of appropriate datasets and poor correlations between metrics and human perceptions. In this work, we attempt to address these issues, and we make two contributions. First, we establish two benchmark datasets, i.e., the BEnchmark Dataset for Dehazing Evaluation (BeDDE) and the EXtension of the BeDDE (exBeDDE), which had been lacking for a long period of time. The BeDDE is used to evaluate dehazing methods via full reference image quality assessment (FR-IQA) metrics. It provides hazy images, clear references, haze level labels, and manually labeled masks that indicate the regions of interest (ROIs) in image pairs. The exBeDDE is used to assess the performance of dehazing evaluation metrics. It provides extra dehazed images and subjective scores from people. To the best of our knowledge, the BeDDE is the first dehazing dataset whose image pairs were collected in natural outdoor scenes without any simulation. Second, we provide a new insight that dehazing involves two separate aspects, i.e., visibility restoration and realness restoration, which should be evaluated independently; thus, to characterize them, we establish two criteria, i.e., the visibility index (VI) and the realness index (RI), respectively. The effectiveness of the criteria is verified through extensive experiments. Furthermore, 14 representative dehazing methods are evaluated as baselines using our criteria on BeDDE. Our datasets and relevant code are available at https://github.com/xiaofeng94/BeDDE-for-defogging. Shiyu Zhao 0001, Lin Zhang 0014, Shuaiyi Huang, Ying Shen 0005, Shengjie Zhao 0001 |
IEEE Trans. Image Process. | 3 |
| 2019 | Dynamic Context Correspondence Network for Semantic AlignmentabstractEstablishing semantic correspondence is a core problem in computer vision and remains challenging due to large intra-class variations and lack of annotated data. In this paper, we aim to incorporate global semantic context in a flexible manner to overcome the limitations of prior work that relies on local semantic representations. To this end, we first propose a context-aware semantic representation that incorporates spatial layout for robust matching against local ambiguities. We then develop a novel dynamic fusion strategy based on attention mechanism to weave the advantages of both local and context features by integrating semantic cues from multiple scales. We instantiate our strategy by designing an end-to-end learnable deep network, named as Dynamic Context Correspondence Network (DCCNet). To train the network, we adopt a multi-auxiliary task loss to improve the efficiency of our weakly-supervised learning procedure. Our approach achieves superior or competitive performance over previous methods on several challenging datasets, including PF-Pascal, PF-Willow, and TSS, demonstrating its effectiveness and generality. Shuaiyi Huang, Qiuyue Wang, Songyang Zhang 0001, Shipeng Yan, Xuming He 0001 |
ICCV | 1 |
| 2019 | Evaluation of Defogging: A Real-World Benchmark Dataset, A New Criterion and BaselinesabstractModern defogging methods are able to achieve very comparable results whose differences are too subtle for people to qualitatively judge. On the other hand, existing quantitative evaluation methods are also not convincing due to a lack of proper datasets. In this work, we attempt to address these issues and establish a long-term lacking benchmark dataset, namely BeDDE (BEnchmark Dataset for Defogging Evaluation), for evaluating the performance of defogging algorithms. To our knowledge, BeDDE is the first real-world dataset comprising foggy images with their registered clear counterparts. Using BeDDE, we set up a new criterion for evaluating defogging methods where VSI, a full reference image quality assessment metric, is calculated and averaged on registered ROIs of all image pairs. The evaluation results of the proposed criterion correlate well with human judgements. 10 state-of-the-art defogging methods are evaluated as baselines on BeDDE. BeDDE is available online. Shiyu Zhao 0001, Lin Zhang 0014, Shuaiyi Huang, Ying Shen 0005, Shengjie Zhao 0001, Yukai Yang |
ICME | 3 |
| 2017 | Structured Attentions for Visual Question AnsweringabstractVisual attention, which assigns weights to image regions according to their relevance to a question, is considered as an indispensable part by most Visual Question Answering models. Although the questions may involve complex rela- tions among multiple regions, few attention models can ef- fectively encode such cross-region relations. In this paper, we demonstrate the importance of encoding such relations by showing the limited effective receptive field of ResNet on two datasets, and propose to model the visual attention as a multivariate distribution over a grid-structured Con- ditional Random Field on image regions. We demonstrate how to convert the iterative inference algorithms, Mean Field and Loopy Belief Propagation, as recurrent layers of an end-to-end neural network. We empirically evalu- ated our model on 3 datasets, in which it surpasses the best baseline model of the newly released CLEVR dataset [13] by 9.5%, and the best published model on the VQA dataset [3] by 1.25%. Source code is available at https://github.com/zhuchen03/vqa-sva. Yanpeng Zhao, Shuaiyi Huang, Kewei Tu |
ICCV | 3 |