Yuan-Ting Hu

dblp:155/9881 · DBLP profile ↗
← Back
16ranked-venue papers
9as first author
7since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 8 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 7 first-author · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
15 papers
3D vision · 28% Video understanding and tracking · 25% Segmentation and scene understanding · 23%
Computer graphics and multimedia
3 papers
Image and video processing · 100%

Topics — the 30 heaviest of 41, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Video understanding and tracking
video object segmentation
1.342019
SAIL-VOS: Semantic Amodal Instance Level Video Object Segmentation - A Synthetic Dataset and Baselines · CVPR 2019
VideoMatch: Matching Based Video Object Segmentation · ECCV (8) 2018
Unsupervised Video Object Segmentation Using Motion Saliency-Guided Spatio-Temporal Propagation · ECCV (1) 2018
Computer vision › Segmentation and scene understanding
image segmentation
0.912025
SAM 2: Segment Anything in Images and Videos · ICLR 2025
Computer vision › Video understanding and tracking › video object segmentation
promptable video segmentation
0.912025
SAM 2: Segment Anything in Images and Videos · ICLR 2025
Computer vision › Segmentation and scene understanding
prompt-based segmentation
0.912025
SAM 2: Segment Anything in Images and Videos · ICLR 2025
Computer vision › Segmentation and scene understanding
video segmentation
0.912025
SAM 2: Segment Anything in Images and Videos · ICLR 2025
Computer vision › Image recognition and object detection
image classification
0.822023
Hiera: A Hierarchical Vision Transformer without the Bells-and-Whistles · ICML 2023
Total Variation Optimization Layers for Computer Vision · CVPR 2022
Computer vision › 3D vision
3d human reconstruction
0.712023
Occupancy Planes for Single-View RGB-D Human Reconstruction · AAAI 2023
Machine learning › Deep learning architectures and training › transformer › vision transformer
hierarchical vision transformer
0.712023
Hiera: A Hierarchical Vision Transformer without the Bells-and-Whistles · ICML 2023
Computer vision › 3D vision › 3d shape representation
implicit function
0.712023
Occupancy Planes for Single-View RGB-D Human Reconstruction · AAAI 2023
Computer vision › 3D vision › 3d shape reconstruction
object shape reconstruction
0.712023
Surface Snapping Optimization Layer for Single Image Object Shape Reconstruction · ICML 2023
Computer vision › 3D vision › 3d scene understanding
semantic scene completion
0.712023
Occupancy Planes for Single-View RGB-D Human Reconstruction · AAAI 2023
Computer vision › 3D vision › 3d reconstruction
single-view 3d reconstruction
0.712023
Surface Snapping Optimization Layer for Single Image Object Shape Reconstruction · ICML 2023
Computer vision › Video understanding and tracking
video classification
0.712023
Hiera: A Hierarchical Vision Transformer without the Bells-and-Whistles · ICML 2023
Machine learning › Optimization for machine learning › differentiable optimization
optimization layer
0.612022
Total Variation Optimization Layers for Computer Vision · CVPR 2022
Image and video processing › image restoration
image denoising
0.612022
Total Variation Optimization Layers for Computer Vision · CVPR 2022
Image and video processing › regularization
total variation regularization
0.612022
Total Variation Optimization Layers for Computer Vision · CVPR 2022
Computer vision › 3D vision
3d reconstruction
0.512021
SAIL-VOS 3D: A Synthetic Dataset and Baselines for Object Detection and 3D Mesh Reconstruction From Video Data · CVPR 2021
Computer vision › Video understanding and tracking
video object detection
0.512021
SAIL-VOS 3D: A Synthetic Dataset and Baselines for Object Detection and 3D Mesh Reconstruction From Video Data · CVPR 2021
Computer vision › Segmentation and scene understanding
semantic segmentation
0.522019
SAIL-VOS: Semantic Amodal Instance Level Video Object Segmentation - A Synthetic Dataset and Baselines · CVPR 2019
MaskRNN: Instance Level Video Object Segmentation · NIPS 2017
Computer vision › Video understanding and tracking › video reconstruction
video inpainting
0.412020
Proposal-Based Video Completion · ECCV (27) 2020
Image and video processing › video restoration
video inpainting
0.412020
Proposal-Based Video Completion · ECCV (27) 2020
Computer vision › Face, body and person analysis › human pose estimation
3d pose estimation
0.412019
Chirality Nets for Human Pose Regression · NeurIPS 2019
Computer vision › Segmentation and scene understanding › instance segmentation
amodal instance segmentation
0.412019
SAIL-VOS: Semantic Amodal Instance Level Video Object Segmentation - A Synthetic Dataset and Baselines · CVPR 2019
Machine learning › Generative modeling
generative adversarial network
0.412019
Max-Sliced Wasserstein Distance and Its Use for GANs · CVPR 2019
Computer vision › Face, body and person analysis
human pose estimation
0.412019
Chirality Nets for Human Pose Regression · NeurIPS 2019
Computer vision › Video understanding and tracking › action recognition
skeleton-based action recognition
0.412019
Chirality Nets for Human Pose Regression · NeurIPS 2019
Machine learning › Optimization for machine learning › optimal transport
sliced wasserstein distance
0.412019
Max-Sliced Wasserstein Distance and Its Use for GANs · CVPR 2019
Computer vision › Segmentation and scene understanding › video segmentation
video semantic segmentation
0.412019
SAIL-VOS: Semantic Amodal Instance Level Video Object Segmentation - A Synthetic Dataset and Baselines · CVPR 2019
Computer vision › Segmentation and scene understanding › image segmentation
unsupervised segmentation
0.312018
Unsupervised Video Object Segmentation Using Motion Saliency-Guided Spatio-Temporal Propagation · ECCV (1) 2018
Computer vision › 3D vision
3d scene understanding
0.322021
SAIL-VOS 3D: A Synthetic Dataset and Baselines for Object Detection and 3D Mesh Reconstruction From Video Data · CVPR 2021
SAIL-VOS: Semantic Amodal Instance Level Video Object Segmentation - A Synthetic Dataset and Baselines · CVPR 2019

Methods — techniques the papers use, named apart from their topics

projected newton method · 1.1deep network layer · 1.1transformer · 0.9streaming memory · 0.9vision transformer · 0.7per-point classification · 0.7optimization layer · 0.7masked autoencoder pretraining · 0.7implicit function learning · 0.7conjugate gradient · 0.7total variation · 0.6unsupervised learning · 0.2one-class SVM · 0.2geodesic distance · 0.2
YearPublicationVenuePosition
2025 SAM 2: Segment Anything in Images and Videos
abstract
We present Segment Anything Model 2 (SAM 2), a foundation model towards solving promptable visual segmentation in images and videos. We build a data engine, which improves model and data via user interaction, to collect the largest video segmentation dataset to date. Our model is a simple transformer architecture with streaming memory for real-time video processing. SAM 2 trained on our data provides strong performance across a wide range of tasks. In video segmentation, we observe better accuracy, using 3x fewer interactions than prior approaches. In image segmentation, our model is more accurate and 6x faster than the Segment Anything Model (SAM). We believe that our data, model, and insights will serve as a significant milestone for video segmentation and related perception tasks. We are releasing our main model, the dataset, an interactive demo and code.
Nikhila Ravi, Valentin Gabeur, Yuan-Ting Hu, Ronghang Hu, Chaitanya Ryali, Tengyu Ma 0005, Haitham Khedr, Roman Rädle, Chloé Rolland, Laura Gustafson, Eric Mintun, Junting Pan, Kalyan Vasudev Alwala, Nicolas Carion, Chao-Yuan Wu, Ross B. Girshick, Piotr Dollár, Christoph Feichtenhofer
ICLR3
2023 Occupancy Planes for Single-View RGB-D Human Reconstruction
abstract
Single-view RGB-D human reconstruction with implicit functions is often formulated as per-point classification. Specifically, a set of 3D locations within the view-frustum of the camera are first projected independently onto the image and a corresponding feature is subsequently extracted for each 3D location. The feature of each 3D location is then used to classify independently whether the corresponding 3D point is inside or outside the observed object. This procedure leads to sub-optimal results because correlations between predictions for neighboring locations are only taken into account implicitly via the extracted features. For more accurate results we propose the occupancy planes (OPlanes) representation, which enables to formulate single-view RGB-D human reconstruction as occupancy prediction on planes which slice through the camera's view frustum. Such a representation provides more flexibility than voxel grids and enables to better leverage correlations than per-point classification. On the challenging S3D data we observe a simple classifier based on the OPlanes representation to yield compelling results, especially in difficult situations with partial occlusions due to other objects and partial visibility, which haven't been addressed by prior work.
Xiaoming Zhao 0001, Yuan-Ting Hu, Zhongzheng Ren, Alexander G. Schwing
AAAI2
2023 Surface Snapping Optimization Layer for Single Image Object Shape Reconstruction
abstract
Reconstructing the 3D shape of objects observed in a single image is a challenging task. Recent approaches rely on visual cues extracted from a given image learned from a deep net. In this work, we leverage recent advances in monocular scene understanding to incorporate an additional geometric cue of surface normals. For this, we proposed a novel optimization layer that encourages the face normals of the reconstructed shape to be aligned with estimated surface normals. We develop a computationally efficient conjugate-gradient-based method that avoids the computation of a high-dimensional sparse matrix. We show this framework to achieve compelling shape reconstruction results on the challenging Pix3D and ShapeNet datasets.
Yuan-Ting Hu, Alexander G. Schwing, Raymond A. Yeh
ICML1
2023 Hiera: A Hierarchical Vision Transformer without the Bells-and-Whistles
abstract
Modern hierarchical vision transformers have added several vision-specific components in the pursuit of supervised classification performance. While these components lead to effective accuracies and attractive FLOP counts, the added complexity actually makes these transformers slower than their vanilla ViT counterparts. In this paper, we argue that this additional bulk is unnecessary. By pretraining with a strong visual pretext task (MAE), we can strip out all the bells-and-whistles from a state-of-the-art multi-stage vision transformer without losing accuracy. In the process, we create Hiera, an extremely simple hierarchical vision transformer that is more accurate than previous models while being significantly faster both at inference and during training. We evaluate Hiera on a variety of tasks for image and video recognition. Our code and models are available at https://github.com/facebookresearch/hiera.
Chaitanya Ryali, Yuan-Ting Hu, Daniel Bolya, Chen Wei 0005, Haoqi Fan 0001, Po-Yao Huang 0001, Vaibhav Aggarwal, Arkabandhu Chowdhury, Omid Poursaeed, Judy Hoffman, Jitendra Malik, Yanghao Li, Christoph Feichtenhofer
ICML2
2022 Equivariance Discovery by Learned Parameter-Sharing
abstract
Designing equivariance as an inductive bias into deep-nets has been a prominent approach to build effective models, e.g., a convolutional neural network incorporates translation equivariance. However, incorporating these inductive biases requires knowledge about the equivariance properties of the data, which may not be available, e.g., when encountering a new domain. To address this, we study how to "discover interpretable equivariances" from data. Specifically, we formulate this discovery process as an optimization problem over a model’s parameter-sharing schemes. We propose to use the partition distance to empirically quantify the accuracy of the recovered equivariance. Also, we theoretically analyze the method for Gaussian data and provide a bound on the mean squared gap between the studied discovery scheme and the oracle scheme. Empirically, we show that the approach recovers known equivariances, such as permutations and shifts, on sum of numbers and spatially-invariant data.
Raymond A. Yeh, Yuan-Ting Hu, Mark Hasegawa-Johnson, Alexander G. Schwing
AISTATS2
2022 Total Variation Optimization Layers for Computer Vision
abstract
Optimization within a layer of a deep-net has emerged as a new direction for deep-net layer design. However, there are two main challenges when applying these layers to computer vision tasks: (a) which optimization problem within a layer is useful?; (b) how to ensure that computation within a layer remains efficient? To study question (a), in this work, we propose total variation (TV) minimization as a layer for computer vision. Motivated by the success of total variation in image processing, we hypothesize that TV as a layer provides useful inductive bias for deep-nets too. We study this hypothesis on five computer vision tasks: image classification, weakly supervised object localization, edge-preserving smoothing, edge detection, and image denoising, improving over existing baselines. To achieve these results we had to address question (b): we developed a GPU-based projected-Newton method which is 37× faster than existing solutions.
Raymond A. Yeh, Yuan-Ting Hu, Zhongzheng Ren, Alexander G. Schwing
CVPR2
2021 SAIL-VOS 3D: A Synthetic Dataset and Baselines for Object Detection and 3D Mesh Reconstruction From Video Data
abstract
Extracting detailed 3D information of objects from video data is an important goal for holistic scene understanding. While recent methods have shown impressive results when reconstructing meshes of objects from a single image, results often remain ambiguous as part of the object is unobserved. Moreover, existing image-based datasets for mesh reconstruction don’t permit to study models which integrate temporal information. To alleviate both concerns we present SAIL-VOS 3D: a synthetic video dataset with frame-by-frame mesh annotations which extends SAIL-VOS. We also develop first baselines for reconstruction of 3D meshes from video data via temporal models. We demonstrate efficacy of the proposed baseline on SAIL-VOS 3D and Pix3D, showing that temporal information improves reconstruction quality. Resources and additional information are available at http://sailvos.web.illinois.edu.
Yuan-Ting Hu, Jiahong Wang, Raymond A. Yeh, Alexander G. Schwing
CVPR1
2020 Proposal-Based Video Completion
Yuan-Ting Hu, Nicolas Ballas, Kristen Grauman, Alexander G. Schwing
ECCV (27)1
2019 Max-Sliced Wasserstein Distance and Its Use for GANs
abstract
Generative adversarial nets (GANs) and variational auto-encoders have significantly improved our distribution modeling capabilities, showing promise for dataset augmentation, image-to-image translation and feature learning. However, to model high-dimensional distributions, sequential training and stacked architectures are common, increasing the number of tunable hyper-parameters as well as the training time. Nonetheless, the sample complexity of the distance metrics remains one of the factors affecting GAN training. We first show that the recently proposed sliced Wasserstein distance has compelling sample complexity properties when compared to the Wasserstein distance. To further improve the sliced Wasserstein distance we then analyze its `projection complexity' and develop the max-sliced Wasserstein distance which enjoys compelling sample complexity while reducing projection complexity, albeit necessitating a max estimation. We finally illustrate that the proposed distance trains GANs on high-dimensional images up to a resolution of 256x256 easily.
Ishan Deshpande, Yuan-Ting Hu, Ruoyu Sun 0001, Ayis Pyrros, Nasir Siddiqui, Oluwasanmi Koyejo, Zhizhen Zhao 0001, David A. Forsyth, Alexander G. Schwing
CVPR2
2019 SAIL-VOS: Semantic Amodal Instance Level Video Object Segmentation - A Synthetic Dataset and Baselines
abstract
We introduce SAIL-VOS (Semantic Amodal Instance Level Video Object Segmentation), a new dataset aiming to stimulate semantic amodal segmentation research. Humans can effortlessly recognize partially occluded objects and reliably estimate their spatial extent beyond the visible. However, few modern computer vision techniques are capable of reasoning about occluded parts of an object. This is partly due to the fact that very few image datasets and no video dataset exist which permit development of those methods. To address this issue, we present a synthetic dataset extracted from the photo-realistic game GTA-V. Each frame is accompanied with densely annotated, pixel-accurate visible and amodal segmentation masks with semantic labels. More than 1.8M objects are annotated resulting in 100 times more annotations than existing datasets. We demonstrate the challenges of the dataset by quantifying the performance of several baselines. Data and additional material is available at http://sailvos.web.illinois.edu.
Yuan-Ting Hu, Hong-Shuo Chen, Kexin Hui, Jia-Bin Huang 0001, Alexander G. Schwing
CVPR1
2019 Chirality Nets for Human Pose Regression
abstract
We propose Chirality Nets, a family of deep nets that is equivariant to the “chirality transform,” i.e., the transformation to create a chiral pair. Through parameter sharing, odd and even symmetry, we propose and prove variants of standard building blocks of deep nets that satisfy the equivariance property, including fully connected layers, convolutional layers, batch-normalization, and LSTM/GRU cells. The proposed layers lead to a more data efficient representation and a reduction in computation by exploiting symmetry. We evaluate chirality nets on the task of human pose regression, which naturally exploits the left/right mirroring of the human body. We study three pose regression tasks: 3D pose estimation from video, 2D pose forecasting, and skeleton based activity recognition. Our approach achieves/matches state-of-the-art results, with more significant gains on small datasets and limited-data settings.
Raymond A. Yeh, Yuan-Ting Hu, Alexander G. Schwing
NeurIPS2
2018 Unsupervised Video Object Segmentation Using Motion Saliency-Guided Spatio-Temporal Propagation
Yuan-Ting Hu, Jia-Bin Huang 0001, Alexander G. Schwing
ECCV (1)1
2018 VideoMatch: Matching Based Video Object Segmentation
Yuan-Ting Hu, Jia-Bin Huang 0001, Alexander G. Schwing
ECCV (8)1
2017 MaskRNN: Instance Level Video Object Segmentation
abstract
Instance level video object segmentation is an important technique for video editing and compression. To capture the temporal coherence, in this paper, we develop MaskRNN, a recurrent neural net approach which fuses in each frame the output of two deep nets for each object instance - a binary segmentation net providing a mask and a localization net providing a bounding box. Due to the recurrent component and the localization component, our method is able to take advantage of long-term temporal structures of the video data as well as rejecting outliers. We validate the proposed algorithm on three challenging benchmark datasets, the DAVIS-2016 dataset, the DAVIS-2017 dataset, and the Segtrack v2 dataset, achieving state-of-the-art performance on all of them.
Yuan-Ting Hu, Jia-Bin Huang 0001, Alexander G. Schwing
NIPS1
2016 Progressive Feature Matching with Alternate Descriptor Selection and Correspondence Enrichment
abstract
We address two difficulties in establishing an accurate system for image matching. First, image matching relies on the descriptor for feature extraction, but the optimal descriptor often varies from image to image, or even patch to patch. Second, conventional matching approaches carry out geometric checking on a small set of correspondence candidates due to the concern of efficiency. It may result in restricted performance in recall. We aim at tackling the two issues by integrating adaptive descriptor selection and progressive candidate enrichment into image matching. We consider that the two integrated components are complementary: The high-quality matching yielded by adaptively selected descriptors helps in exploring more plausible candidates, while the enriched candidate set serves as a better reference for descriptor selection. It motivates us to formulate image matching as a joint optimization problem, in which adaptive descriptor selection and progressive correspondence enrichment are alternately conducted. Our approach is comprehensively evaluated and compared with the state-of-the-art approaches on two benchmarks. The promising results manifest its effectiveness.
Yuan-Ting Hu, Yen-Yu Lin
CVPR1
2015 Matching Images With Multiple Descriptors: An Unsupervised Approach for Locally Adaptive Descriptor Selection
abstract
With the aim to improve the performance of feature matching, we present an unsupervised approach for adaptive description selection in the space of homographies. Inspired by the observation that the homographies of correct feature correspondences vary smoothly along the spatial domain, our approach stands on the unsupervised nature of feature matching, and can choose a good descriptor locally for matching each feature point, instead of using one global descriptor. To this end, the homography space serves as the domain for selecting various heterogeneous descriptors. Correspondences obtained by any descriptors are considered as points in the space, and their geometric coherence and spatial continuity are measured via computing the geodesic distances. In this way, mutual verification across different descriptors is allowed, and correct correspondences will be highlighted with a high degree of consistency short geodesic distances here. It follows that one-class SVM can be applied to identifying these correct correspondences, and achieves adaptive descriptor selection. The proposed approach is comprehensively compared with the state-of-the-art approaches, and evaluated on five benchmarks of image matching. The promising results manifest its effectiveness.
Yuan-Ting Hu, Yen-Yu Lin, Hsin-Yi Chen, Kuang-Jui Hsu, Bing-Yu Chen 0004
IEEE Trans. Image Process.1