Hang Su 0005

dblp:26/5371-5 · DBLP profile ↗
← Back
10ranked-venue papers
3as first author
5since 2021 · last 2026
0000-0001-8770-8754ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 6 · 3 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 L4P: Towards Unified Low-Level 4D Vision Perception
abstract
The spatio-temporal relationship between the pixels of a video carries critical information for low-level 4D perception tasks. A single model that reasons about it should be able to solve several such tasks well. Yet, most state-of-theart methods rely on architectures specialized for the task at hand. We present L4P, a feedforward, general-purpose architecture that solves low-level 4D perception tasks in a unified framework. LAP leverages a pre-trained ViT-based video encoder and combines it with per-task heads that are lightweight and therefore do not require extensive training. Despite its general and feedforward formulation, our method is competitive with existing specialized methods on both dense tasks, such as depth or optical flow estimation, and sparse tasks, such as 2D/3D tracking. Moreover, it solves all tasks at once in a time comparable to that of single-task methods.
Abhishek Badki, Hang Su 0005, Bowen Wen, Orazio Gallo
3DV2
2025 Zero-Shot Monocular Scene Flow Estimation in the Wild
abstract
Large models have shown generalization across datasets for many low-level vision tasks, like depth estimation, but no such general models exist for scene flow. Even though scene flow prediction has wide potential, its practical use is limited because of the lack of generalization of current predictive models. We identify three key challenges and propose solutions for each. First, we create a method that jointly estimates geometry and motion for accurate prediction. Second, we alleviate scene flow data scarcity with a data recipe that affords us 1M annotated training samples across diverse synthetic scenes. Third, we evaluate different parameterizations for scene flow prediction and adopt a natural and effective parameterization. Our model outperforms existing methods as well as baselines built on large-scale models in terms of 3D end-point error, and shows zero-shot generalization to the casually captured videos from DAVIS and the robotic manipulation scenes from RoboTAP. Overall, our approach makes scene flow prediction more practical in-the-wild.Website: https://research.nvidia.com/labs/lpr/zeromsf/
Yiqing Liang, Abhishek Badki, Hang Su 0005, James Tompkin 0001, Orazio Gallo
CVPR3
2025 HybridDepth: Robust Metric Depth Fusion by Leveraging Depth from Focus and Single-Image Priors
abstract
We propose Hybriddepth, a robust depth estimation pipeline that addresses key challenges in depth estimation, including scale ambiguity, hardware heterogene-ity, and generalizability. Hybriddepth leverages focal stack, data conveniently accessible in common mobile de-vices, to produce accurate metric depth maps. By incorpo-rating depth priors afforded by recent advances in single-image depth estimation, our model achieves a higher level of structural detail compared to existing methods. We test our pipeline as an end-to-end system, with a newly developed mobile client to capture focal stacks, which are then sent to a GPU-powered server for depth estimation. Comprehensive quantitative and qualitative analyses demonstrate that Hybriddepth outperforms state-of-the-art (SOTA) models on common datasets such as DDFF12 and NYU Depth V2. Hybriddepth also shows strong zero-shot generalization. When trained on NYU Depth V2, Hybriddepth surpasses SOTA models in zero-shot per-formance on ARKitScenes and delivers more structurally accurate depth maps on Mobile Depth. The code is avail-able at https://github.com/cake-labIHybridDepth/.
Ashkan Ganj, Hang Su 0005, Tian Guo 0001
WACV2
2024 FoVA-Depth: Field-of-View Agnostic Depth Estimation for Cross-Dataset Generalization
abstract
Wide field-of-view (FoV) cameras efficiently capture large portions of the scene, which makes them attractive in multiple domains, such as automotive and robotics. For such applications, estimating depth from multiple images is a critical task, and therefore, a large amount of ground truth (GT) data is available. Unfortunately, most of the GT data is for pinhole cameras, making it impossible to properly train depth estimation models for large-FoV cameras. We propose the first method to train a stereo depth estimation model on the widely available pinhole data, and to generalize it to data captured with larger FoVs. Our intuition is simple: We warp the training data to a canonical, large-FoV representation and augment it to allow a single network to reason about diverse types of distortions that otherwise would prevent generalization. We show strong generalization ability of our approach on both indoor and outdoor datasets, which was not possible with previous methods.
Daniel Lichy, Hang Su 0005, Abhishek Badki, Jan Kautz, Orazio Gallo
3DV2
2024 BlobGEN-3D: Compositional 3D-Consistent Freeview Image Generation with 3D Blobs
Chao Liu 0064, Weili Nie, Sifei Liu, Abhishek Badki, Hang Su 0005, Morteza Mardani, Benjamin Eckart, Arash Vahdat
SIGGRAPH Asia5
2019 Pixel-Adaptive Convolutional Neural Networks
abstract
Convolutions are the fundamental building blocks of CNNs. The fact that their weights are spatially shared is one of the main reasons for their widespread use, but it is also a major limitation, as it makes convolutions content-agnostic. We propose a pixel-adaptive convolution (PAC) operation, a simple yet effective modification of standard convolutions, in which the filter weights are multiplied with a spatially varying kernel that depends on learnable, local pixel features. PAC is a generalization of several popular filtering techniques and thus can be used for a wide range of use cases. Specifically, we demonstrate state-of-the-art performance when PAC is used for deep joint image upsampling. PAC also offers an effective alternative to fully-connected CRF (Full-CRF), called PAC-CRF, which performs competitively compared to Full-CRF, while being considerably faster. In addition, we also demonstrate that PAC can be used as a drop-in replacement for convolution layers in pre-trained networks, resulting in consistent performance improvements.
Hang Su 0005, Varun Jampani, Deqing Sun, Orazio Gallo, Erik G. Learned-Miller, Jan Kautz
CVPR1
2018 SPLATNet: Sparse Lattice Networks for Point Cloud Processing
abstract
We present a network architecture for processing point clouds that directly operates on a collection of points represented as a sparse set of samples in a high-dimensional lattice. Naively applying convolutions on this lattice scales poorly, both in terms of memory and computational cost, as the size of the lattice increases. Instead, our network uses sparse bilateral convolutional layers as building blocks. These layers maintain efficiency by using indexing structures to apply convolutions only on occupied parts of the lattice, and allow flexible specifications of the lattice structure enabling hierarchical and spatially-aware feature learning, as well as joint 2D-3D reasoning. Both point-based and image-based representations can be easily incorporated in a network with such layers and the resulting model can be trained in an end-to-end manner. We present results on 3D segmentation tasks where our approach outperforms existing state-of-the-art techniques.
Hang Su 0005, Varun Jampani, Deqing Sun, Subhransu Maji, Evangelos Kalogerakis, Ming-Hsuan Yang 0001, Jan Kautz
CVPR1
2017 End-to-End Face Detection and Cast Grouping in Movies Using Erdös-Rényi Clustering
abstract
We present an end-to-end system for detecting and clustering faces by identity in full-length movies. Unlike works that start with a predefined set of detected faces, we consider the end-to-end problem of detection and clustering together. We make three separate contributions. First, we combine a state-of-the-art face detector with a generic tracker to extract high quality face tracklets. We then introduce a novel clustering method, motivated by the classic graph theory results of Erdös and Rényi. It is based on the observations that large clusters can be fully connected by joining just a small fraction of their point pairs, while just a single connection between two different people can lead to poor clustering results. This suggests clustering using a verification system with very few false positives but perhaps moderate recall. We introduce a novel verification method, rank-1 counts verification, that has this property, and use it in a link-based clustering scheme. Finally, we define a novel end-to-end detection and clustering evaluation metric allowing us to assess the accuracy of the entire end-to-end system. We present state-of-the-art results on multiple video data sets and also on standard face databases.
SouYoung Jin, Hang Su 0005, Chris Stauffer, Erik G. Learned-Miller
ICCV2
2015 Multi-view Convolutional Neural Networks for 3D Shape Recognition
abstract
A longstanding question in computer vision concerns the representation of 3D shapes for recognition: should 3D shapes be represented with descriptors operating on their native 3D formats, such as voxel grid or polygon mesh, or can they be effectively represented with view-based descriptors? We address this question in the context of learning to recognize 3D shapes from a collection of their rendered views on 2D images. We first present a standard CNN architecture trained to recognize the shapes' rendered views independently of each other, and show that a 3D shape can be recognized even from a single view at an accuracy far higher than using state-of-the-art 3D shape descriptors. Recognition rates further increase when multiple views of the shapes are provided. In addition, we present a novel CNN architecture that combines information from multiple views of a 3D shape into a single and compact shape descriptor offering even better recognition performance. The same architecture can be applied to accurately recognize human hand-drawn sketches of shapes. We conclude that a collection of 2D views can be highly informative for 3D shape recognition and is amenable to emerging CNN architectures and their derivatives.
Hang Su 0005, Subhransu Maji, Evangelos Kalogerakis, Erik G. Learned-Miller
ICCV1
2014 The SUN Attribute Database: Beyond Categories for Deeper Scene Understanding
Genevieve Patterson, Hang Su 0005, James Hays
Int. J. Comput. Vis.3