Peiyun Hu

dblp:123/2935 · DBLP profile ↗
← Back
14ranked-venue papers
5as first author
5since 2021 · last 2025
0000-0002-8498-0653ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 5 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 4 first-author · 3 since 2021Systems, architecture and hardware · 1Security and privacy · 1Human-computer interaction and ubiquitous computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
11 papers
3D vision · 54% Autonomous driving · 13% Motion planning and robot control · 8%
Computer graphics and multimedia
1 paper
Multimedia analysis and retrieval · 100%

Topics — the 29 heaviest of 34, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Robotics › Autonomous driving
perception
1.332025
Differentiable Raycasting for Self-Supervised Occupancy Forecasting · ECCV (38) 2022
Active Perception Using Light Curtains for Autonomous Driving · ECCV (5) 2020
Lidar Panoptic Segmentation in an Open World · Int. J. Comput. Vis. 2025
Robotics › Motion planning and robot control
motion planning
1.222023
Point Cloud Forecasting as a Proxy for 4D Occupancy Forecasting · CVPR 2023
Safe Local Motion Planning With Self-Supervised Freespace Forecasting · CVPR 2021
Computer vision › 3D vision
3d scene understanding
0.912025
Lidar Panoptic Segmentation in an Open World · Int. J. Comput. Vis. 2025
Computer vision › 3D vision › point cloud segmentation › LiDAR segmentation
LiDAR panoptic segmentation
0.912025
Lidar Panoptic Segmentation in an Open World · Int. J. Comput. Vis. 2025
Computer vision › 3D vision › 3d human pose estimation
multi-person 3d pose estimation
0.912025
CoMotion: Concurrent Multi-person 3D Motion · ICLR 2025
Computer vision › Segmentation and scene understanding
panoptic segmentation
0.912025
Lidar Panoptic Segmentation in an Open World · Int. J. Comput. Vis. 2025
Computer vision › 3D vision › pose estimation
pose tracking
0.912025
CoMotion: Concurrent Multi-person 3D Motion · ICLR 2025
Computer vision › 3D vision › 3d scene understanding › semantic scene completion
4d occupancy forecasting
0.712023
Point Cloud Forecasting as a Proxy for 4D Occupancy Forecasting · CVPR 2023
Computer vision › 3D vision › 3d scene understanding
semantic scene completion
0.612022
Differentiable Raycasting for Self-Supervised Occupancy Forecasting · ECCV (38) 2022
Computer vision › 3D vision
3d object detection
0.412020
What You See is What You Get: Exploiting Visibility for 3D Object Detection · CVPR 2020
Computer vision › 3D vision › 3d scene modeling › scene representation
3d scene representation
0.412020
What You See is What You Get: Exploiting Visibility for 3D Object Detection · CVPR 2020
Robotics › Robot navigation and mapping
active perception
0.412020
Active Perception Using Light Curtains for Autonomous Driving · ECCV (5) 2020
Computer vision › 3D vision › range sensing
depth sensing
0.412020
Active Perception Using Light Curtains for Autonomous Driving · ECCV (5) 2020
Computer vision › 3D vision › 3d object detection › point cloud object detection
LiDAR-based 3D object detection
0.412020
What You See is What You Get: Exploiting Visibility for 3D Object Detection · CVPR 2020
Computer vision › 3D vision › 3d shape representation › volumetric representation
voxel-based representation
0.412020
What You See is What You Get: Exploiting Visibility for 3D Object Detection · CVPR 2020
Machine learning › Efficient and distributed learning
active learning
0.412019
Active Learning with Partial Feedback · ICLR (Poster) 2019
Computer vision › Face, body and person analysis
face detection
0.312017
Finding Tiny Faces · CVPR 2017
Computer vision › Image recognition and object detection
object detection
0.312017
Finding Tiny Faces · CVPR 2017
Computer vision › Face, body and person analysis › face detection
small face detection
0.312017
Finding Tiny Faces · CVPR 2017
Computer vision › Video understanding and tracking › multi-object tracking
multi-person tracking
0.312025
CoMotion: Concurrent Multi-person 3D Motion · ICLR 2025
Machine learning › Probabilistic and Bayesian machine learning › hierarchical modeling › hierarchical bayesian model
hierarchical probabilistic model
0.212016
Bottom-Up and Top-Down Reasoning with Hierarchical Rectified Gaussians · CVPR 2016
Computer vision › 3D vision › low-level vision › feature detection
keypoint detection
0.212016
Bottom-Up and Top-Down Reasoning with Hierarchical Rectified Gaussians · CVPR 2016
Computer vision › 3D vision › point cloud processing › point cloud video understanding
point cloud forecasting
0.212023
Point Cloud Forecasting as a Proxy for 4D Occupancy Forecasting · CVPR 2023
Computer vision › 3D vision › 3d scene understanding › dynamic scene understanding
scene forecasting
0.112021
Safe Local Motion Planning With Self-Supervised Freespace Forecasting · CVPR 2021
Web and social media mining
social network analysis
0.112012
Can we understand van gogh's mood?: learning to infer affects from images in social networks · ACM Multimedia 2012
Multimedia analysis and retrieval › affective computing
affective multimedia analysis
0.112012
Understanding the emotional impact of images · ACM Multimedia 2012
Multimedia analysis and retrieval › affective computing
image emotion analysis
0.112012
Understanding the emotional impact of images · ACM Multimedia 2012
Computer vision › Image recognition and object detection › object detection
contextual reasoning
0.112017
Finding Tiny Faces · CVPR 2017
Machine learning › Probabilistic and Bayesian machine learning › structured models › graphical models
factor graphs
0.012012
Can we understand van gogh's mood?: learning to infer affects from images in social networks · ACM Multimedia 2012

Methods — techniques the papers use, named apart from their topics

self-supervised learning · 1.7pseudo-labeling · 0.9panoptic segmentation · 0.9monocular 3d pose estimation · 0.9LiDAR point cloud processing · 0.9LiDAR rendering · 0.7differentiable ray casting · 0.6learning-based planner · 0.5synthetic data augmentation · 0.4ray casting · 0.4semi-supervised learning · 0.3factor graph model · 0.3
YearPublicationVenuePosition
2025 CoMotion: Concurrent Multi-person 3D Motion
abstract
We introduce an approach for detecting and tracking detailed 3D poses of multiple people from a single monocular camera stream. Our system maintains temporally coherent predictions in crowded scenes filled with difficult poses and occlusions. Our model performs both strong per-frame detection and a learned pose update to track people from frame to frame. Rather than match detections across time, poses are updated directly from a new input image, which enables online tracking through occlusion. We train on numerous image and video datasets leveraging pseudo-labeled annotations to produce a model that matches state-of-the-art systems in 3D pose estimation accuracy while being faster and more accurate in tracking multiple people through time.
Alejandro Newell, Peiyun Hu, Lahav Lipson, Stephan R. Richter, Vladlen Koltun
ICLR2
2025 Lidar Panoptic Segmentation in an Open World
Anirudh Srinivasan Chakravarthy, Meghana Reddy Ganesina, Peiyun Hu, Laura Leal-Taixé, Shu Kong, Deva Ramanan, Aljosa Osep
Int. J. Comput. Vis.3
2023 Point Cloud Forecasting as a Proxy for 4D Occupancy Forecasting
abstract
Predicting how the world can evolve in the future is crucial for motion planning in autonomous systems. Classical methods are limited because they rely on costly human annotations in the form of semantic class labels, bounding boxes, and tracks or HD maps of cities to plan their motion - and thus are difficult to scale to large unlabeled datasets. One promising self-supervised task is 3D point cloud forecasting [11, 18–20] from unannotated LiDAR sequences. We show that this task requires algorithms to implicitly capture (1) sensor extrinsics (i.e., the egomotion of the autonomous vehicle), (2) sensor intrinsics (i.e., the sampling pattern specific to the particular LiDAR sensor), and (3) the shape and motion of other objects in the scene. But autonomous systems should make predictions about the world and not their sensors! To this end, we factor out (1) and (2) by recasting the task as one of spacetime (4D) occupancy forecasting. But because it is expensive to obtain ground-truth 4D occupancy, we “render” point cloud data from 4D occupancy predictions given sensor extrinsics and intrinsics, allowing one to train and test occupancy algorithms with unannotated LiDAR sequences. This also allows one to evaluate and compare point cloud forecasting algorithms across diverse datasets, sensors, and vehicles.
Tarasha Khurana, Peiyun Hu, David Held, Deva Ramanan
CVPR2
2022 Differentiable Raycasting for Self-Supervised Occupancy Forecasting
Tarasha Khurana, Peiyun Hu, Achal Dave, Jason Ziglar, David Held, Deva Ramanan
ECCV (38)2
2021 Safe Local Motion Planning With Self-Supervised Freespace Forecasting
abstract
Safe local motion planning for autonomous driving in dynamic environments requires forecasting how the scene evolves. Practical autonomy stacks adopt a semantic object-centric representation of a dynamic scene and build object detection, tracking, and prediction modules to solve forecasting. However, training these modules comes at an enormous human cost of manually annotated objects across frames. In this work, we explore future freespace as an alternative representation to support motion planning. Our key intuition is that it is important to avoid straying into occupied space regardless of what is occupying it. Importantly, computing ground-truth future freespace is annotation-free. First, we explore freespace forecasting as a self-supervised learning task. We then demonstrate how to use forecasted freespace to identify collision-prone plans from off-the-shelf motion planners. Finally, we propose future freespace as an additional source of annotation-free supervision. We demonstrate how to integrate such supervision into the learning-based planners. Experimental results on nuScenes and CARLA suggest both approaches lead to a significant reduction in collision rates.1
Peiyun Hu, Aaron Huang, John M. Dolan, David Held, Deva Ramanan
CVPR1
2020 What You See is What You Get: Exploiting Visibility for 3D Object Detection
abstract
Recent advances in 3D sensing have created unique challenges for computer vision. One fundamental challenge is finding a good representation for 3D sensor data. Most popular representations (such as PointNet) are proposed in the context of processing truly 3D data (e.g. points sampled from mesh models), ignoring the fact that 3D sensored data such as a LiDAR sweep is in fact 2.5D. We argue that representing 2.5D data as collections of (x,y,z) points fundamentally destroys hidden information about freespace. In this paper, we demonstrate such knowledge can be efficiently recovered through 3D raycasting and readily incorporated into batch-based gradient learning. We describe a simple approach to augmenting voxel-based networks with visibility: we add a voxelized visibility map as an additional input stream. In addition, we show that visibility can be combined with two crucial modifications common to state-of-the-art 3D detectors: synthetic data augmentation of virtual objects and temporal aggregation of LiDAR sweeps over multiple time frames. On the NuScenes 3D detection benchmark, we show that, by adding an additional stream for visibility input, we can significantly improve the overall detection accuracy of a state-of-the-art 3D detector.
Peiyun Hu, Jason Ziglar, David Held, Deva Ramanan
CVPR1
2020 Active Perception Using Light Curtains for Autonomous Driving
Siddharth Ancha, Yaadhav Raaj, Peiyun Hu, Srinivasa G. Narasimhan, David Held
ECCV (5)3
2019 Active Learning with Partial Feedback
Peiyun Hu, Zachary C. Lipton, Anima Anandkumar, Deva Ramanan
ICLR (Poster)1
2019 Inferring Distributions Over Depth from a Single Image
abstract
When building a geometric scene understanding system for autonomous vehicles, it is crucial to know when the system might fail. Most contemporary approaches cast the problem as depth regression, whose output is a depth value for each pixel. Such approaches cannot diagnose when failures might occur. One attractive alternative is a deep Bayesian network, which captures uncertainty in both model parameters and ambiguous sensor measurements. However, estimating uncertainties is often slow and the distributions are often limited to be uni-modal. In this paper, we recast the continuous problem of depth regression as discrete binary classification, whose output is an un-normalized distribution over possible depths for each pixel. Such output allows one to reliably and efficiently capture multi-modal depth distributions in ambiguous cases, such as depth discontinuities and reflective surfaces. Results on standard benchmarks show that our method produces accurate depth predictions and significantly better uncertainty estimations than prior art while running near real-time. Finally, by making use of uncertainties of the predicted distribution, we significantly reduce streak-like artifacts and improves accuracy as well as memory efficiency in 3D map reconstruction. Video and code can be found on the project website1.
Gengshan Yang, Peiyun Hu, Deva Ramanan
IROS2
2017 Finding Tiny Faces
abstract
Though tremendous strides have been made in object recognition, one of the remaining open challenges is detecting small objects. We explore three aspects of the problem in the context of finding small faces: the role of scale invariance, image resolution, and contextual reasoning. While most recognition approaches aim to be scale-invariant, the cues for recognizing a 3px tall face are fundamentally different than those for recognizing a 300px tall face. We take a different approach and train separate detectors for different scales. To maintain efficiency, detectors are trained in a multi-task fashion: they make use of features extracted from multiple layers of single (deep) feature hierarchy. While training detectors for large objects is straightforward, the crucial challenge remains training detectors for small objects. We show that context is crucial, and define templates that make use of massively-large receptive fields (where 99% of the template extends beyond the object of interest). Finally, we explore the role of scale in pre-trained deep networks, providing ways to extrapolate networks tuned for limited scales to rather extreme ranges. We demonstrate state-of-the-art results on massively-benchmarked face datasets (FDDB and WIDER FACE). In particular, when compared to prior art on WIDER FACE, our results reduce error by a factor of 2 (our models produce an AP of 82% while prior art ranges from 29-64%).
Peiyun Hu, Deva Ramanan
CVPR1
2017 Unconstrained Face Detection and Open-Set Face Recognition Challenge
abstract
Face detection and recognition benchmarks have shifted toward more difficult environments. The challenge presented in this paper addresses the next step in the direction of automatic detection and identification of people from outdoor surveillance cameras. While face detection has shown remarkable success in images collected from the web, surveillance cameras include more diverse occlusions, poses, weather conditions and image blur. Although face verification or closed-set face identification have surpassed human capabilities on some datasets, open-set identification is much more complex as it needs to reject both unknown identities and false accepts from the face detector. We show that unconstrained face detection can approach high detection rates albeit with moderate false accept rates. By contrast, open-set face recognition is currently weak and requires much more attention.
Manuel Günther, Peiyun Hu, Christian Herrmann 0001, Chi-Ho Chan, Min Jiang 0003, Shufan Yang, Akshay Raj Dhamija, Deva Ramanan, Jürgen Beyerer, Josef Kittler, Mohamad Al Jazaery, Mohammad Iqbal Nouyed, Guodong Guo, Cezary Stankiewicz, Terrance E. Boult
IJCB2
2016 Bottom-Up and Top-Down Reasoning with Hierarchical Rectified Gaussians
abstract
Convolutional neural nets (CNNs) have demonstrated remarkable performance in recent history. Such approaches tend to work in a "unidirectional" bottom-up feed-forward fashion. However, practical experience and biological evidence tells us that feedback plays a crucial role, particularly for detailed spatial understanding tasks. This work explores "bidirectional" architectures that also reason with top-down feedback: neural units are influenced by both lower and higher-level units. We do so by treating units as rectified latent variables in a quadratic energy function, which can be seen as a hierarchical Rectified Gaussian model (RGs) [39]. We show that RGs can be optimized with a quadratic program (QP), that can in turn be optimized with a recurrent neural network (with rectified linear units). This allows RGs to be trained with GPU-optimized gradient descent. From a theoretical perspective, RGs help establish a connection between CNNs and hierarchical probabilistic models. From a practical perspective, RGs are well suited for detailed spatial tasks that can benefit from top-down reasoning. We illustrate them on the challenging task of keypoint localization under occlusions, where local bottom-up evidence may be misleading. We demonstrate state-of-the-art results on challenging benchmarks.
Peiyun Hu, Deva Ramanan
CVPR1
2012 Can we understand van gogh's mood?: learning to infer affects from images in social networks
abstract
Can we understand van Gogh's mood from his artworks? For many years, people have tried to capture van Gogh's affects from his artworks so as to understand the essential meaning behind the images and catch on why van Gogh created these works. In this paper, we study the problem of inferring affects from images in social networks. In particular, we aim to answer: What are the fundamental features that reflect the affects of the authors in images? How the social network information can be leveraged to help detect these affects? We propose a semi-supervised framework to formulate the problem into a factor graph model. Experiments on 20,000 random-download Flickr images show that our method can achieve a precision of 49% with a recall of 24% on inferring authors'affects into 16 categories. Finally, we demonstrate the effectiveness of the proposed method on automatically understanding van Gogh's Mood from his artworks, and inferring the trend of public affects around special event.
Jia Jia 0001, Sen Wu 0001, Xiaohui Wang 0004, Peiyun Hu, Lianhong Cai, Jie Tang 0001
ACM Multimedia4
2012 Understanding the emotional impact of images
abstract
No abstract available.
Xiaohui Wang 0004, Jia Jia 0001, Peiyun Hu, Sen Wu 0001, Jie Tang 0001, Lianhong Cai
ACM Multimedia3