Wongun Choi

dblp:04/8611 · DBLP profile ↗
← Back
22ranked-venue papers
9as first author
3since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 18 · 9 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 18 · 6 first-author · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
18 papers
Video understanding and tracking · 35% Image recognition and object detection · 16% 3D vision · 14%
Theoretical computer science
1 paper
Graph algorithms and graph theory · 100%

Topics — the 30 heaviest of 37, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Video understanding and tracking
multi-object tracking
1.462021
Learning a Proposal Classifier for Multiple Object Tracking · CVPR 2021
Deep Network Flow for Multi-object Tracking · CVPR 2017
Near-Online Multi-target Tracking with Aggregated Local Flow Descriptor · ICCV 2015
Robotics › Autonomous driving
trajectory prediction
0.722019
Multi-Agent Tensor Fusion for Contextual Trajectory Prediction · CVPR 2019
DESIRE: Distant Future Prediction in Dynamic Scenes with Interacting Agents · CVPR 2017
Computer vision › 3D vision › multi-view geometry
camera geometry
0.612022
Ray3D: ray-based 3D human pose estimation for monocular absolute 3D localization · CVPR 2022
Computer vision › Face, body and person analysis
human pose estimation
0.612022
Ray3D: ray-based 3D human pose estimation for monocular absolute 3D localization · CVPR 2022
Computer vision › Image recognition and object detection
object detection
0.522017
Learning Efficient Object Detection Models with Knowledge Distillation · NIPS 2017
Exploit All the Layers: Fast and Accurate CNN Object Detector with Scale Dependent Pooling and Cascaded Rejection Classifiers · CVPR 2016
Computer vision › Video understanding and tracking › multi-object tracking
data association
0.522017
Deep Network Flow for Multi-object Tracking · CVPR 2017
Near-Online Multi-target Tracking with Aggregated Local Flow Descriptor · ICCV 2015
Machine learning › Graph learning › graph neural network
graph convolutional network
0.512021
Learning a Proposal Classifier for Multiple Object Tracking · CVPR 2021
Computer vision › Image recognition and object detection › object detection › proposal-based detection
proposal classification
0.512021
Learning a Proposal Classifier for Multiple Object Tracking · CVPR 2021
Computer vision › Segmentation and scene understanding › scene understanding
indoor scene understanding
0.422015
Indoor Scene Understanding with Geometric and Semantic Contexts · Int. J. Comput. Vis. 2015
Understanding Indoor Scenes Using 3D Geometric Phrases · CVPR 2013
Knowledge, reasoning and agents › Multi-agent systems › agent interaction › multi-agent interaction
multi-agent interaction modeling
0.412019
Multi-Agent Tensor Fusion for Contextual Trajectory Prediction · CVPR 2019
Computer vision › Video understanding and tracking › activity recognition
group activity recognition
0.432014
Understanding Collective Activitiesof People from Videos · IEEE Trans. Pattern Anal. Mach. Intell. 2014
Learning context for collective activity recognition · CVPR 2011
A Unified Framework for Multi-target Tracking and Collective Activity Recognition · ECCV (4) 2012
Computer vision › Video understanding and tracking › multi-object tracking
multi-person tracking
0.422014
Understanding Collective Activitiesof People from Videos · IEEE Trans. Pattern Anal. Mach. Intell. 2014
A General Framework for Tracking Multiple People from a Moving Camera · IEEE Trans. Pattern Anal. Mach. Intell. 2013
Computer vision › Image recognition and object detection › object detection
efficient object detection
0.312017
Learning Efficient Object Detection Models with Knowledge Distillation · NIPS 2017
Machine learning › Efficient and distributed learning › model compression
knowledge distillation
0.312017
Learning Efficient Object Detection Models with Knowledge Distillation · NIPS 2017
Machine learning › Efficient and distributed learning
model compression
0.312017
Learning Efficient Object Detection Models with Knowledge Distillation · NIPS 2017
Graph algorithms and graph theory
graph algorithms
0.312017
Deep Network Flow for Multi-object Tracking · CVPR 2017
Graph algorithms and graph theory › graph algorithms
network flow
0.312017
Deep Network Flow for Multi-object Tracking · CVPR 2017
Computer vision › Image recognition and object detection › object detection › deep learning object detection
CNN-based detection
0.212016
Exploit All the Layers: Fast and Accurate CNN Object Detector with Scale Dependent Pooling and Cascaded Rejection Classifiers · CVPR 2016
Computer vision › 3D vision
3d object detection
0.212015
Data-driven 3D Voxel Patterns for object category recognition · CVPR 2015
Computer vision › 3D vision › spatial understanding
geometric context
0.212015
Indoor Scene Understanding with Geometric and Semantic Contexts · Int. J. Comput. Vis. 2015
Computer vision › Segmentation and scene understanding
instance segmentation
0.212015
Data-driven 3D Voxel Patterns for object category recognition · CVPR 2015
Computer vision › 3D vision
object pose estimation
0.212015
Data-driven 3D Voxel Patterns for object category recognition · CVPR 2015
Computer vision › Video understanding and tracking › object tracking
online tracking
0.212015
Near-Online Multi-target Tracking with Aggregated Local Flow Descriptor · ICCV 2015
Computer vision › Video understanding and tracking › crowd analysis
group detection
0.212014
Discovering Groups of People in Images · ECCV (4) 2014
Computer vision › Video understanding and tracking › multi-object tracking
tracklet association
0.212014
Understanding Collective Activitiesof People from Videos · IEEE Trans. Pattern Anal. Mach. Intell. 2014
Computer vision › Video understanding and tracking
activity recognition
0.222012
Learning context for collective activity recognition · CVPR 2011
A Unified Framework for Multi-target Tracking and Collective Activity Recognition · ECCV (4) 2012
Computer vision › 3D vision › motion estimation
ego-motion estimation
0.212013
A General Framework for Tracking Multiple People from a Moving Camera · IEEE Trans. Pattern Anal. Mach. Intell. 2013
Computer vision › Segmentation and scene understanding
scene understanding
0.212013
Understanding Indoor Scenes Using 3D Geometric Phrases · CVPR 2013
Machine learning › Time series and sequential data › time series modeling
probabilistic forecasting
0.112019
Multi-Agent Tensor Fusion for Contextual Trajectory Prediction · CVPR 2019
Machine learning › Deep learning architectures and training
teacher-student framework
0.112017
Learning Efficient Object Detection Models with Knowledge Distillation · NIPS 2017

Methods — techniques the papers use, named apart from their topics

ray-based representation · 0.6graph convolutional network · 0.5graph clustering · 0.5recurrent decoding · 0.4convolutional fusion · 0.4adversarial loss · 0.4smoothed network flow · 0.3scene context fusion · 0.3inverse optimal control · 0.3conditional variational autoencoder · 0.3backpropagation through optimization · 0.3RNN encoder-decoder · 0.3
YearPublicationVenuePosition
2024 3SD: Self-Supervised Saliency Detection With No Labels
abstract
We present a conceptually simple self-supervised method for saliency detection. Our method generates and uses pseudo-ground truth labels for training. The generated pseudo-GT labels don’t require any kind of human annotations (e.g., pixel-wise labels or weak labels like scribbles). Recent works show that features extracted from classification tasks provide important saliency cues like structure and semantic information of salient objects in the image. Our method, called 3SD, exploits this idea by adding a branch for a self-supervised classification task in parallel with salient object detection, to obtain class activation maps (CAM maps). These CAM maps along with the edges of the input image are used to generate the pseudo-GT saliency maps to train our 3SD network. Specifically, we propose a contrastive learning-based training on multiple image patches for the classification task. We show the multi-patch classification with contrastive loss improves the quality of the CAM maps compared to naive classification on the entire image. Experiments on six benchmark datasets demonstrate that without any labels, our 3SD method outperforms all existing weakly supervised and unsupervised methods, and its performance is on par with the fully-supervised methods.
Rajeev Yasarla, Renliang Weng, Wongun Choi, Vishal M. Patel, Amir Sadeghian
WACV3
2022 Ray3D: ray-based 3D human pose estimation for monocular absolute 3D localization
abstract
In this paper, we propose a novel monocular ray-based 3D (Ray3D) absolute human pose estimation with calibrated camera. Accurate and generalizable absolute 3D human pose estimation from monocular 2D pose input is an ill-posed problem. To address this challenge, we convert the input from pixel space to 3D normalized rays. This conversion makes our approach robust to camera intrinsic parameter changes. To deal with the in-the-wild camera extrinsic parameter variations, Ray3D explicitly takes the camera extrinsic parameters as an input and jointly models the distribution between the 3D pose rays and camera extrinsic parameters. This novel network design is the key to the outstanding generalizability of Ray3D approach. To have a comprehensive understanding of how the camera intrinsic and extrinsic parameter variations affect the accuracy of absolute 3D key-point localization, we conduct in-depth systematic experiments on three single person 3D benchmarks as well as one synthetic benchmark. These experiments demonstrate that our method significantly outperforms existing state-of-the-art models. Our code and the synthetic dataset are available at https://github.com/YxZhxn/Ray3D.
Fenghai Li, Renliang Weng, Wongun Choi
CVPR4
2021 Learning a Proposal Classifier for Multiple Object Tracking
abstract
The recent trend in multiple object tracking (MOT) is heading towards leveraging deep learning to boost the tracking performance. However, it is not trivial to solve the data-association problem in an end-to-end fashion. In this paper, we propose a novel proposal-based learnable framework, which models MOT as a proposal generation, proposal scoring and trajectory inference paradigm on an affinity graph. This framework is similar to the two-stage object detector Faster RCNN, and can solve the MOT problem in a data-driven way. For proposal generation, we propose an iterative graph clustering method to reduce the computational cost while maintaining the quality of the generated proposals. For proposal scoring, we deploy a trainable graph-convolutional-network (GCN) to learn the structural patterns of the generated proposals and rank them according to the estimated quality scores. For trajectory inference, a simple deoverlapping strategy is adopted to generate tracking output while complying with the constraints that no detection can be assigned to more than one track. We experimentally demonstrate that the proposed method achieves a clear performance improvement in both MOTA and IDF1 with respect to previous state-of-the-art on two public benchmarks. Our code is available at https://github.com/daip13/LPC_MOT.git.
Renliang Weng, Wongun Choi, Changshui Zhang, Zhangping He
CVPR3
2019 Multi-Agent Tensor Fusion for Contextual Trajectory Prediction
abstract
Accurate prediction of others' trajectories is essential for autonomous driving. Trajectory prediction is challenging because it requires reasoning about agents' past movements, social interactions among varying numbers and kinds of agents, constraints from the scene context, and the stochasticity of human behavior. Our approach models these interactions and constraints jointly within a novel Multi-Agent Tensor Fusion (MATF) network. Specifically, the model encodes multiple agents' past trajectories and the scene context into a Multi-Agent Tensor, then applies convolutional fusion to capture multiagent interactions while retaining the spatial structure of agents and the scene context. The model decodes recurrently to multiple agents' future trajectories, using adversarial loss to learn stochastic predictions. Experiments on both highway driving and pedestrian crowd datasets show that the model achieves state-of-the-art prediction accuracy.
Tianyang Zhao 0004, Mathew Monfort, Wongun Choi, Chris L. Baker, Yibiao Zhao, Yizhou Wang 0001, Ying Nian Wu
CVPR4
2019 Memory Warps for Long-Term Online Video Representations and Anticipation
abstract
We propose a novel memory-based online video representation that is efficient, accurate and predictive. This is in contrast to prior works that often rely on computationally heavy 3D convolutions, ignore motion when aligning features over time, or operate in an off-line mode to utilize future frames. In particular, our memory (i) holds the feature representation, (ii) is spatially warped over time to compensate for observer and scene motions, (iii) can carry long-term information, and (iv) enables predicting feature representations in future frames. By exploring a variant that operates at multiple temporal scales, we efficiently learn across even longer time horizons. We apply our online framework to object detection in videos, obtaining a large 2.3 times speed-up and losing only 0.9% mAP on ImageNet-VID dataset, compared to prior works that even use future frames. Finally, we demonstrate the predictive property of our representation in two novel detection setups, where features are propagated over time to (i) significantly enhance a real-time detector by more than 10% mAP in a multi-threaded online setup and to (ii) anticipate objects in future frames.
Wongun Choi, Samuel Schulter, Manmohan Krishna Chandraker
WACV2
2017 DESIRE: Distant Future Prediction in Dynamic Scenes with Interacting Agents
abstract
We introduce a Deep Stochastic IOC RNN Encoder-decoder framework, DESIRE, for the task of future predictions of multiple interacting agents in dynamic scenes. DESIRE effectively predicts future locations of objects in multiple scenes by 1) accounting for the multi-modal nature of the future prediction (i.e., given the same context, future may vary), 2) foreseeing the potential future outcomes and make a strategic prediction based on that, and 3) reasoning not only from the past motion history, but also from the scene context as well as the interactions among the agents. DESIRE achieves these in a single end-to-end trainable neural network model, while being computationally efficient. The model first obtains a diverse set of hypothetical future prediction samples employing a conditional variational auto-encoder, which are ranked and refined by the following RNN scoring-regression module. Samples are scored by accounting for accumulated future rewards, which enables better long-term strategic decisions similar to IOC frameworks. An RNN scene context fusion module jointly captures past motion histories, the semantic scene context and interactions among multiple agents. A feedback mechanism iterates over the ranking and refinement to further boost the prediction accuracy. We evaluate our model on two publicly available datasets: KITTI and Stanford Drone Dataset. Our experiments show that the proposed model significantly improves the prediction accuracy compared to other baseline methods.
Namhoon Lee, Wongun Choi, Paul Vernaza, Christopher B. Choy, Philip Torr 0001, Manmohan Krishna Chandraker
CVPR2
2017 Deep Network Flow for Multi-object Tracking
abstract
Data association problems are an important component of many computer vision applications, with multi-object tracking being one of the most prominent examples. A typical approach to data association involves finding a graph matching or network flow that minimizes a sum of pairwise association costs, which are often either hand-crafted or learned as linear functions of fixed features. In this work, we demonstrate that it is possible to learn features for network-flow-based data association via backpropagation, by expressing the optimum of a smoothed network flow problem as a differentiable function of the pairwise association costs. We apply this approach to multi-object tracking with a network flow formulation. Our experiments demonstrate that we are able to successfully learn all cost functions for the association problem in an end-to-end fashion, which outperform hand-crafted costs in all settings. The integration and combination of various sources of inputs becomes easy and the cost functions can be learned entirely from data, alleviating tedious hand-designing of costs.
Samuel Schulter, Paul Vernaza, Wongun Choi, Manmohan Krishna Chandraker
CVPR3
2017 Learning Efficient Object Detection Models with Knowledge Distillation
abstract
Despite significant accuracy improvement in convolutional neural networks (CNN) based object detectors, they often require prohibitive runtimes to process an image for real-time applications. State-of-the-art models often use very deep networks with a large number of floating point operations. Efforts such as model compression learn compact models with fewer number of parameters, but with much reduced accuracy. In this work, we propose a new framework to learn compact and fast ob- ject detection networks with improved accuracy using knowledge distillation [20] and hint learning [34]. Although knowledge distillation has demonstrated excellent improvements for simpler classification setups, the complexity of detection poses new challenges in the form of regression, region proposals and less voluminous la- bels. We address this through several innovations such as a weighted cross-entropy loss to address class imbalance, a teacher bounded loss to handle the regression component and adaptation layers to better learn from intermediate teacher distribu- tions. We conduct comprehensive empirical evaluation with different distillation configurations over multiple datasets including PASCAL, KITTI, ILSVRC and MS-COCO. Our results show consistent improvement in accuracy-speed trade-offs for modern multi-class detection models.
Guobin Chen, Wongun Choi, Xiang Yu 0002, Tony X. Han, Manmohan Krishna Chandraker
NIPS2
2017 Subcategory-Aware Convolutional Neural Networks for Object Proposals and Detection
abstract
In Convolutional Neural Network (CNN)-based object detection methods, region proposal becomes a bottleneck when objects exhibit significant scale variation, occlusion or truncation. In addition, these methods mainly focus on 2D object detection and cannot estimate detailed properties of objects. In this paper, we propose subcategory-aware CNNs for object detection. We introduce a novel region proposal network that uses subcategory information to guide the proposal generating process, and a new detection network for joint detection and subcategory classification. By using subcategories related to object pose, we achieve state of-the-art performance on both detection and pose estimation on commonly used benchmarks.
Wongun Choi, Yuanqing Lin, Silvio Savarese
WACV2
2016 Exploit All the Layers: Fast and Accurate CNN Object Detector with Scale Dependent Pooling and Cascaded Rejection Classifiers
abstract
In this paper, we investigate two new strategies to detect objects accurately and efficiently using deep convolutional neural network: 1) scale-dependent pooling and 2) layerwise cascaded rejection classifiers. The scale-dependent pooling (SDP) improves detection accuracy by exploiting appropriate convolutional features depending on the scale of candidate object proposals. The cascaded rejection classifiers (CRC) effectively utilize convolutional features and eliminate negative object proposals in a cascaded manner, which greatly speeds up the detection while maintaining high accuracy. In combination of the two, our method achieves significantly better accuracy compared to other state-of-the-arts in three challenging datasets, PASCAL object detection challenge, KITTI object detection benchmark and newly collected Inner-city dataset, while being more efficient.
Wongun Choi, Yuanqing Lin
CVPR2
2016 Atomic scenes for scalable traffic scene recognition in monocular videos
abstract
The efficacy of advance warning systems (AWS) in automobiles can be significantly enhanced by semantic recognition of traffic scenes that pose a potential danger. However, the complexity of road scenes and the need for real-time solutions pose key challenges. This paper proposes a novel framework for monocular traffic scene recognition, relying on a decomposition into high-order and atomic scenes to meet those challenges. High-order scenes carry semantic meaning useful for AWS applications, while atomic scenes are easy to learn and represent elemental behaviors based on 3D localization of individual traffic participants. Atomic scenes allow our framework to be scalable, since a few of them combine to influence prediction for a wide array of high-order scenes. We propose a novel hierarchical model that captures co-occurence and mutual exclusion relationships while incorporating both low-level trajectory features and high-level scene features, with parameters learned using a structured support vector machine. We propose efficient inference that exploits the structure of our model to obtain real-time rates. Further, we demonstrate experiments in a large-scale dataset for scene recognition that consists of challenging traffic videos of inner-city scenes with ground truth annotations of scene types and object bounding boxes, as well as state-of-the-art 3D object localization outputs. Our experiments show the advantages of our approach relative to several baselines on a novel Inner-City dataset.
Chao-Yeh Chen, Wongun Choi, Manmohan Krishna Chandraker
WACV2
2015 Data-driven 3D Voxel Patterns for object category recognition
abstract
Despite the great progress achieved in recognizing objects as 2D bounding boxes in images, it is still very challenging to detect occluded objects and estimate the 3D properties of multiple objects from a single image. In this paper, we propose a novel object representation, 3D Voxel Pattern (3DVP), that jointly encodes the key properties of objects including appearance, 3D shape, viewpoint, occlusion and truncation. We discover 3DVPs in a data-driven way, and train a bank of specialized detectors for a dictionary of 3DVPs. The 3DVP detectors are capable of detecting objects with specific visibility patterns and transferring the meta-data from the 3DVPs to the detected objects, such as 2D segmentation mask, 3D pose as well as occlusion or truncation boundaries. The transferred meta-data allows us to infer the occlusion relationship among objects, which in turn provides improved object recognition results. Experiments are conducted on the KITTI detection benchmark [17] and the outdoor-scene dataset [41]. We improve state-of-the-art results on car detection and pose estimation with notable margins (6% in difficult data of KITTI). We also verify the ability of our method in accurately segmenting objects from the background and localizing them in 3D.
Wongun Choi, Yuanqing Lin, Silvio Savarese
CVPR2
2015 Near-Online Multi-target Tracking with Aggregated Local Flow Descriptor
abstract
In this paper, we tackle two key aspects of multiple target tracking problem: 1) designing an accurate affinity measure to associate detections and 2) implementing an efficient and accurate (near) online multiple target tracking algorithm. As for the first contribution, we introduce a novel Aggregated Local Flow Descriptor (ALFD) that encodes the relative motion pattern between a pair of temporally distant detections using long term interest point trajectories (IPTs). Leveraging on the IPTs, the ALFD provides a robust affinity measure for estimating the likelihood of matching detections regardless of the application scenarios. As for another contribution, we present a Near-Online Multi-target Tracking (NOMT) algorithm. The tracking problem is formulated as a data-association between targets and detections in a temporal window, that is performed repeatedly at every frame. While being efficient, NOMT achieves robustness via integrating multiple cues including ALFD metric, target dynamics, appearance similarity, and long term trajectory regularization into the model. Our ablative analysis verifies the superiority of the ALFD metric over the other conventional affinity metrics. We run a comprehensive experimental evaluation on two challenging tracking datasets, KITTI [16] and MOT [2] datasets. The NOMT method combined with ALFD metric achieves the best accuracy in both datasets with significant margins (about 10% higher MOTA) over the state-of-the-art.
Wongun Choi
ICCV1
2015 Indoor Scene Understanding with Geometric and Semantic Contexts
Wongun Choi, Yu-Wei Chao, Caroline Pantofaru, Silvio Savarese
Int. J. Comput. Vis.1
2014 Discovering Groups of People in Images
Wongun Choi, Yu-Wei Chao, Caroline Pantofaru, Silvio Savarese
ECCV (4)1
2014 Understanding Collective Activitiesof People from Videos
abstract
This paper presents a principled framework for analyzing collective activities at different levels of semantic granularity from videos. Our framework is capable of jointly tracking multiple individuals, recognizing activities performed by individuals in isolation (i.e., atomic activities such as walking or standing), recognizing the interactions between pairs of individuals (i.e., interaction activities) as well as understanding the activities of group of individuals (i.e., collective activities). A key property of our work is that it can coherently combine bottom-up information stemming from detections or fragments of tracks (or tracklets) with top-down evidence. Top-down evidence is provided by a newly proposed descriptor that captures the coherent behavior of groups of individuals in a spatial-temporal neighborhood of the sequence. Top-down evidence provides contextual information for establishing accurate associations between detections or tracklets across frames and, thus, for obtaining more robust tracking results. Bottom-up evidence percolates upwards so as to automatically infer collective activity labels. Experimental results on two challenging data sets demonstrate our theoretical claims and indicate that our model achieves enhances tracking results and the best collective classification results to date.
Wongun Choi, Silvio Savarese
IEEE Trans. Pattern Anal. Mach. Intell.1
2013 Understanding Indoor Scenes Using 3D Geometric Phrases
abstract
Visual scene understanding is a difficult problem interleaving object detection, geometric reasoning and scene classification. We present a hierarchical scene model for learning and reasoning about complex indoor scenes which is computationally tractable, can be learned from a reasonable amount of training data, and avoids oversimplification. At the core of this approach is the 3D Geometric Phrase Model which captures the semantic and geometric relationships between objects which frequently co-occur in the same 3D spatial configuration. Experiments show that this model effectively explains scene semantics, geometry and object groupings from a single image, while also improving individual object detections.
Wongun Choi, Yu-Wei Chao, Caroline Pantofaru, Silvio Savarese
CVPR1
2013 Breaking the Chain: Liberation from the Temporal Markov Assumption for Tracking Human Poses
abstract
We present an approach to multi-target tracking that has expressive potential beyond the capabilities of chain-shaped hidden Markov models, yet has significantly reduced complexity. Our framework, which we call tracking-by-selection, is similar to tracking-by-detection in that it separates the tasks of detection and tracking, but it shifts temporal reasoning from the tracking stage to the detection stage. The core feature of tracking-by-selection is that it reasons about path hypotheses that traverse the entire video instead of a chain of single-frame object hypotheses. A traditional chain-shaped tracking-by-detection model is only able to promote consistency between one frame and the next. In tracking-by-selection, path hypotheses exist across time, and encouraging long-term temporal consistency is as simple as rewarding path hypotheses with consistent image features. One additional advantage of tracking-by-selection is that it results in a dramatically simplified model that can be solved exactly. We adapt an existing tracking-by-detection model to the tracking-by-selection framework, and show improved performance on a challenging dataset.
Ryan Tokola, Wongun Choi, Silvio Savarese
ICCV2
2013 A General Framework for Tracking Multiple People from a Moving Camera
abstract
In this paper, we present a general framework for tracking multiple, possibly interacting, people from a mobile vision platform. To determine all of the trajectories robustly and in a 3D coordinate system, we estimate both the camera's ego-motion and the people's paths within a single coherent framework. The tracking problem is framed as finding the MAP solution of a posterior probability, and is solved using the reversible jump Markov chain Monte Carlo (RJ-MCMC) particle filtering method. We evaluate our system on challenging datasets taken from moving cameras, including an outdoor street scene video dataset, as well as an indoor RGB-D dataset collected in an office. Experimental evidence shows that the proposed method can robustly estimate a camera's motion from dynamic scenes and stably track people who are moving independently or interacting.
Wongun Choi, Caroline Pantofaru, Silvio Savarese
IEEE Trans. Pattern Anal. Mach. Intell.1
2012 A Unified Framework for Multi-target Tracking and Collective Activity Recognition
Wongun Choi, Silvio Savarese
ECCV (4)1
2011 Learning context for collective activity recognition
abstract
In this paper we present a framework for the recognition of collective human activities. A collective activity is defined or reinforced by the existence of coherent behavior of individuals in time and space. We call such coherent behavior `Crowd Context'. Examples of collective activities are “queuing in a line” or “talking”. Following, we propose to recognize collective activities using the crowd context and introduce a new scheme for learning it automatically. Our scheme is constructed upon a Random Forest structure which randomly samples variable volume spatio-temporal regions to pick the most discriminating attributes for classification. Unlike previous approaches, our algorithm automatically finds the optimal configuration of spatio-temporal bins, over which to sample the evidence, by randomization. This enables a methodology for modeling crowd context. We employ a 3D Markov Random Field to regularize the classification and localize collective activities in the scene. We demonstrate the flexibility and scalability of the proposed framework in a number of experiments and show that our method outperforms state-of-the art action classification techniques.
Wongun Choi, Khuram Shahid, Silvio Savarese
CVPR1
2010 Multiple Target Tracking in World Coordinate with Single, Minimally Calibrated Camera
Wongun Choi, Silvio Savarese
ECCV (4)1