Chunhui Gu

dblp:52/7659 · DBLP profile ↗
← Back
8ranked-venue papers
4as first author
0since 2021 · last 2018
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 4 first-authorArtificial intelligence and machine learning · 7 · 4 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
7 papers
Image recognition and object detection · 38% Video understanding and tracking · 29% Segmentation and scene understanding · 27%

Topics — the 17 heaviest of 19, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Video understanding and tracking
action recognition
0.312018
AVA: A Video Dataset of Spatio-Temporally Localized Atomic Visual Actions · CVPR 2018
Computer vision › Video understanding and tracking › action detection
spatio-temporal action localization
0.312018
AVA: A Video Dataset of Spatio-Temporally Localized Atomic Visual Actions · CVPR 2018
Computer vision › Image recognition and object detection
object detection
0.222012
Multi-component Models for Object Detection · ECCV (4) 2012
Recognition using regions · CVPR 2009
Computer vision › Image recognition and object detection › object detection
region-based detection
0.222012
Semantic segmentation using regions and parts · CVPR 2012
Recognition using regions · CVPR 2009
Computer vision › Segmentation and scene understanding
semantic segmentation
0.112012
Semantic segmentation using regions and parts · CVPR 2012
Computer vision › Segmentation and scene understanding › image segmentation › binary segmentation
foreground-background segmentation
0.112010
Figure-ground segmentation improves handled object recognition in egocentric video · CVPR 2010
Computer vision › Image recognition and object detection
object recognition
0.112010
Figure-ground segmentation improves handled object recognition in egocentric video · CVPR 2010
Computer vision › Segmentation and scene understanding › image segmentation
hierarchical segmentation
0.112009
Context by region ancestry · ICCV 2009
Computer vision › Segmentation and scene understanding
image segmentation
0.112009
Context by region ancestry · ICCV 2009
Computer vision › Image recognition and object detection › image classification
object classification
0.112009
Recognition using regions · CVPR 2009
Computer vision › Segmentation and scene understanding
object segmentation
0.112009
Recognition using regions · CVPR 2009
Computer vision › Segmentation and scene understanding
scene understanding
0.112009
Context by region ancestry · ICCV 2009
Computer vision › Vision and language
visual context modeling
0.112009
Context by region ancestry · ICCV 2009
Computer vision › Image recognition and object detection › object detection
component-based detection
0.012012
Multi-component Models for Object Detection · ECCV (4) 2012
Computer vision › Image recognition and object detection › object detection
part-based object detection
0.012012
Semantic segmentation using regions and parts · CVPR 2012
Computer vision › Video understanding and tracking
egocentric video
0.012010
Figure-ground segmentation improves handled object recognition in egocentric video · CVPR 2010
Computer vision › Image recognition and object detection › object detection › hough transform
hough voting
0.012009
Recognition using regions · CVPR 2009

Methods — techniques the papers use, named apart from their topics

action localization · 0.3region-based detection · 0.1pixel classification · 0.1part models · 0.1mixture model · 0.1deformable part model · 0.1optical flow · 0.1max-margin classifier · 0.1SIFT · 0.1HOG · 0.1
YearPublicationVenuePosition
2018 AVA: A Video Dataset of Spatio-Temporally Localized Atomic Visual Actions
abstract
This paper introduces a video dataset of spatio-temporally localized Atomic Visual Actions (AVA). The AVA dataset densely annotates 80 atomic visual actions in 437 15-minute video clips, where actions are localized in space and time, resulting in 1.59M action labels with multiple labels per person occurring frequently. The key characteristics of our dataset are: (1) the definition of atomic visual actions, rather than composite actions; (2) precise spatio-temporal annotations with possibly multiple annotations for each person; (3) exhaustive annotation of these atomic actions over 15-minute video clips; (4) people temporally linked across consecutive segments; and (5) using movies to gather a varied set of action representations. This departs from existing datasets for spatio-temporal action recognition, which typically provide sparse annotations for composite actions in short video clips. AVA, with its realistic scene and action complexity, exposes the intrinsic difficulty of action recognition. To benchmark this, we present a novel approach for action localization that builds upon the current state-of-the-art methods, and demonstrates better performance on JHMDB and UCF101-24 categories. While setting a new state of the art on existing datasets, the overall results on AVA are low at 15.8% mAP, underscoring the need for developing new approaches for video understanding.
Chunhui Gu, Chen Sun 0002, David A. Ross, Carl Vondrick, Caroline Pantofaru, Yeqing Li, Sudheendra Vijayanarasimhan, George Toderici, Susanna Ricco, Rahul Sukthankar, Cordelia Schmid, Jitendra Malik
CVPR1
2018 Towards A Semantic Perceptual Image Metric
abstract
We present a full reference, perceptual image metric based on VGG-16, an artificial neural network trained on object classification. We fit the metric to a new database based on 140k unique images annotated with ground truth by human raters who received minimal instruction. The resulting metric shows competitive performance on TID 2013, a database widely used to assess image quality assessments methods. More interestingly, it shows strong responses to objects potentially carrying semantic relevance such as faces and text, which we demonstrate using a visualization technique and ablation experiments. In effect, the metric appears to model a higher influence of semantic context on judgments, which we observe particularly in untrained raters. As the vast majority of users of image processing systems are unfamiliar with Image Quality Assessment (IQA) tasks, these findings may have significant impact on real-world applications of perceptual metrics.
Troy T. Chinen, Jona Ballé, Chunhui Gu, Sung Jin Hwang, Sergey Ioffe, Nicholas Johnston, Thomas K. Leung, David Minnen, Sean M. O'Malley, Charles Rosenberg 0001, George Toderici
ICIP3
2012 Semantic segmentation using regions and parts
abstract
We address the problem of segmenting and recognizing objects in real world images, focusing on challenging articulated categories such as humans and other animals. For this purpose, we propose a novel design for region-based object detectors that integrates efficiently top-down information from scanning-windows part models and global appearance cues. Our detectors produce class-specific scores for bottom-up regions, and then aggregate the votes of multiple overlapping candidates through pixel classification. We evaluate our approach on the PASCAL segmentation challenge, and report competitive performance with respect to current leading techniques. On VOC2010, our method obtains the best results in 6/20 categories and the highest performance on articulated objects.
Pablo Andrés Arbeláez, Bharath Hariharan, Chunhui Gu, Saurabh Gupta 0001, Lubomir D. Bourdev, Jitendra Malik
CVPR3
2012 Multi-component Models for Object Detection
Chunhui Gu, Pablo Andrés Arbeláez, Yuanqing Lin, Kai Yu 0001, Jitendra Malik
ECCV (4)1
2010 Figure-ground segmentation improves handled object recognition in egocentric video
abstract
Identifying handled objects, i.e. objects being manipulated by a user, is essential for recognizing the person's activities. An egocentric camera as worn on the body enjoys many advantages such as having a natural first-person view and not needing to instrument the environment. It is also a challenging setting, where background clutter is known to be a major source of problems and is difficult to handle with the camera constantly and arbitrarily moving. In this work we develop a bottom-up motion-based approach to robustly segment out foreground objects in egocentric video and show that it greatly improves object recognition accuracy. Our key insight is that egocentric video of object manipulation is a special domain and many domain-specific cues can readily help. We compute dense optical flow and fit it into multiple affine layers. We then use a max-margin classifier to combine motion with empirical knowledge of object location and background movement as well as temporal cues of support region and color appearance. We evaluate our segmentation algorithm on the large Intel Egocentric Object Recognition dataset with 42 objects and 100K frames. We show that, when combined with temporal integration, figure-ground segmentation improves the accuracy of a SIFT-based recognition system from 33% to 60%, and that of a latent-HOG system from 64% to 86%.
Xiaofeng Ren, Chunhui Gu
CVPR2
2010 Discriminative Mixture-of-Templates for Viewpoint Classification
Chunhui Gu, Xiaofeng Ren
ECCV (5)1
2009 Recognition using regions
abstract
This paper presents a unified framework for object detection, segmentation, and classification using regions. Region features are appealing in this context because: (1) they encode shape and scale information of objects naturally; (2) they are only mildly affected by background clutter. Regions have not been popular as features due to their sensitivity to segmentation errors. In this paper, we start by producing a robust bag of overlaid regions for each image using Arbeldez et al., CVPR 2009. Each region is represented by a rich set of image cues (shape, color and texture). We then learn region weights using a max-margin framework. In detection and segmentation, we apply a generalized Hough voting scheme to generate hypotheses of object locations, scales and support, followed by a verification classifier and a constrained segmenter on each hypothesis. The proposed approach significantly outperforms the state of the art on the ETHZ shape database(87.1% average detection rate compared to Ferrari et al. 's 67.2%), and achieves competitive performance on the Caltech 101 database.
Chunhui Gu, Joseph J. Lim, Pablo Andrés Arbeláez, Jitendra Malik
CVPR1
2009 Context by region ancestry
abstract
In this paper, we introduce a new approach for modeling visual context. For this purpose, we consider the leaves of a hierarchical segmentation tree as elementary units. Each leaf is described by features of its ancestral set, the regions on the path linking the leaf to the root. We construct region trees by using a high-performance segmentation method. We then learn the importance of different descriptors (e.g. color, texture, shape) of the ancestors for classification. We report competitive results on the MSRC segmentation dataset and the MIT scene dataset, showing that region ancestry efficiently encodes information about discriminative parts, objects and scenes.
Joseph J. Lim, Pablo Andrés Arbeláez, Chunhui Gu, Jitendra Malik
ICCV3