Guanhang Wu

dblp:187/8081 · DBLP profile ↗
← Back
7ranked-venue papers
0as first author
2since 2021 · last 2022
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Video understanding and tracking · 27% Transfer learning and domain adaptation · 21% Image recognition and object detection · 17%
Interdisciplinary, comprehensive, and emerging computing
2 papers
Smart cities and intelligent transportation · 62% Environmental and earth informatics · 38%

Topics — the 13 heaviest of 18, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Video understanding and tracking
action recognition
0.612022
Learning from Weakly-Labeled Web Videos via Exploring Sub-concepts · AAAI 2022
Machine learning › Transfer learning and domain adaptation
domain shift
0.612022
The Auto Arborist Dataset: A Large-Scale Benchmark for Multiview Urban Forest Monitoring Under Domain Shift · CVPR 2022
Machine learning › Representation and self-supervised learning
pre-training
0.612022
Learning from Weakly-Labeled Web Videos via Exploring Sub-concepts · AAAI 2022
Computer vision › Video understanding and tracking › action recognition › action recognition under limited supervision
weakly supervised action recognition
0.612022
Learning from Weakly-Labeled Web Videos via Exploring Sub-concepts · AAAI 2022
Computer vision › Video understanding and tracking
temporal context
0.412020
Context R-CNN: Long Term Temporal Context for Per-Camera Object Detection · CVPR 2020
Machine learning › Trustworthy machine learning › robustness › adversarial robustness
adversarial training
0.312018
Adversarial Multiple Source Domain Adaptation · NeurIPS 2018
Machine learning › Transfer learning and domain adaptation
domain adaptation
0.312018
Adversarial Multiple Source Domain Adaptation · NeurIPS 2018
Machine learning › Representation and self-supervised learning › representation learning › invariant representation learning
domain-invariant representation
0.312018
Adversarial Multiple Source Domain Adaptation · NeurIPS 2018
Machine learning › Transfer learning and domain adaptation › domain adaptation
multi-source domain adaptation
0.312018
Adversarial Multiple Source Domain Adaptation · NeurIPS 2018
Computer vision › Image recognition and object detection › object counting
crowd counting
0.312017
FCN-rLSTM: Deep Spatio-Temporal Neural Networks for Vehicle Counting in City Cameras · ICCV 2017
Computer vision › Image recognition and object detection › object counting
vehicle counting
0.312017
FCN-rLSTM: Deep Spatio-Temporal Neural Networks for Vehicle Counting in City Cameras · ICCV 2017
Smart cities and intelligent transportation › mobility data analysis
traffic analytics
0.312017
Understanding Traffic Density from Large-Scale Web Camera Data · CVPR 2017
Machine learning › Generative modeling
generative adversarial network
0.112018
Adversarial Multiple Source Domain Adaptation · NeurIPS 2018

Methods — techniques the papers use, named apart from their topics

street-level imagery · 1.1multiview imagery · 1.1aerial imagery · 1.1teacher-student learning · 0.6pseudo-labeling · 0.6knowledge distillation · 0.6memory bank · 0.4attention-based temporal aggregation · 0.43d convolution · 0.4adversarial neural network · 0.3
YearPublicationVenuePosition
2022 Learning from Weakly-Labeled Web Videos via Exploring Sub-concepts
abstract
Learning visual knowledge from massive weakly-labeled web videos has attracted growing research interests thanks to the large corpus of easily accessible video data on the Internet. However, for video action recognition, the action of interest might only exist in arbitrary clips of untrimmed web videos, resulting in high label noises in the temporal space. To address this challenge, we introduce a new method for pre-training video action recognition models using queried web videos. Instead of trying to filter out potential noises, we propose to provide fine-grained supervision signals by defining the concept of Sub-Pseudo Label (SPL). Specifically, SPL spans out a new set of meaningful "middle ground" label space constructed by extrapolating the original weak labels during video querying and the prior knowledge distilled from a teacher model. Consequently, SPL provides enriched supervision for video models to learn better representations and improves data utilization efficiency of untrimmed videos. We validate the effectiveness of our method on four video action recognition datasets and a weakly-labeled image dataset. Experiments show that SPL outperforms several existing pre-training strategies and the learned representations lead to competitive results on several benchmarks.
Guanhang Wu, Xuehan Xiong, Chen-Yu Lee, Zhichao Lu, Yun Fu 0001, Tomas Pfister
AAAI3
2022 The Auto Arborist Dataset: A Large-Scale Benchmark for Multiview Urban Forest Monitoring Under Domain Shift
abstract
Generalization to novel domains is a fundamental chal-lenge for computer vision. Near-perfect accuracy on bench-marks is common, but these models do not work as expected when deployed outside of the training distribution. To build computer vision systems that truly solve real-world prob-lems at global scale, we need benchmarks that fully capture real-world complexity, including geographic domain shift, long-tailed distributions, and data noise. We propose urban forest monitoring as an ideal testbed for studying and improving upon these computer vision challenges, while working towards filling a crucial environ-mental and societal need. Urban forests provide significant benefits to urban societies. However, planning and main-taining these forests is expensive. One particularly costly aspect of urban forest management is monitoring the ex-isting trees in a city: e.g., tracking tree locations, species, and health. Monitoring efforts are currently based on tree censuses built by human experts, costing cities millions of dollars per census and thus collected infrequently. Previous investigations into automating urban forest monitoring focused on small datasets from single cities, covering only common categories. To address these short-comings, we introduce a new large-scale dataset that joins public tree censuses from 23 cities with a large collection of street level and aerial imagery. Our Auto Arborist dataset contains over 2.5M trees and 344 genera and is >2 or-ders of magnitude larger than the closest dataset in the literature. We introduce baseline results on our dataset across modalities as well as metrics for the detailed analy-sis of generalization with respect to geographic distribution shifts, vital for such a system to be deployed at-scale.
Sara Beery, Guanhang Wu, Trevor Edwards, Filip Pavetic, Bo Majewski, Shreyasee Mukherjee, Stanley Chan, John Morgan, Vivek Rathod, Jonathan Huang
CVPR2
2020 Context R-CNN: Long Term Temporal Context for Per-Camera Object Detection
abstract
In static monitoring cameras, useful contextual information can stretch far beyond the few seconds typical video understanding models might see: subjects may exhibit similar behavior over multiple days, and background objects remain static. Due to power and storage constraints, sampling frequencies are low, often no faster than one frame per second, and sometimes are irregular due to the use of a motion trigger. In order to perform well in this setting, models must be robust to irregular sampling rates. In this paper we propose a method that leverages temporal context from the unlabeled frames of a novel camera to improve performance at that camera. Specifically, we propose an attention-based approach that allows our model, Context R-CNN, to index into a long term memory bank constructed on a per-camera basis and aggregate contextual features from other frames to boost object detection performance on the current frame. We apply Context R-CNN to two settings: (1) species detection using camera traps, and (2) vehicle detection in traffic cameras, showing in both settings that Context R-CNN leads to performance gains over strong baselines. Moreover, we show that increasing the contextual time horizon leads to improved results. When applied to camera trap data from the Snapshot Serengeti dataset, Context R-CNN with context from up to a month of images outperforms a single-frame baseline by 17.9% mAP, and outperforms S3D (a 3d convolution based baseline) by 11.2% mAP.
Sara Beery, Guanhang Wu, Vivek Rathod, Ronny Votel, Jonathan Huang
CVPR2
2018 Adversarial Multiple Source Domain Adaptation
abstract
While domain adaptation has been actively researched, most algorithms focus on the single-source-single-target adaptation setting. In this paper we propose new generalization bounds and algorithms under both classification and regression settings for unsupervised multiple source domain adaptation. Our theoretical analysis naturally leads to an efficient learning strategy using adversarial neural networks: we show how to interpret it as learning feature representations that are invariant to the multiple domain shifts while still being discriminative for the learning task. To this end, we propose multisource domain adversarial networks (MDAN) that approach domain adaptation by optimizing task-adaptive generalization bounds. To demonstrate the effectiveness of MDAN, we conduct extensive experiments showing superior adaptation performance on both classification and regression problems: sentiment analysis, digit classification, and vehicle counting.
Han Zhao 0002, Shanghang Zhang, Guanhang Wu, José M. F. Moura, João Paulo Costeira, Geoffrey J. Gordon
NeurIPS3
2017 Understanding Traffic Density from Large-Scale Web Camera Data
Shanghang Zhang, Guanhang Wu, João Paulo Costeira, José M. F. Moura
CVPR2
2017 FCN-rLSTM: Deep Spatio-Temporal Neural Networks for Vehicle Counting in City Cameras
abstract
In this paper, we develop deep spatio-temporal neural networks to sequentially count vehicles from low quality videos captured by city cameras (citycams). Citycam videos have low resolution, low frame rate, high occlusion and large perspective, making most existing methods lose their efficacy. To overcome limitations of existing methods and incorporate the temporal information of traffic video, we design a novel FCN-rLSTM network to jointly estimate vehicle density and vehicle count by connecting fully convolutional neural networks (FCN) with long short term memory networks (LSTM) in a residual learning fashion. Such design leverages the strengths of FCN for pixel-level prediction and the strengths of LSTM for learning complex temporal dynamics. The residual learning connection reformulates the vehicle count regression as learning residual functions with reference to the sum of densities in each frame, which significantly accelerates the training of networks. To preserve feature map resolution, we propose a Hyper-Atrous combination to integrate atrous convolution in FCN and combine feature maps of different convolution layers. FCN-rLSTM enables refined feature representation and a novel end-to-end trainable mapping from pixels to vehicle count. We extensively evaluated the proposed method on different counting tasks with three datasets, with experimental results demonstrating their effectiveness and robustness. In particular, FCN-rLSTM reduces the mean absolute error (MAE) from 5.31 to 4.21 on TRANCOS; and reduces the MAE from 2.74 to 1.53 on WebCamT. Training process is accelerated by 5 times on average.
Shanghang Zhang, Guanhang Wu, João Paulo Costeira, José M. F. Moura
ICCV2
2016 Scalable Vision System for Mouse Homecage Ethology
Ghadi Salem, Jonathan Krynitsky, Brett Kirkland, Eugene Lin, Aaron Chan, Simeon Anfinrud, Sarah Anderson, Marcial Garmendia-Cedillos, Rhamy Belayachi, Juan Alonso-Cruz, Joshua Yu, Anthony Iano-Fletcher, George Dold, Tom Talbot, Alexxai V. Kravitz, James B. Mitchell, Guanhang Wu, John U. Dennis, Monson H. Hayes III, Kristin Branson, Thomas Pohida
ACIVS17