EDBT 2026 Demo / reviewers in the wild / expert
Guanhang Wu
dblp:187/8081
· DBLP profile ↗
7ranked-venue papers
0as first author
2since 2021 · last 2022
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Video understanding and tracking · 27% Transfer learning and domain adaptation · 21% Image recognition and object detection · 17% | |
| Interdisciplinary, comprehensive, and emerging computing
2 papers |
Smart cities and intelligent transportation · 62% Environmental and earth informatics · 38% |
Topics — the 13 heaviest of 18, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Video understanding and tracking
action recognition |
0.6 | 1 | 2022 | Learning from Weakly-Labeled Web Videos via Exploring Sub-concepts · AAAI 2022 |
Machine learning › Transfer learning and domain adaptation
domain shift |
0.6 | 1 | 2022 | The Auto Arborist Dataset: A Large-Scale Benchmark for Multiview Urban Forest Monitoring Under Domain Shift · CVPR 2022 |
Machine learning › Representation and self-supervised learning
pre-training |
0.6 | 1 | 2022 | Learning from Weakly-Labeled Web Videos via Exploring Sub-concepts · AAAI 2022 |
Computer vision › Video understanding and tracking › action recognition › action recognition under limited supervision
weakly supervised action recognition |
0.6 | 1 | 2022 | Learning from Weakly-Labeled Web Videos via Exploring Sub-concepts · AAAI 2022 |
Computer vision › Video understanding and tracking
temporal context |
0.4 | 1 | 2020 | Context R-CNN: Long Term Temporal Context for Per-Camera Object Detection · CVPR 2020 |
Machine learning › Trustworthy machine learning › robustness › adversarial robustness
adversarial training |
0.3 | 1 | 2018 | Adversarial Multiple Source Domain Adaptation · NeurIPS 2018 |
Machine learning › Transfer learning and domain adaptation
domain adaptation |
0.3 | 1 | 2018 | Adversarial Multiple Source Domain Adaptation · NeurIPS 2018 |
Machine learning › Representation and self-supervised learning › representation learning › invariant representation learning
domain-invariant representation |
0.3 | 1 | 2018 | Adversarial Multiple Source Domain Adaptation · NeurIPS 2018 |
Machine learning › Transfer learning and domain adaptation › domain adaptation
multi-source domain adaptation |
0.3 | 1 | 2018 | Adversarial Multiple Source Domain Adaptation · NeurIPS 2018 |
Computer vision › Image recognition and object detection › object counting
crowd counting |
0.3 | 1 | 2017 | FCN-rLSTM: Deep Spatio-Temporal Neural Networks for Vehicle Counting in City Cameras · ICCV 2017 |
Computer vision › Image recognition and object detection › object counting
vehicle counting |
0.3 | 1 | 2017 | FCN-rLSTM: Deep Spatio-Temporal Neural Networks for Vehicle Counting in City Cameras · ICCV 2017 |
Smart cities and intelligent transportation › mobility data analysis
traffic analytics |
0.3 | 1 | 2017 | Understanding Traffic Density from Large-Scale Web Camera Data · CVPR 2017 |
Machine learning › Generative modeling
generative adversarial network |
0.1 | 1 | 2018 | Adversarial Multiple Source Domain Adaptation · NeurIPS 2018 |
Methods — techniques the papers use, named apart from their topics
street-level imagery · 1.1multiview imagery · 1.1aerial imagery · 1.1teacher-student learning · 0.6pseudo-labeling · 0.6knowledge distillation · 0.6memory bank · 0.4attention-based temporal aggregation · 0.43d convolution · 0.4adversarial neural network · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | Learning from Weakly-Labeled Web Videos via Exploring Sub-conceptsabstractLearning visual knowledge from massive weakly-labeled web videos has attracted growing research interests thanks to the large corpus of easily accessible video data on the Internet. However, for video action recognition, the action of interest might only exist in arbitrary clips of untrimmed web videos, resulting in high label noises in the temporal space. To address this challenge, we introduce a new method for pre-training video action recognition models using queried web videos. Instead of trying to filter out potential noises, we propose to provide fine-grained supervision signals by defining the concept of Sub-Pseudo Label (SPL). Specifically, SPL spans out a new set of meaningful "middle ground" label space constructed by extrapolating the original weak labels during video querying and the prior knowledge distilled from a teacher model. Consequently, SPL provides enriched supervision for video models to learn better representations and improves data utilization efficiency of untrimmed videos. We validate the effectiveness of our method on four video action recognition datasets and a weakly-labeled image dataset. Experiments show that SPL outperforms several existing pre-training strategies and the learned representations lead to competitive results on several benchmarks. Guanhang Wu, Xuehan Xiong, Chen-Yu Lee, Zhichao Lu, Yun Fu 0001, Tomas Pfister |
AAAI | 3 |
| 2022 | The Auto Arborist Dataset: A Large-Scale Benchmark for Multiview Urban Forest Monitoring Under Domain ShiftabstractGeneralization to novel domains is a fundamental chal-lenge for computer vision. Near-perfect accuracy on bench-marks is common, but these models do not work as expected when deployed outside of the training distribution. To build computer vision systems that truly solve real-world prob-lems at global scale, we need benchmarks that fully capture real-world complexity, including geographic domain shift, long-tailed distributions, and data noise. We propose urban forest monitoring as an ideal testbed for studying and improving upon these computer vision challenges, while working towards filling a crucial environ-mental and societal need. Urban forests provide significant benefits to urban societies. However, planning and main-taining these forests is expensive. One particularly costly aspect of urban forest management is monitoring the ex-isting trees in a city: e.g., tracking tree locations, species, and health. Monitoring efforts are currently based on tree censuses built by human experts, costing cities millions of dollars per census and thus collected infrequently. Previous investigations into automating urban forest monitoring focused on small datasets from single cities, covering only common categories. To address these short-comings, we introduce a new large-scale dataset that joins public tree censuses from 23 cities with a large collection of street level and aerial imagery. Our Auto Arborist dataset contains over 2.5M trees and 344 genera and is >2 or-ders of magnitude larger than the closest dataset in the literature. We introduce baseline results on our dataset across modalities as well as metrics for the detailed analy-sis of generalization with respect to geographic distribution shifts, vital for such a system to be deployed at-scale. Sara Beery, Guanhang Wu, Trevor Edwards, Filip Pavetic, Bo Majewski, Shreyasee Mukherjee, Stanley Chan, John Morgan, Vivek Rathod, Jonathan Huang |
CVPR | 2 |
| 2020 | Context R-CNN: Long Term Temporal Context for Per-Camera Object DetectionabstractIn static monitoring cameras, useful contextual information can stretch far beyond the few seconds typical video understanding models might see: subjects may exhibit similar behavior over multiple days, and background objects remain static. Due to power and storage constraints, sampling frequencies are low, often no faster than one frame per second, and sometimes are irregular due to the use of a motion trigger. In order to perform well in this setting, models must be robust to irregular sampling rates. In this paper we propose a method that leverages temporal context from the unlabeled frames of a novel camera to improve performance at that camera. Specifically, we propose an attention-based approach that allows our model, Context R-CNN, to index into a long term memory bank constructed on a per-camera basis and aggregate contextual features from other frames to boost object detection performance on the current frame. We apply Context R-CNN to two settings: (1) species detection using camera traps, and (2) vehicle detection in traffic cameras, showing in both settings that Context R-CNN leads to performance gains over strong baselines. Moreover, we show that increasing the contextual time horizon leads to improved results. When applied to camera trap data from the Snapshot Serengeti dataset, Context R-CNN with context from up to a month of images outperforms a single-frame baseline by 17.9% mAP, and outperforms S3D (a 3d convolution based baseline) by 11.2% mAP. Sara Beery, Guanhang Wu, Vivek Rathod, Ronny Votel, Jonathan Huang |
CVPR | 2 |
| 2018 | Adversarial Multiple Source Domain AdaptationabstractWhile domain adaptation has been actively researched, most algorithms focus on the single-source-single-target adaptation setting. In this paper we propose new generalization bounds and algorithms under both classification and regression settings for unsupervised multiple source domain adaptation. Our theoretical analysis naturally leads to an efficient learning strategy using adversarial neural networks: we show how to interpret it as learning feature representations that are invariant to the multiple domain shifts while still being discriminative for the learning task. To this end, we propose multisource domain adversarial networks (MDAN) that approach domain adaptation by optimizing task-adaptive generalization bounds. To demonstrate the effectiveness of MDAN, we conduct extensive experiments showing superior adaptation performance on both classification and regression problems: sentiment analysis, digit classification, and vehicle counting. Han Zhao 0002, Shanghang Zhang, Guanhang Wu, José M. F. Moura, João Paulo Costeira, Geoffrey J. Gordon |
NeurIPS | 3 |
| 2017 | Understanding Traffic Density from Large-Scale Web Camera Data
Shanghang Zhang, Guanhang Wu, João Paulo Costeira, José M. F. Moura |
CVPR | 2 |
| 2017 | FCN-rLSTM: Deep Spatio-Temporal Neural Networks for Vehicle Counting in City CamerasabstractIn this paper, we develop deep spatio-temporal neural networks to sequentially count vehicles from low quality videos captured by city cameras (citycams). Citycam videos have low resolution, low frame rate, high occlusion and large perspective, making most existing methods lose their efficacy. To overcome limitations of existing methods and incorporate the temporal information of traffic video, we design a novel FCN-rLSTM network to jointly estimate vehicle density and vehicle count by connecting fully convolutional neural networks (FCN) with long short term memory networks (LSTM) in a residual learning fashion. Such design leverages the strengths of FCN for pixel-level prediction and the strengths of LSTM for learning complex temporal dynamics. The residual learning connection reformulates the vehicle count regression as learning residual functions with reference to the sum of densities in each frame, which significantly accelerates the training of networks. To preserve feature map resolution, we propose a Hyper-Atrous combination to integrate atrous convolution in FCN and combine feature maps of different convolution layers. FCN-rLSTM enables refined feature representation and a novel end-to-end trainable mapping from pixels to vehicle count. We extensively evaluated the proposed method on different counting tasks with three datasets, with experimental results demonstrating their effectiveness and robustness. In particular, FCN-rLSTM reduces the mean absolute error (MAE) from 5.31 to 4.21 on TRANCOS; and reduces the MAE from 2.74 to 1.53 on WebCamT. Training process is accelerated by 5 times on average. Shanghang Zhang, Guanhang Wu, João Paulo Costeira, José M. F. Moura |
ICCV | 2 |
| 2016 | Scalable Vision System for Mouse Homecage Ethology
Ghadi Salem, Jonathan Krynitsky, Brett Kirkland, Eugene Lin, Aaron Chan, Simeon Anfinrud, Sarah Anderson, Marcial Garmendia-Cedillos, Rhamy Belayachi, Juan Alonso-Cruz, Joshua Yu, Anthony Iano-Fletcher, George Dold, Tom Talbot, Alexxai V. Kravitz, James B. Mitchell, Guanhang Wu, John U. Dennis, Monson H. Hayes III, Kristin Branson, Thomas Pohida |
ACIVS | 17 |