Yuxin Wu 0004

dblp:90/8974-4 · DBLP profile ↗
← Back
8ranked-venue papers
3as first author
0since 2021 · last 2020
0000-0002-7348-6609ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 3 first-authorGraphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
8 papers
Deep learning architectures and training · 24% Representation and self-supervised learning · 14% Segmentation and scene understanding · 14%
Computer graphics and multimedia
1 paper
Rendering · 100%
Human-computer interaction and pervasive computing
1 paper
Games and playful interaction · 100%

Topics — the 16 heaviest of 18, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Deep learning architectures and training › normalization
group normalization
0.822020
Group Normalization · Int. J. Comput. Vis. 2020
Group Normalization · ECCV (13) 2018
Machine learning › Representation and self-supervised learning
contrastive learning
0.412020
Momentum Contrast for Unsupervised Visual Representation Learning · CVPR 2020
Computer vision › Segmentation and scene understanding
image segmentation
0.412020
PointRend: Image Segmentation As Rendering · CVPR 2020
Machine learning › Representation and self-supervised learning › contrastive learning › self-supervised contrastive learning
momentum contrastive learning
0.412020
Momentum Contrast for Unsupervised Visual Representation Learning · CVPR 2020
Machine learning › Deep learning architectures and training › normalization
normalization layers
0.412020
Group Normalization · Int. J. Comput. Vis. 2020
Computer vision › Segmentation and scene understanding › interactive segmentation
point-based segmentation
0.412020
PointRend: Image Segmentation As Rendering · CVPR 2020
Machine learning › Learning paradigms › unsupervised learning
unsupervised visual representation learning
0.412020
Momentum Contrast for Unsupervised Visual Representation Learning · CVPR 2020
Rendering
point-based rendering
0.412020
PointRend: Image Segmentation As Rendering · CVPR 2020
Machine learning › Trustworthy machine learning › robustness
adversarial robustness
0.412019
Feature Denoising for Improving Adversarial Robustness · CVPR 2019
Machine learning › Trustworthy machine learning
feature denoising
0.412019
Feature Denoising for Improving Adversarial Robustness · CVPR 2019
Robotics › Robot navigation and mapping › visual navigation
semantic visual navigation
0.412019
Bayesian Relational Memory for Semantic Visual Navigation · ICCV 2019
Machine learning › Deep learning architectures and training
normalization
0.312018
Group Normalization · ECCV (13) 2018
Machine learning › Reinforcement learning
actor-critic methods
0.312017
Training Agent for First-Person Shooter Game with Actor-Critic Curriculum Learning · ICLR (Poster) 2017
Machine learning › Learning paradigms
curriculum learning
0.312017
Training Agent for First-Person Shooter Game with Actor-Critic Curriculum Learning · ICLR (Poster) 2017
Knowledge, reasoning and agents › Multi-agent systems › multi-agent environments
real-time strategy games
0.312017
ELF: An Extensive, Lightweight and Flexible Research Platform for Real-time Strategy Games · NIPS 2017
Games and playful interaction
game AI
0.312017
Training Agent for First-Person Shooter Game with Actor-Critic Curriculum Learning · ICLR (Poster) 2017

Methods — techniques the papers use, named apart from their topics

point-based rendering · 0.9iterative subdivision · 0.9curriculum learning · 0.9queue-based dictionary · 0.4moving-averaged encoder · 0.4dictionary look-up · 0.4non-local means · 0.4graph-based memory · 0.4bayesian inference · 0.4adversarial training · 0.4actor-critic · 0.3
YearPublicationVenuePosition
2020 Momentum Contrast for Unsupervised Visual Representation Learning
abstract
We present Momentum Contrast (MoCo) for unsupervised visual representation learning. From a perspective on contrastive learning as dictionary look-up, we build a dynamic dictionary with a queue and a moving-averaged encoder. This enables building a large and consistent dictionary on-the-fly that facilitates contrastive unsupervised learning. MoCo provides competitive results under the common linear protocol on ImageNet classification. More importantly, the representations learned by MoCo transfer well to downstream tasks. MoCo can outperform its supervised pre-training counterpart in 7 detection/segmentation tasks on PASCAL VOC, COCO, and other datasets, sometimes surpassing it by large margins. This suggests that the gap between unsupervised and supervised representation learning has been largely closed in many vision tasks.
Kaiming He, Haoqi Fan 0001, Yuxin Wu 0004, Saining Xie, Ross B. Girshick
CVPR3
2020 PointRend: Image Segmentation As Rendering
abstract
We present a new method for efficient high-quality image segmentation of objects and scenes. By analogizing classical computer graphics methods for efficient rendering with over- and undersampling challenges faced in pixel labeling tasks, we develop a unique perspective of image segmentation as a rendering problem. From this vantage, we present the PointRend (Point-based Rendering) neural network module: a module that performs point-based segmentation predictions at adaptively selected locations based on an iterative subdivision algorithm. PointRend can be flexibly applied to both instance and semantic segmentation tasks by building on top of existing state-of-the-art models. While many concrete implementations of the general idea are possible, we show that a simple design already achieves excellent results. Qualitatively, PointRend outputs crisp object boundaries in regions that are over-smoothed by previous methods. Quantitatively, PointRend yields significant gains on COCO and Cityscapes, for both instance and semantic segmentation. PointRend's efficiency enables output resolutions that are otherwise impractical in terms of memory or computation compared to existing approaches. Code has been made available at https://github.com/facebookresearch/detectron2/tree/master/projects/PointRend.
Alexander Kirillov, Yuxin Wu 0004, Kaiming He, Ross B. Girshick
CVPR2
2020 Group Normalization
Yuxin Wu 0004, Kaiming He
Int. J. Comput. Vis.1
2019 Feature Denoising for Improving Adversarial Robustness
abstract
Adversarial attacks to image classification systems present challenges to convolutional networks and opportunities for understanding them. This study suggests that adversarial perturbations on images lead to noise in the features constructed by these networks. Motivated by this observation, we develop new network architectures that increase adversarial robustness by performing feature denoising. Specifically, our networks contain blocks that denoise the features using non-local means or other filters; the entire networks are trained end-to-end. When combined with adversarial training, our feature denoising networks substantially improve the state-of-the-art in adversarial robustness in both white-box and black-box attack settings. On ImageNet, under 10-iteration PGD white-box attacks where prior art has 27.9% accuracy, our method achieves 55.7%; even under extreme 2000-iteration PGD white-box attacks, our method secures 42.6% accuracy. Our method was ranked first in Competition on Adversarial Attacks and Defenses (CAAD) 2018 --- it achieved 50.6% classification accuracy on a secret, ImageNet-like test dataset against 48 unknown attackers, surpassing the runner-up approach by ~10%. Code is available at https://github.com/facebookresearch/ImageNet-Adversarial-Training.
Cihang Xie, Yuxin Wu 0004, Laurens van der Maaten, Alan L. Yuille, Kaiming He
CVPR2
2019 Bayesian Relational Memory for Semantic Visual Navigation
abstract
We introduce a new memory architecture, Bayesian Relational Memory (BRM), to improve the generalization ability for semantic visual navigation agents in unseen environments, where an agent is given a semantic target to navigate towards. BRM takes the form of a probabilistic relation graph over semantic entities (e.g., room types), which allows (1) capturing the layout prior from training environments, i.e., prior knowledge, (2) estimating posterior layout at test time, i.e., memory update, and (3) efficient planning for navigation, altogether. We develop a BRM agent consisting of a BRM module for producing sub-goals and a goal-conditioned locomotion module for control. When testing in unseen environments, the BRM agent outperforms baselines that do not explicitly utilize the probabilistic relational memory structure.
Yi Wu 0013, Yuxin Wu 0004, Aviv Tamar, Stuart Russell 0001, Georgia Gkioxari, Yuandong Tian
ICCV2
2018 Group Normalization
Yuxin Wu 0004, Kaiming He
ECCV (13)1
2017 Training Agent for First-Person Shooter Game with Actor-Critic Curriculum Learning
Yuxin Wu 0004, Yuandong Tian
ICLR (Poster)1
2017 ELF: An Extensive, Lightweight and Flexible Research Platform for Real-time Strategy Games
abstract
In this paper, we propose ELF, an Extensive, Lightweight and Flexible platform for fundamental reinforcement learning research. Using ELF, we implement a highly customizable real-time strategy (RTS) engine with three game environments (Mini-RTS, Capture the Flag and Tower Defense). Mini-RTS, as a miniature version of StarCraft, captures key game dynamics and runs at 165K frame-per-second (FPS) on a laptop. When coupled with modern reinforcement learning methods, the system can train a full-game bot against built-in AIs end-to-end in one day with 6 CPUs and 1 GPU. In addition, our platform is flexible in terms of environment-agent communication topologies, choices of RL methods, changes in game parameters, and can host existing C/C++-based game environments like ALE. Using ELF, we thoroughly explore training parameters and show that a network with Leaky ReLU and Batch Normalization coupled with long-horizon training and progressive curriculum beats the rule-based built-in AI more than 70% of the time in the full game of Mini-RTS. Strong performance is also achieved on the other two games. In game replays, we show our agents learn interesting strategies. ELF, along with its RL platform, is open-sourced at https://github.com/facebookresearch/ELF.
Yuandong Tian, Qucheng Gong, Wenling Shang, Yuxin Wu 0004, C. Lawrence Zitnick
NIPS4