Andrew G. Howard

dblp:139/0987 · also Andrew Howard 0002 · DBLP profile ↗
← Back
17ranked-venue papers
4as first author
5since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 16 · 4 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 1 first-author · 4 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
13 papers
Efficient and distributed learning · 52% Segmentation and scene understanding · 16% Image recognition and object detection · 15%
Computer architecture, parallel and distributed computing, and storage systems
2 papers
Hardware accelerators and domain-specific architectures · 71% Performance modeling and evaluation · 29%

Topics — the 29 heaviest of 33, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning
model compression
2.462024
MobileNetV4: Universal Models for the Mobile Ecosystem · ECCV (40) 2024
Searching for MobileNetV3 · ICCV 2019
NetAdapt: Platform-Aware Neural Network Adaptation for Mobile Applications · ECCV (10) 2018
Machine learning › Efficient and distributed learning › automated machine learning
neural architecture search
0.822019
Searching for MobileNetV3 · ICCV 2019
MnasNet: Platform-Aware Neural Architecture Search for Mobile · CVPR 2019
Hardware accelerators and domain-specific architectures › machine learning accelerator › DNN inference
low-precision DNN inference
0.812024
PikeLPN: Mitigating Overlooked Inefficiencies of Low-Precision Neural Networks · CVPR 2024
Hardware accelerators and domain-specific architectures
machine learning accelerator
0.812024
PikeLPN: Mitigating Overlooked Inefficiencies of Low-Precision Neural Networks · CVPR 2024
Computer vision › Segmentation and scene understanding › image segmentation
efficient segmentation
0.712023
ReMaX: Relaxing for Better Training on Efficient Panoptic Segmentation · NeurIPS 2023
Machine learning › Efficient and distributed learning
inference efficiency
0.712023
ReMaX: Relaxing for Better Training on Efficient Panoptic Segmentation · NeurIPS 2023
Computer vision › Segmentation and scene understanding
panoptic segmentation
0.712023
ReMaX: Relaxing for Better Training on Efficient Panoptic Segmentation · NeurIPS 2023
Computer vision › Image recognition and object detection
object localization
0.612022
On Label Granularity and Object Localization · ECCV (10) 2022
Machine learning › Efficient and distributed learning › model compression
quantization
0.622024
Quantization and Training of Neural Networks for Efficient Integer-Arithmetic-Only Inference · CVPR 2018
PikeLPN: Mitigating Overlooked Inefficiencies of Low-Precision Neural Networks · CVPR 2024
Computer vision › Segmentation and scene understanding
semantic segmentation
0.522019
Searching for MobileNetV3 · ICCV 2019
MobileNetV2: Inverted Residuals and Linear Bottlenecks · CVPR 2018
Computer vision › Image recognition and object detection › image classification
mobile image classification
0.412019
Searching for MobileNetV3 · ICCV 2019
Machine learning › Learning paradigms
multi-task learning
0.412019
K for the Price of 1: Parameter-efficient Multi-task and Transfer Learning · ICLR (Poster) 2019
Machine learning › Efficient and distributed learning
parameter-efficient fine-tuning
0.412019
K for the Price of 1: Parameter-efficient Multi-task and Transfer Learning · ICLR (Poster) 2019
Machine learning › Transfer learning and domain adaptation
parameter-efficient transfer learning
0.412019
K for the Price of 1: Parameter-efficient Multi-task and Transfer Learning · ICLR (Poster) 2019
Machine learning › Deep learning architectures and training › efficient deep learning
efficient neural network architecture
0.312018
MobileNetV2: Inverted Residuals and Linear Bottlenecks · CVPR 2018
Computer vision › Image recognition and object detection › image classification
fine-grained image classification
0.312018
Large Scale Fine-Grained Categorization and Domain-Specific Transfer Learning · CVPR 2018
Machine learning › Efficient and distributed learning › model compression
lightweight neural network
0.312018
MobileNetV2: Inverted Residuals and Linear Bottlenecks · CVPR 2018
Machine learning › Efficient and distributed learning › model compression
pruning
0.312018
NetAdapt: Platform-Aware Neural Network Adaptation for Mobile Applications · ECCV (10) 2018
Machine learning › Efficient and distributed learning › model compression › quantization
quantization-aware training
0.312018
Quantization and Training of Neural Networks for Efficient Integer-Arithmetic-Only Inference · CVPR 2018
Hardware accelerators and domain-specific architectures › efficient inference
efficient neural network inference
0.312018
NetAdapt: Platform-Aware Neural Network Adaptation for Mobile Applications · ECCV (10) 2018
Computer vision › Image recognition and object detection › efficient visual recognition
mobile image recognition
0.212024
MobileNetV4: Universal Models for the Mobile Ecosystem · ECCV (40) 2024
Machine learning › Deep learning architectures and training › transformer
masked transformer
0.212023
ReMaX: Relaxing for Better Training on Efficient Panoptic Segmentation · NeurIPS 2023
Machine learning › Deep learning architectures and training
convolutional neural network
0.112019
MnasNet: Platform-Aware Neural Architecture Search for Mobile · CVPR 2019
Computer vision › Image recognition and object detection
image classification
0.112018
Quantization and Training of Neural Networks for Efficient Integer-Arithmetic-Only Inference · CVPR 2018
Computer vision › Image recognition and object detection
object detection
0.112018
MobileNetV2: Inverted Residuals and Linear Bottlenecks · CVPR 2018
Machine learning › Learning theory › margin-based learning
large margin classification
0.112007
Learning Monotonic Transformations for Classification · NIPS 2007
Machine learning › Optimization for machine learning › convex relaxation
semidefinite programming relaxation
0.112007
Learning Monotonic Transformations for Classification · NIPS 2007
Machine learning › Kernel, tree and ensemble methods
kernel methods
0.012004
Probability Product Kernels · J. Mach. Learn. Res. 2004
Machine learning › Representation and self-supervised learning
similarity measure
0.012004
Probability Product Kernels · J. Mach. Learn. Res. 2004

Methods — techniques the papers use, named apart from their topics

quantization · 1.5double quantization · 1.5distribution-heterogeneous quantization · 1.5neural architecture search · 0.8training-time relaxation · 0.7loss relaxation · 0.7platform-aware search · 0.4latency measurement · 0.4hardware-aware NAS · 0.4factorized hierarchical search space · 0.4latency-aware optimization · 0.3automated network adaptation · 0.3
YearPublicationVenuePosition
2024 PikeLPN: Mitigating Overlooked Inefficiencies of Low-Precision Neural Networks
abstract
Low-precision quantization is recognized for its efficacy in neural network optimization. Our analysis reveals that non-quantized elementwise operations which are prevalent in layers such as parameterized activation functions, batch normalization, and quantization scaling dominate the inference cost of low-precision models. These non-quantized elementwise operations are commonly overlooked in SOTA efficiency metrics such as Arithmetic Computation Effort (ACE) [46]. In this paper, we propose ACEv2- an extended version of ACE which offers a better alignment with the inference cost of quantized models and their energy consumption on ML hardware. Moreover, we introduce PikeLPN11Pike is a slim fast fish, LPN stands for Low-Precision Network., a model that addresses these efficiency issues by applying quantization to both elementwise operations and multiply-accumulate operations. In particular, we present a novel quantization technique for batch normalization layers named QuantNorm which allows for quantizing the batch normalization parameters without compromising the model performance. Additionally, we propose applying Double Quantization where the quantization scaling parameters are quantized. Furthermore, we recognize and resolve the issue of distribution mismatch in Separable Convolution layers by introducing Distribution-Heterogeneous Quantization which enables quantizing them to low-precision. PikeLPN achieves Pareto-optimality in efficiency-accuracy trade-off with up to 3× efficiency improvement compared to SOTA low-precision models.
Marina Neseem, Conor McCullough, Randy Hsin, Chas Leichner, Shan Li 0001, In Suk Chong, Andrew G. Howard, Lukasz Lew, Sherief Reda, Ville-Mikko Rautio, Daniele Moro
CVPR7
2024 MobileNetV4: Universal Models for the Mobile Ecosystem
Danfeng Qin, Chas Leichner, Manolis Delakis, Marco Fornoni, Shixin Luo, Colby R. Banbury, Chengxi Ye, Berkin Akin, Vaibhav Aggarwal, Tenghui Zhu, Daniele Moro, Andrew G. Howard
ECCV (40)14
2023 ReMaX: Relaxing for Better Training on Efficient Panoptic Segmentation
abstract
This paper presents a new mechanism to facilitate the training of mask transformers for efficient panoptic segmentation, democratizing its deployment. We observe that due to the high complexity in the training objective of panoptic segmentation, it will inevitably lead to much higher penalization on false positive. Such unbalanced loss makes the training process of the end-to-end mask-transformer based architectures difficult, especially for efficient models. In this paper, we present ReMaX that adds relaxation to mask predictions and class predictions during the training phase for panoptic segmentation. We demonstrate that via these simple relaxation techniques during training, our model can be consistently improved by a clear margin without any extra computational cost on inference. By combining our method with efficient backbones like MobileNetV3-Small, our method achieves new state-of-the-art results for efficient panoptic segmentation on COCO, ADE20K and Cityscapes. Code and pre-trained checkpoints will be available at https://github.com/google-research/deeplab2.
Shuyang Sun, Andrew G. Howard, Qihang Yu, Philip Torr 0001, Liang-Chieh Chen
NeurIPS3
2022 On Label Granularity and Object Localization
Elijah Cole, Kimberly Wilber, Grant Van Horn, Marco Fornoni, Pietro Perona, Serge J. Belongie, Andrew G. Howard, Oisin Mac Aodha
ECCV (10)8
2021 Multi-path Neural Networks for On-device Multi-domain Visual Classification
abstract
Learning multiple domains/tasks with a single model is important for improving data efficiency and lowering inference cost for numerous vision tasks, especially on resource-constrained mobile devices. However, hand-crafting a multi-domain/task model can be both tedious and challenging. This paper proposes a novel approach to automatically learn a multi-path network for multi-domain visual classification on mobile devices. The proposed multi-path network is learned from neural architecture search by applying one reinforcement learning controller for each domain to select the best path in the super-network created from a MobileNetV3-like search space. An adaptive balanced domain prioritization algorithm is proposed to balance optimizing the joint model on multiple domains simultaneously. The determined multi-path model selectively shares parameters across domains in shared nodes while keeping domain-specific parameters within non-shared nodes in individual domain paths. This approach effectively reduces the total number of parameters and FLOPS, encouraging positive knowledge transfer while mitigating negative interference across domains. Extensive evaluations on the Visual Decathlon dataset demonstrate that the proposed multi-path model achieves state-of-the-art performance in terms of accuracy, model size, and FLOPS against other approaches using MobileNetV3-like architectures. Furthermore, the proposed method improves average accuracy over learning single-domain models individually, and reduces the total number of parameters and FLOPS by 78% and 32% respectively, compared to the approach that simply bundles single-domain models for multi-domain learning.
Qifei Wang, Junjie Ke, Joshua Greaves, Grace Chu, Gabriel Bender, Luciano Sbaiz, Alec Go, Andrew G. Howard, Ming-Hsuan Yang 0001, Jeff Gilbert, Peyman Milanfar, Feng Yang 0008
WACV8
2020 SpotPatch: Parameter-Efficient Transfer Learning for Mobile Object Detection
Keren Ye, Adriana Kovashka, Mark Sandler 0002, Menglong Zhu, Andrew G. Howard, Marco Fornoni
ACCV (6)5
2019 MnasNet: Platform-Aware Neural Architecture Search for Mobile
abstract
Designing convolutional neural networks (CNN) for mobile devices is challenging because mobile models need to be small and fast, yet still accurate. Although significant efforts have been dedicated to design and improve mobile CNNs on all dimensions, it is very difficult to manually balance these trade-offs when there are so many architectural possibilities to consider. In this paper, we propose an automated mobile neural architecture search (MNAS) approach, which explicitly incorporate model latency into the main objective so that the search can identify a model that achieves a good trade-off between accuracy and latency. Unlike previous work, where latency is considered via another, often inaccurate proxy (e.g., FLOPS), our approach directly measures real-world inference latency by executing the model on mobile phones. To further strike the right balance between flexibility and search space size, we propose a novel factorized hierarchical search space that encourages layer diversity throughout the network. Experimental results show that our approach consistently outperforms state-of-the-art mobile CNN models across multiple vision tasks. On the ImageNet classification task, our MnasNet achieves 75.2% top-1 accuracy with 78ms latency on a Pixel phone, which is 1.8× faster than MobileNetV2 with 0.5% higher accuracy and 2.3× faster than NASNet with 1.2% higher accuracy. Our MnasNet also achieves better mAP quality than MobileNets for COCO object detection. Code is at https://github.com/tensorflow/tpu/tree/master/models/official/mnasnet.
Mingxing Tan, Bo Chen 0019, Ruoming Pang, Vijay Vasudevan, Mark Sandler 0002, Andrew G. Howard, Quoc V. Le
CVPR6
2019 Searching for MobileNetV3
abstract
We present the next generation of MobileNets based on a combination of complementary search techniques as well as a novel architecture design. MobileNetV3 is tuned to mobile phone CPUs through a combination of hardware-aware network architecture search (NAS) complemented by the NetAdapt algorithm and then subsequently improved through novel architecture advances. This paper starts the exploration of how automated search algorithms and network design can work together to harness complementary approaches improving the overall state of the art. Through this process we create two new MobileNet models for release: MobileNetV3-Large and MobileNetV3-Small which are targeted for high and low resource use cases. These models are then adapted and applied to the tasks of object detection and semantic segmentation. For the task of semantic segmentation (or any dense pixel prediction), we propose a new efficient segmentation decoder Lite Reduced Atrous Spatial Pyramid Pooling (LR-ASPP). We achieve new state of the art results for mobile classification, detection and segmentation. MobileNetV3-Large is 3.2% more accurate on ImageNet classification while reducing latency by 20% compared to MobileNetV2. MobileNetV3-Small is 6.6% more accurate compared to a MobileNetV2 model with comparable latency. MobileNetV3-Large detection is over 25% faster at roughly the same accuracy as MobileNetV2 on COCO detection. MobileNetV3-Large LRASPP is 34% faster than MobileNetV2 R-ASPP at similar accuracy for Cityscapes segmentation.
Andrew G. Howard, Ruoming Pang, Hartwig Adam, Quoc V. Le, Mark Sandler 0002, Bo Chen 0019, Liang-Chieh Chen, Mingxing Tan, Grace Chu, Vijay Vasudevan, Yukun Zhu
ICCV1
2019 K for the Price of 1: Parameter-efficient Multi-task and Transfer Learning
Pramod Kaushik Mudrakarta, Mark Sandler 0002, Andrey Zhmoginov, Andrew G. Howard
ICLR (Poster)4
2018 Large Scale Fine-Grained Categorization and Domain-Specific Transfer Learning
abstract
Transferring the knowledge learned from large scale datasets (e.g., ImageNet) via fine-tuning offers an effective solution for domain-specific fine-grained visual categorization (FGVC) tasks (e.g., recognizing bird species or car make & model). In such scenarios, data annotation often calls for specialized domain knowledge and thus is difficult to scale. In this work, we first tackle a problem in large scale FGVC. Our method won first place in iNaturalist 2017 large scale species classification challenge. Central to the success of our approach is a training scheme that uses higher image resolution and deals with the long-tailed distribution of training data. Next, we study transfer learning via fine-tuning from large scale datasets to small scale, domain-specific FGVC datasets. We propose a measure to estimate domain similarity via Earth Mover's Distance and demonstrate that transfer learning benefits from pre-training on a source domain that is similar to the target domain by this measure. Our proposed transfer learning outperforms ImageNet pre-training and obtains state-of-the-art results on multiple commonly used FGVC datasets.
Yin Cui, Yang Song 0009, Chen Sun 0002, Andrew G. Howard, Serge J. Belongie
CVPR4
2018 Quantization and Training of Neural Networks for Efficient Integer-Arithmetic-Only Inference
abstract
The rising popularity of intelligent mobile devices and the daunting computational cost of deep learning-based models call for efficient and accurate on-device inference schemes. We propose a quantization scheme that allows inference to be carried out using integer-only arithmetic, which can be implemented more efficiently than floating point inference on commonly available integer-only hardware. We also co-design a training procedure to preserve end-to-end model accuracy post quantization. As a result, the proposed quantization scheme improves the tradeoff between accuracy and on-device latency. The improvements are significant even on MobileNets, a model family known for run-time efficiency, and are demonstrated in ImageNet classification and COCO detection on popular CPUs.
Benoit Jacob, Skirmantas Kligys, Bo Chen 0019, Menglong Zhu, Matthew Tang, Andrew G. Howard, Hartwig Adam, Dmitry Kalenichenko
CVPR6
2018 MobileNetV2: Inverted Residuals and Linear Bottlenecks
abstract
In this paper we describe a new mobile architecture, MobileNetV2, that improves the state of the art performance of mobile models on multiple tasks and benchmarks as well as across a spectrum of different model sizes. We also describe efficient ways of applying these mobile models to object detection in a novel framework we call SSDLite. Additionally, we demonstrate how to build mobile semantic segmentation models through a reduced form of DeepLabv3 which we call Mobile DeepLabv3. is based on an inverted residual structure where the shortcut connections are between the thin bottleneck layers. The intermediate expansion layer uses lightweight depthwise convolutions to filter features as a source of non-linearity. Additionally, we find that it is important to remove non-linearities in the narrow layers in order to maintain representational power. We demonstrate that this improves performance and provide an intuition that led to this design. Finally, our approach allows decoupling of the input/output domains from the expressiveness of the transformation, which provides a convenient framework for further analysis. We measure our performance on ImageNet [1] classification, COCO object detection [2], VOC image segmentation [3]. We evaluate the trade-offs between accuracy, and number of operations measured by multiply-adds (MAdd), as well as actual latency, and the number of parameters.
Mark Sandler 0002, Andrew G. Howard, Menglong Zhu, Andrey Zhmoginov, Liang-Chieh Chen
CVPR2
2018 NetAdapt: Platform-Aware Neural Network Adaptation for Mobile Applications
Tien-Ju Yang, Andrew G. Howard, Bo Chen 0019, Alec Go, Mark Sandler 0002, Vivienne Sze, Hartwig Adam
ECCV (10)2
2009 Transformation Learning Via Kernel Alignment
abstract
This article proposes an algorithm to automatically learn useful transformations of data to improve accuracy in supervised classification tasks. These transformations take the form of a mixture of base transformations and are learned by maximizing the kernel alignment criterion. Because the proposed optimization is nonconvex, a semidefinite relaxation is derived to find an approximate global solution. This new convex algorithm learns kernels made up of a matrix mixture of transformations. This formulation yields a simpler optimization while achieving comparable or improved accuracies to previous transformation learning algorithms based on maximizing the margin. Remarkably, the new optimization problem does not slow down with the availability of additional data allowing it to scale to large datasets. One application of this method is learning monotonic transformations constructed from a base set of truncated ramp functions. These monotonic transformations permit a nonlinear filtering of the input to the classifier. The effectiveness of the method is demonstrated on synthetic data, text data and image data.
Andrew G. Howard, Tony Jebara
ICMLA1
2007 Learning Monotonic Transformations for Classification
abstract
A discriminative method is proposed for learning monotonic transforma- tions of the training data while jointly estimating a large-margin classi(cid:12)er. In many domains such as document classi(cid:12)cation, image histogram classi(cid:12)- cation and gene microarray experiments, (cid:12)xed monotonic transformations can be useful as a preprocessing step. However, most classi(cid:12)ers only explore these transformations through manual trial and error or via prior domain knowledge. The proposed method learns monotonic transformations auto- matically while training a large-margin classi(cid:12)er without any prior knowl- edge of the domain. A monotonic piecewise linear function is learned which transforms data for subsequent processing by a linear hyperplane classi(cid:12)er. Two algorithmic implementations of the method are formalized. The (cid:12)rst solves a convergent alternating sequence of quadratic and linear programs until it obtains a locally optimal solution. An improved algorithm is then derived using a convex semide(cid:12)nite relaxation that overcomes initializa- tion issues in the greedy optimization problem. The e(cid:11)ectiveness of these learned transformations on synthetic problems, text data and image data is demonstrated.
Andrew G. Howard, Tony Jebara
NIPS1
2004 Dynamical Systems Trees
Andrew G. Howard, Tony Jebara
UAI1
2004 Probability Product Kernels
Tony Jebara, Risi Kondor, Andrew G. Howard
J. Mach. Learn. Res.3