Lukasz Lew

dblp:52/5311 · DBLP profile ↗
← Back
5ranked-venue papers
0as first author
4since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Databases, data management, data science and information retrieval · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Efficient and distributed learning · 82% Machine translation · 18%
Computer architecture, parallel and distributed computing, and storage systems
2 papers
Hardware accelerators and domain-specific architectures · 69% Performance modeling and evaluation · 31%
Databases, data mining, and information retrieval
1 paper
Machine learning and data management · 100%

Topics — the 8 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning
model compression
1.532024
Binarized Neural Machine Translation · NeurIPS 2023
PokeBNN: A Binary Pursuit of Lightweight Accuracy · CVPR 2022
PikeLPN: Mitigating Overlooked Inefficiencies of Low-Precision Neural Networks · CVPR 2024
Machine learning › Efficient and distributed learning › model compression
quantization
0.922024
Binarized Neural Machine Translation · NeurIPS 2023
PikeLPN: Mitigating Overlooked Inefficiencies of Low-Precision Neural Networks · CVPR 2024
Hardware accelerators and domain-specific architectures › machine learning accelerator › DNN inference
low-precision DNN inference
0.812024
PikeLPN: Mitigating Overlooked Inefficiencies of Low-Precision Neural Networks · CVPR 2024
Hardware accelerators and domain-specific architectures
machine learning accelerator
0.812024
PikeLPN: Mitigating Overlooked Inefficiencies of Low-Precision Neural Networks · CVPR 2024
Natural language and speech › Machine translation
neural machine translation
0.712023
Binarized Neural Machine Translation · NeurIPS 2023
Machine learning › Efficient and distributed learning › model compression › quantization › quantized neural network
binary neural network
0.612022
PokeBNN: A Binary Pursuit of Lightweight Accuracy · CVPR 2022
Machine learning and data management
machine learning lifecycle management
0.312017
TFX: A TensorFlow-Based Production-Scale Machine Learning Platform · KDD 2017
Machine learning and data management › machine learning systems
machine learning platform
0.312017
TFX: A TensorFlow-Based Production-Scale Machine Learning Platform · KDD 2017

Methods — techniques the papers use, named apart from their topics

quantization · 1.5double quantization · 1.5distribution-heterogeneous quantization · 1.5pokeconv · 1.1arithmetic computation effort · 1.1scaling law study · 0.7residual connections · 0.7layernorm · 0.7binarization · 0.7
YearPublicationVenuePosition
2024 PikeLPN: Mitigating Overlooked Inefficiencies of Low-Precision Neural Networks
abstract
Low-precision quantization is recognized for its efficacy in neural network optimization. Our analysis reveals that non-quantized elementwise operations which are prevalent in layers such as parameterized activation functions, batch normalization, and quantization scaling dominate the inference cost of low-precision models. These non-quantized elementwise operations are commonly overlooked in SOTA efficiency metrics such as Arithmetic Computation Effort (ACE) [46]. In this paper, we propose ACEv2- an extended version of ACE which offers a better alignment with the inference cost of quantized models and their energy consumption on ML hardware. Moreover, we introduce PikeLPN11Pike is a slim fast fish, LPN stands for Low-Precision Network., a model that addresses these efficiency issues by applying quantization to both elementwise operations and multiply-accumulate operations. In particular, we present a novel quantization technique for batch normalization layers named QuantNorm which allows for quantizing the batch normalization parameters without compromising the model performance. Additionally, we propose applying Double Quantization where the quantization scaling parameters are quantized. Furthermore, we recognize and resolve the issue of distribution mismatch in Separable Convolution layers by introducing Distribution-Heterogeneous Quantization which enables quantizing them to low-precision. PikeLPN achieves Pareto-optimality in efficiency-accuracy trade-off with up to 3× efficiency improvement compared to SOTA low-precision models.
Marina Neseem, Conor McCullough, Randy Hsin, Chas Leichner, Shan Li 0001, In Suk Chong, Andrew G. Howard, Lukasz Lew, Sherief Reda, Ville-Mikko Rautio, Daniele Moro
CVPR8
2023 Binarized Neural Machine Translation
abstract
The rapid scaling of language models is motivating research using low-bitwidth quantization. In this work, we propose a novel binarization technique for Transformers applied to machine translation (BMT), the first of its kind. We identify and address the problem of inflated dot-product variance when using one-bit weights and activations. Specifically, BMT leverages additional LayerNorms and residual connections to improve binarization quality. Experiments on the WMT dataset show that a one-bit weight-only Transformer can achieve the same quality as a float one, while being 16$\times$ smaller in size. One-bit activations incur varying degrees of quality drop, but mitigated by the proposed architectural changes. We further conduct a scaling law study using production-scale translation datasets, which shows that one-bit weight Transformers scale and generalize well in both in-domain and out-of-domain settings. Implementation in JAX/Flax will be open sourced.
Yichi Zhang 0006, Ankush Garg, Lukasz Lew, Behrooz Ghorbani, Zhiru Zhang, Orhan Firat
NeurIPS4
2022 PokeBNN: A Binary Pursuit of Lightweight Accuracy
abstract
Optimization of Top-1 ImageNet promotes enormous networks that may be impractical in inference settings. Binary neural networks (BNNs) have the potential to significantly lower the compute intensity but existing models suffer from low quality. To overcome this deficiency, we propose Poke- Conv, a binary convolution block which improves quality of BNNs by techniques such as adding multiple residual paths, and tuning the activation function. We apply it to ResNet-50 and optimize ResNet's initial convolutional layer which is hard to binarize. We name the resulting network family PokeBNN11Poke/pnki/is pronounced similarly to pocket. PokeConv, PokeBNN, and Pokemon are abbreviations of Pocket Convolution, Pocket Binary Neural Network, and Pocket Monster, respectively.. These techniques are chosen to yield favorable improvements in both top-1 accuracy and the network's cost. In order to enable joint optimization of the cost together with accuracy, we define arithmetic computation effort (ACE), a hardware- and energy-inspired cost metric for quantized and binarized networks. We also identify a need to optimize an under-explored hyper-parameter controlling the binarization gradient approximation. We establish a new, strong state-of-the-art (SOTA) on top-1 accuracy together with commonly-used CPU64 cost, ACE cost and network size metrics. ReActNet-Adam [33], the previous SOTA in BNNs, achieved a 70.5% top-1 accuracy with 7.9 ACE. A small variant of PokeBNN achieves 70.5% top-1 with 2.6 ACE, more than 3x reduction in cost; a larger PokeBNN achieves 75.6% top-1 with 7.8 ACE, more than 5% improvement in accuracy without increasing the cost. PokeBNN implementation in JAX/Flax [6, 18] and re-production instructions are open sourced.22Source code and reproduction instructions are available in AQT repos-itory: github.com/google/aqt.
Yichi Zhang 0006, Zhiru Zhang, Lukasz Lew
CVPR3
2022 4-bit Conformer with Native Quantization Aware Training for Speech Recognition
abstract
Reducing the latency and model size has always been a significant research problem for live Automatic Speech Recognition (ASR) application scenarios. Along this direction, model quantization has become an increasingly popular approach to compress neural networks and reduce computation cost. Most of the existing practical ASR systems apply post-training 8-bit quantization. To achieve a higher compression rate without introducing additional performance regression, in this study, we propose to develop 4-bit ASR models with native quantization aware training, which leverages native integer operations to effectively optimize both training and inference. We conducted two experiments on state-of-the-art Conformer-based ASR models to evaluate our proposed quantization technique. First, we explored the impact of different precisions for both weight and activation quantization on the LibriSpeech dataset, and obtained a lossless 4-bit Conformer model with 7.7x size reduction compared to the float32 model. Following this, we for the first time investigated and revealed the viability of 4-bit quantization on a practical ASR system that is trained with large-scale datasets, and produced a lossless Conformer ASR model with mixed 4-bit and 8-bit weights that has 5x size reduction compared to the float32 model.
Shaojin Ding, Phoenix Meadowlark, Yanzhang He, Lukasz Lew, Shivani Agrawal, Oleg Rybakov
INTERSPEECH4
2017 TFX: A TensorFlow-Based Production-Scale Machine Learning Platform
abstract
Creating and maintaining a platform for reliably producing and deploying machine learning models requires careful orchestration of many components---a learner for generating models based on training data, modules for analyzing and validating both data as well as models, and finally infrastructure for serving models in production. This becomes particularly challenging when data changes over time and fresh models need to be produced continuously. Unfortunately, such orchestration is often done ad hoc using glue code and custom scripts developed by individual teams for specific use cases, leading to duplicated effort and fragile systems with high technical debt.
Denis Baylor, Eric Breck, Heng-Tze Cheng, Noah Fiedel, Chuan Yu Foo, Zakaria Haque, Salem Haykal, Mustafa Ispir, Vihan Jain, Levent Koc 0001, Chiu Yuen Koo, Lukasz Lew, Clemens Mewald, Akshay Naresh Modi, Neoklis Polyzotis, Sukriti Ramesh, Sudip Roy 0002, Steven Euijong Whang, Martin Wicke, Jarek Wilkiewicz, Martin Zinkevich
KDD12