EDBT 2026 Demo / reviewers in the wild / expert
Lukasz Lew
dblp:52/5311
· DBLP profile ↗
5ranked-venue papers
0as first author
4since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Databases, data management, data science and information retrieval · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Efficient and distributed learning · 82% Machine translation · 18% | |
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
Hardware accelerators and domain-specific architectures · 69% Performance modeling and evaluation · 31% | |
| Databases, data mining, and information retrieval
1 paper |
Machine learning and data management · 100% |
Topics — the 8 heaviest of 10, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Efficient and distributed learning
model compression |
1.5 | 3 | 2024 | Binarized Neural Machine Translation · NeurIPS 2023 PokeBNN: A Binary Pursuit of Lightweight Accuracy · CVPR 2022 PikeLPN: Mitigating Overlooked Inefficiencies of Low-Precision Neural Networks · CVPR 2024 |
Machine learning › Efficient and distributed learning › model compression
quantization |
0.9 | 2 | 2024 | Binarized Neural Machine Translation · NeurIPS 2023 PikeLPN: Mitigating Overlooked Inefficiencies of Low-Precision Neural Networks · CVPR 2024 |
Hardware accelerators and domain-specific architectures › machine learning accelerator › DNN inference
low-precision DNN inference |
0.8 | 1 | 2024 | PikeLPN: Mitigating Overlooked Inefficiencies of Low-Precision Neural Networks · CVPR 2024 |
Hardware accelerators and domain-specific architectures
machine learning accelerator |
0.8 | 1 | 2024 | PikeLPN: Mitigating Overlooked Inefficiencies of Low-Precision Neural Networks · CVPR 2024 |
Natural language and speech › Machine translation
neural machine translation |
0.7 | 1 | 2023 | Binarized Neural Machine Translation · NeurIPS 2023 |
Machine learning › Efficient and distributed learning › model compression › quantization › quantized neural network
binary neural network |
0.6 | 1 | 2022 | PokeBNN: A Binary Pursuit of Lightweight Accuracy · CVPR 2022 |
Machine learning and data management
machine learning lifecycle management |
0.3 | 1 | 2017 | TFX: A TensorFlow-Based Production-Scale Machine Learning Platform · KDD 2017 |
Machine learning and data management › machine learning systems
machine learning platform |
0.3 | 1 | 2017 | TFX: A TensorFlow-Based Production-Scale Machine Learning Platform · KDD 2017 |
Methods — techniques the papers use, named apart from their topics
quantization · 1.5double quantization · 1.5distribution-heterogeneous quantization · 1.5pokeconv · 1.1arithmetic computation effort · 1.1scaling law study · 0.7residual connections · 0.7layernorm · 0.7binarization · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | PikeLPN: Mitigating Overlooked Inefficiencies of Low-Precision Neural NetworksabstractLow-precision quantization is recognized for its efficacy in neural network optimization. Our analysis reveals that non-quantized elementwise operations which are prevalent in layers such as parameterized activation functions, batch normalization, and quantization scaling dominate the inference cost of low-precision models. These non-quantized elementwise operations are commonly overlooked in SOTA efficiency metrics such as Arithmetic Computation Effort (ACE) [46]. In this paper, we propose ACEv2- an extended version of ACE which offers a better alignment with the inference cost of quantized models and their energy consumption on ML hardware. Moreover, we introduce PikeLPN11Pike is a slim fast fish, LPN stands for Low-Precision Network., a model that addresses these efficiency issues by applying quantization to both elementwise operations and multiply-accumulate operations. In particular, we present a novel quantization technique for batch normalization layers named QuantNorm which allows for quantizing the batch normalization parameters without compromising the model performance. Additionally, we propose applying Double Quantization where the quantization scaling parameters are quantized. Furthermore, we recognize and resolve the issue of distribution mismatch in Separable Convolution layers by introducing Distribution-Heterogeneous Quantization which enables quantizing them to low-precision. PikeLPN achieves Pareto-optimality in efficiency-accuracy trade-off with up to 3× efficiency improvement compared to SOTA low-precision models. Marina Neseem, Conor McCullough, Randy Hsin, Chas Leichner, Shan Li 0001, In Suk Chong, Andrew G. Howard, Lukasz Lew, Sherief Reda, Ville-Mikko Rautio, Daniele Moro |
CVPR | 8 |
| 2023 | Binarized Neural Machine TranslationabstractThe rapid scaling of language models is motivating research using low-bitwidth quantization.
In this work, we propose a novel binarization technique for Transformers applied to machine translation (BMT), the first of its kind. We identify and address the problem of inflated dot-product variance when using one-bit weights and activations. Specifically, BMT leverages additional LayerNorms and residual connections to improve binarization quality. Experiments on the WMT dataset show that a one-bit weight-only Transformer can achieve the same quality as a float one, while being 16$\times$ smaller in size. One-bit activations incur varying degrees of quality drop, but mitigated by the proposed architectural changes. We further conduct a scaling law study using production-scale translation datasets, which shows that one-bit weight Transformers scale and generalize well in both in-domain and out-of-domain settings. Implementation in JAX/Flax will be open sourced. Yichi Zhang 0006, Ankush Garg, Lukasz Lew, Behrooz Ghorbani, Zhiru Zhang, Orhan Firat |
NeurIPS | 4 |
| 2022 | PokeBNN: A Binary Pursuit of Lightweight AccuracyabstractOptimization of Top-1 ImageNet promotes enormous networks that may be impractical in inference settings. Binary neural networks (BNNs) have the potential to significantly lower the compute intensity but existing models suffer from low quality. To overcome this deficiency, we propose Poke- Conv, a binary convolution block which improves quality of BNNs by techniques such as adding multiple residual paths, and tuning the activation function. We apply it to ResNet-50 and optimize ResNet's initial convolutional layer which is hard to binarize. We name the resulting network family PokeBNN11Poke/pnki/is pronounced similarly to pocket. PokeConv, PokeBNN, and Pokemon are abbreviations of Pocket Convolution, Pocket Binary Neural Network, and Pocket Monster, respectively.. These techniques are chosen to yield favorable improvements in both top-1 accuracy and the network's cost. In order to enable joint optimization of the cost together with accuracy, we define arithmetic computation effort (ACE), a hardware- and energy-inspired cost metric for quantized and binarized networks. We also identify a need to optimize an under-explored hyper-parameter controlling the binarization gradient approximation. We establish a new, strong state-of-the-art (SOTA) on top-1 accuracy together with commonly-used CPU64 cost, ACE cost and network size metrics. ReActNet-Adam [33], the previous SOTA in BNNs, achieved a 70.5% top-1 accuracy with 7.9 ACE. A small variant of PokeBNN achieves 70.5% top-1 with 2.6 ACE, more than 3x reduction in cost; a larger PokeBNN achieves 75.6% top-1 with 7.8 ACE, more than 5% improvement in accuracy without increasing the cost. PokeBNN implementation in JAX/Flax [6, 18] and re-production instructions are open sourced.22Source code and reproduction instructions are available in AQT repos-itory: github.com/google/aqt. Yichi Zhang 0006, Zhiru Zhang, Lukasz Lew |
CVPR | 3 |
| 2022 | 4-bit Conformer with Native Quantization Aware Training for Speech RecognitionabstractReducing the latency and model size has always been a significant research problem for live Automatic Speech Recognition (ASR) application scenarios. Along this direction, model quantization has become an increasingly popular approach to compress neural networks and reduce computation cost. Most of the existing practical ASR systems apply post-training 8-bit quantization. To achieve a higher compression rate without introducing additional performance regression, in this study, we propose to develop 4-bit ASR models with native quantization aware training, which leverages native integer operations to effectively optimize both training and inference. We conducted two experiments on state-of-the-art Conformer-based ASR models to evaluate our proposed quantization technique. First, we explored the impact of different precisions for both weight and activation quantization on the LibriSpeech dataset, and obtained a lossless 4-bit Conformer model with 7.7x size reduction compared to the float32 model. Following this, we for the first time investigated and revealed the viability of 4-bit quantization on a practical ASR system that is trained with large-scale datasets, and produced a lossless Conformer ASR model with mixed 4-bit and 8-bit weights that has 5x size reduction compared to the float32 model. Shaojin Ding, Phoenix Meadowlark, Yanzhang He, Lukasz Lew, Shivani Agrawal, Oleg Rybakov |
INTERSPEECH | 4 |
| 2017 | TFX: A TensorFlow-Based Production-Scale Machine Learning PlatformabstractCreating and maintaining a platform for reliably producing and deploying machine learning models requires careful orchestration of many components---a learner for generating models based on training data, modules for analyzing and validating both data as well as models, and finally infrastructure for serving models in production. This becomes particularly challenging when data changes over time and fresh models need to be produced continuously. Unfortunately, such orchestration is often done ad hoc using glue code and custom scripts developed by individual teams for specific use cases, leading to duplicated effort and fragile systems with high technical debt. Denis Baylor, Eric Breck, Heng-Tze Cheng, Noah Fiedel, Chuan Yu Foo, Zakaria Haque, Salem Haykal, Mustafa Ispir, Vihan Jain, Levent Koc 0001, Chiu Yuen Koo, Lukasz Lew, Clemens Mewald, Akshay Naresh Modi, Neoklis Polyzotis, Sukriti Ramesh, Sudip Roy 0002, Steven Euijong Whang, Martin Wicke, Jarek Wilkiewicz, Martin Zinkevich |
KDD | 12 |