VLDB 2026 Research / reviewers in the wild / expert
Adrian Lancucki
dblp:140/7631
· DBLP profile ↗
17ranked-venue papers
6as first author
9since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 4 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Inference-Time Hyper-Scaling with KV Cache CompressionabstractInference-time scaling trades efficiency for increased reasoning accuracy by generating longer or more parallel sequences. However, in Transformer LLMs, generation cost is bottlenecked by the size of the key–value (KV) cache, rather than the number of generated tokens. Hence, we explore inference-time hyper-scaling: by compressing the KV cache, we can generate more tokens within the same compute budget and further improve the accuracy of scaled inference. The success of this approach, however, hinges on the ability of compression methods to preserve accuracy even at high compression ratios. To make hyper-scaling practical, we introduce Dynamic Memory Sparsification (DMS), a novel method for sparsifying KV caches that only requires 1K training steps to achieve 8× compression, while maintaining better accuracy than training-free sparse attention. Instead of prematurely discarding cached tokens, DMS delays token eviction, implicitly merging representations and preserving critical information. We demonstrate the effectiveness of inference-time hyper-scaling with DMS on multiple families of LLMs, showing that it boosts accuracy for comparable inference latency and memory load. For instance, we enhance Qwen-R1 32B by 9.1 points on AIME 24, 7.6 on GPQA, and 9.6 on LiveCodeBench on average for an equivalent number of memory reads. Adrian Lancucki, Konrad Staniszewski, Piotr Nawrot, Edoardo Maria Ponti |
NeurIPS | 1 |
| 2024 | Dynamic Memory Compression: Retrofitting LLMs for Accelerated InferenceabstractTransformers have emerged as the backbone of large language models (LLMs). However, generation remains inefficient due to the need to store in memory a cache of key–value representations for past tokens, whose size scales linearly with the input sequence length and batch size. As a solution, we propose Dynamic Memory Compression (DMC), a method for on-line key–value cache compression at inference time. Most importantly, the model learns to apply different compression ratios in different heads and layers. We retrofit pre-trained LLMs such as Llama 2 (7B, 13B and 70B) into DMC Transformers, achieving up to $\sim 3.7 \times$ throughput increase during auto-regressive inference on an NVIDIA H100 GPU. DMC is applied via continued pre-training on a negligible percentage of the original data without adding any extra parameters. We find that DMC preserves the original downstream performance with up to 4$\times$ cache compression, outperforming up-trained grouped-query attention (GQA) and key–value eviction policies (H$_2$O, TOVA). GQA and DMC can be even combined to obtain compounded gains. As a result DMC fits longer contexts and larger batches within any given memory budget. We release the DMC code and models at https://github.com/NVIDIA/Megatron-LM/tree/DMC. Piotr Nawrot, Adrian Lancucki, Marcin Chochowski, David Tarjan, Edoardo Maria Ponti |
ICML | 2 |
| 2023 | Efficient Transformers with Dynamic Token PoolingabstractTransformers achieve unrivalled performance in modelling language, but remain inefficient in terms of memory and time complexity.A possible remedy is to reduce the sequence length in the intermediate layers by pooling fixed-length segments of tokens.Nevertheless, natural units of meaning, such as words or phrases, display varying sizes.To address this mismatch, we equip language models with a dynamic-pooling mechanism, which predicts segment boundaries in an autoregressive fashion.We compare several methods to infer boundaries, including end-to-end learning through stochastic re-parameterisation, supervised learning (based on segmentations from subword tokenizers or spikes in conditional entropy), as well as linguistically motivated boundaries.We perform character-level evaluation on texts from multiple datasets and morphologically diverse languages.The results demonstrate that dynamic pooling, which jointly segments and models language, is both faster and more accurate than vanilla Transformers and fixed-length pooling within the same computational budget. Piotr Nawrot, Jan Chorowski, Adrian Lancucki, Edoardo Maria Ponti |
ACL (1) | 3 |
| 2022 | One TTS Alignment to Rule Them AllabstractSpeech-to-text alignment is a critical component of neural text-to-speech (TTS) models. Autoregressive TTS models typically use an attention mechanism to learn these alignments on-line. However, these alignments tend to be brittle and often fail to generalize to long utterances and out-of-domain text, leading to missing or repeating words. Most non-autoregressive end-to-end TTS models rely on durations extracted from external sources. In this paper we leverage the alignment mechanism proposed in RAD-TTS and demonstrate its applicability to wide variety of neural TTS models. The alignment learning framework combines the forward-sum algorithm, Viterbi algorithm, and an efficient static prior. In our experiments, the framework improves all tested TTS architectures, both autoregressive (Flowtron, Tacotron 2) and non-autoregressive (FastPitch, FastSpeech 2, RAD-TTS). Specifically, it improves alignment convergence speed, simplifies the training pipeline by eliminating need for external aligners, enhances robustness to errors on long utterances and improves the perceived speech synthesis quality, as judged by human evaluators. Rohan Badlani, Adrian Lancucki, Kevin J. Shih, Rafael Valle, Wei Ping, Bryan Catanzaro |
ICASSP | 2 |
| 2022 | Contrastive Prediction Strategies for Unsupervised Segmentation and Categorization of Phonemes and WordsabstractWe identify a performance trade-off between the tasks of phoneme categorization and phoneme and word segmentation in several self-supervised learning algorithms based on Contrastive Predictive Coding (CPC). Our experiments suggest that context building networks, albeit necessary for high performance on categorization tasks, harm segmentation performance by causing a temporal shift on the learned representations. Aiming to tackle this trade-off, we take inspiration from the leading approaches on segmentation and propose multi-level Aligned CPC (mACPC). It builds on Aligned CPC (ACPC), a variant of CPC which exhibits the best performance on categorization tasks, and incorporates multi-level modeling and optimization for detection of spectral changes. Our methods improve in all tested categorization metrics and achieve state-of-the-art performance in word segmentation. Santiago Cuervo, Maciej Grabias, Jan Chorowski, Grzegorz Ciesielski, Adrian Lancucki, Pawel Rychlikowski, Ricard Marxer |
ICASSP | 5 |
| 2022 | Variable-rate hierarchical CPC leads to acoustic unit discovery in speechabstractThe success of deep learning comes from its ability to capture the hierarchical structure of data by learning high-level representations defined in terms of low-level ones. In this paper we explore self-supervised learning of hierarchical representations of speech by applying multiple levels of Contrastive Predictive Coding (CPC). We observe that simply stacking two CPC models does not yield significant improvements over single-level architectures. Inspired by the fact that speech is often described as a sequence of discrete units unevenly distributed in time, we propose a model in which the output of a low-level CPC module is non-uniformly downsampled to directly minimize the loss of a high-level CPC module. The latter is designed to also enforce a prior of separability and discreteness in its representations by enforcing dissimilarity of successive high-level representations through focused negative sampling, and by quantization of the prediction targets. Accounting for the structure of the speech signal improves upon single-level CPC features and enhances the disentanglement of the learned representations, as measured by downstream speech recognition tasks, while resulting in a meaningful segmentation of the signal that closely resembles phone boundaries. Santiago Cuervo, Adrian Lancucki, Ricard Marxer, Pawel Rychlikowski, Jan Chorowski |
NeurIPS | 2 |
| 2021 | Fastpitch: Parallel Text-to-Speech with Pitch PredictionabstractWe present FastPitch, a fully-parallel text-to-speech model based on FastSpeech, conditioned on fundamental frequency contours. The model predicts pitch contours during inference. By altering these predictions, the generated speech can be more expressive, better match the semantic of the utterance, and in the end more engaging to the listener. Uniformly increasing or decreasing pitch with FastPitch generates speech that resembles the voluntary modulation of voice. Conditioning on frequency contours improves the overall quality of synthesized speech, making it comparable to state-of-the-art. It does not introduce an overhead, and FastPitch retains the favorable, fully-parallel Transformer architecture, with over 900× real-time factor for mel-spectrogram synthesis of a typical utterance. Adrian Lancucki |
ICASSP | 1 |
| 2021 | Information Retrieval for ZeroSpeech 2021: The Submission by University of WroclawabstractWe present a number of low-resource approaches to the tasks of the Zero Resource Speech Challenge 2021. We build on the unsupervised representations of speech proposed by the organizers as a baseline, derived from CPC and clustered with the k-means algorithm. We demonstrate that simple methods of refining those representations can narrow the gap, or even improve upon the solutions which use a high computational budget. The results lead to the conclusion that the CPC-derived representations are still too noisy for training language models, but stable enough for simpler forms of pattern matching and retrieval. Jan Chorowski, Grzegorz Ciesielski, Jaroslaw Dzikowski, Adrian Lancucki, Ricard Marxer, Mateusz Opala, Piotr Pusz, Pawel Rychlikowski, Michal Stypulkowski |
Interspeech | 4 |
| 2021 | Aligned Contrastive Predictive CodingabstractWe investigate the possibility of forcing a self-supervised model trained using a contrastive predictive loss to extract slowly varying latent representations. Rather than producing individual predictions for each of the future representations, the model emits a sequence of predictions shorter than that of the upcoming representations to which they will be aligned. In this way, the prediction network solves a simpler task of predicting the next symbols, but not their exact timing, while the encoding network is trained to produce piece-wise constant latent codes. We evaluate the model on a speech coding task and demonstrate that the proposed Aligned Contrastive Predictive Coding (ACPC) leads to higher linear phone prediction accuracy and lower ABX error rates, while being slightly faster to train due to the reduced number of prediction heads. Jan Chorowski, Grzegorz Ciesielski, Jaroslaw Dzikowski, Adrian Lancucki, Ricard Marxer, Mateusz Opala, Piotr Pusz, Pawel Rychlikowski, Michal Stypulkowski |
Interspeech | 4 |
| 2020 | Robust Training of Vector Quantized Bottleneck ModelsabstractIn this paper we demonstrate methods for reliable and efficient training of discrete representation using Vector-Quantized Variational Auto-Encoder models (VQ-VAEs). Discrete latent variable models have been shown to learn nontrivial representations of speech, applicable to unsupervised voice conversion and reaching state-of-the-art performance on unit discovery tasks. For unsupervised representation learning, they became viable alternatives to continuous latent variable models such as the Variational Auto-Encoder (VAE). However, training deep discrete variable models is challenging, due to the inherent non-differentiability of the discretization operation. In this paper we focus on VQ-VAE, a state-of-the-art discrete bottleneck model shown to perform on par with its continuous counterparts. It quantizes encoder outputs with on-line k-means clustering. We show that the codebook learning can suffer from poor initialization and non-stationarity of clustered encoder outputs. We demonstrate that these can be successfully overcome by increasing the learning rate for the codebook and periodic date-dependent codeword re-initialization. As a result, we achieve more robust training across different tasks, and significantly increase the usage of latent codewords even for large codebooks. This has practical benefit, for instance, in unsupervised representation learning, where large codebooks may lead to disentanglement of latent representations. Adrian Lancucki, Jan Chorowski, Guillaume Sanchez, Ricard Marxer, Nanxin Chen, Hans J. G. A. Dolfing, Sameer Khurana, Tanel Alumäe, Antoine Laurent |
IJCNN | 1 |
| 2020 | A Convolutional Deep Markov Model for Unsupervised Speech Representation LearningabstractProbabilistic Latent Variable Models (LVMs) provide an alternative to self-supervised learning approaches for linguistic representation learning from speech. LVMs admit an intuitive probabilistic interpretation where the latent structure shapes the information extracted from the signal. Even though LVMs have recently seen a renewed interest due to the introduction of Variational Autoencoders (VAEs), their use for speech representation learning remains largely unexplored. In this work, we propose Convolutional Deep Markov Model (ConvDMM), a Gaussian state-space model with non-linear emission and transition functions modelled by deep neural networks. This unsupervised model is trained using black box variational inference. A deep convolutional neural network is used as an inference network for structured variational approximation. When trained on a large scale speech dataset (LibriSpeech), ConvDMM produces features that significantly outperform multiple self-supervised feature extracting methods on linear phone classification and recognition on the Wall Street Journal dataset. Furthermore, we found that ConvDMM complements self-supervised methods like Wav2Vec and PASE, improving on the results achieved with any of the methods alone. Lastly, we find that ConvDMM features enable learning better phone recognizers than any other features in an extreme low-resource regime with few labeled training examples. Sameer Khurana, Antoine Laurent, Wei-Ning Hsu, Jan Chorowski, Adrian Lancucki, Ricard Marxer, James R. Glass |
INTERSPEECH | 5 |
| 2019 | Towards Using Context-Dependent Symbols in CTC Without State-Tying Decision TreesabstractDeep neural acoustic models benefit from context-dependent (CD) modeling of output symbols. We consider direct training of CTC networks with CD outputs, and identify two issues. The first one is frame-level normalization of probabilities in CTC, which induces strong language modeling behavior that leads to overfitting and interference with external language models. The second one is poor generalization in the presence of numerous lexical units like triphones or tri-chars. We mitigate the former with utterance-level normalization of probabilities. The latter typically requires reducing the CD symbol inventory with state-tying decision trees, which have to be transferred from classical GMM-HMM systems. We replace the trees with a CD symbol embedding network, which saves parameters and ensures generalization to unseen and undersampled CD symbols. The embedding network is trained together with the rest of the acoustic model and removes one of the last cases in which neural systems have to be bootstrapped from GMM-HMM ones. Jan Chorowski, Adrian Lancucki, Bartosz Kostka, Michal Zapotoczny |
INTERSPEECH | 2 |
| 2019 | Lattice Generation in Attention-Based Speech Recognition Models
Michal Zapotoczny, Piotr Pietrzak, Adrian Lancucki, Jan Chorowski |
INTERSPEECH | 3 |
| 2016 | A new evolutionary microRNA marker selection using next-generation sequencing dataabstractNext-generation sequencing allows high-throughput measurements of non-coding RNA expression levels in tissues. Analysis of microRNAs (miRNAs) is particularly effective in differentiation of cancerous tissue samples, based on patterns of their expression levels. The paper presents a wrapper feature selection approach based on t-Distributed Stochastic Neighbor Embedding (t-SNE), Covariance Matrix Adaptation Evolution Strategy (CMA-Es) and Support Vector Machine (SVM). The advantage of t-SNE is amplification of pairwise similarities by the means of t-Student neighborhood function. The attributes are embedded into 1-D space to reveal similarities between the features. Such information is used by CMA-ES through real-valued encoding in order to model pairwise relations between miRNAs with covariance matrices. Finally, the wrapper uses SVM to evaluate the objective, which expresses the tradeoff between classification quality and the desired number of features. The approach is tested on eight different cancer types from The Cancer Genome Atlas. It allows to find small sets of miRNAs to differentiate cancer types from a single tumor class to the normal one with high certainty. Adrian Lancucki, Indrajit Saha, Shib Sankar Bhowmick, Ujjwal Maulik, Piotr Lipinski |
CEC | 1 |
| 2016 | Multiobjective optimization of frequent pattern models in ultra-high frequency time series: Stability versus universalityabstractThis paper concerns knowledge discovery from ultra-high frequency financial time series of order book data from the London Stock Exchange Rebuild Order Book database. It focuses on extracting frequent patterns from order book shapes, which characterize particular stock market states. Because order book shapes vary considerably in time, it is necessary to convert the original price-volume representation to real vectors of a fixed length. In this paper two methods of representing order books are studied: using Gaussian Mixture Models (GMMs), and Gaussian density-based representation (GD). Both methods produce real vectors of a constant length and frequent patterns are defined as codebooks of clusters of these real vectors. To obtain useful knowledge we would like these patterns to represent different states of the market, be stable in time, and be universal. We define universality as periodical reoccurence, which makes historical knowledge usable for future decision making. Because of these time-related requirements a typical spatial clustering is not sufficient. This paper studies the usage of an evolutionary algorithm to obtain patterns with desirable properties. Krzysztof Michalak, Adrian Lancucki, Piotr Lipinski |
CEC | 2 |
| 2015 | A new evolutionary gene selection techniqueabstractMicroarray technology allows to investigate gene expression levels by analyzing high dimensional datasets of few samples. Selection of discriminative, differentially expressed genes from such datasets is important to differentiate, prognose and understand the underlying biological processes. In this regard, the paper presents a new evolutionary gene selection method based on Student-t Stochastic Neighbor Embedding (t-SNE), Differential Evolution (DE) and Support Vector Machine (SVM). Here the underlying classification task of SVM is used as an optimization problem of DE, while t-SNE provides better ordering of genes for selection purpose. Generally, t-SNE is used to reorder the genes in such a way so that similar genes are grouped together and dissimilar genes are kept further apart. These reordered genes are then fragmented into fixed-length partitions. Thereafter, from each partition, a gene is selected randomly to encode the initial population of DE along with the combination of its weight and threshold values in order to participate in fitness computation. In the final generation of DE, a subset of genes is selected based on higher classification accuracy. The proposed technique is tested on six publicly available microarray datasets concerning various cancerous tissues of Homo sapiens and yields a potential set of genes by providing prefect or nearly perfect classification accuracy. Moreover, the superiority of the proposed technique has been demonstrated in comparison with other widely used techniques. Finally, the achieved results have also been justified by a statistical test and allowed us to draw biological conclusions through the identification of Gene Ontologies. Adrian Lancucki, Indrajit Saha, Piotr Lipinski |
CEC | 1 |
| 2014 | Continuous Population-Based Incremental Learning with Mixture Probability Modeling for Dynamic Optimization Problems
Adrian Lancucki, Jan Chorowski, Krzysztof Michalak, Patryk Filipiak, Piotr Lipinski |
IDEAL | 1 |