EDBT 2026 Demo / reviewers in the wild / expert
Giovanni Motta
dblp:51/675
· DBLP profile ↗
29ranked-venue papers
11as first author
7since 2021 · last 2023
0009-0004-6019-9960ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 27 · 9 first-author · 7 since 2021Databases, data management, data science and information retrieval · 10 · 5 first-authorArtificial intelligence and machine learning · 3 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | The Gift of Feedback: Improving ASR Model Quality by Learning from User Corrections Through Federated LearningabstractAutomatic speech recognition (ASR) models are typically trained on large datasets of transcribed speech. As language evolves and new terms come into use, these models can become outdated and stale. In the context of models trained on the server but deployed on edge devices, errors may result from the mismatch between server training data and actual on-device usage. In this work, we seek to continually learn from on-device user corrections through Federated Learning (FL) to address this issue. We explore techniques to target fresh terms that the model has not previously encountered, learn long-tail words, and mitigate catastrophic forgetting. In experimental evaluations, we find that the proposed techniques improve model recognition of fresh terms, while preserving quality on the overall language distribution. Lillian Zhou, Mingqing Chen, Harry Zhang, Rohit Prabhavalkar, Dhruv Guliani, Giovanni Motta, Rajiv Mathews |
ASRU | 7 |
| 2023 | Online Model Compression for Federated Learning with Large ModelsabstractThis paper addresses the challenges of training large neural networks under federated learning settings: high on-device memory usage and communication cost. The proposed Online Model Compression (OMC) provides a framework that stores model parameters in a compressed format and decompresses them only when needed. We use quantization as the compression method in this paper and propose three methods, (1) per-variable transformation, (2) weight-matrix-only quantization, and (3) partial variable quantization, to minimize its impact on model accuracy. Our experiments on two recent neural networks for speech recognition and two different datasets show that OMC can reduce memory usage and communication cost of model parameters by up to 59% while attaining comparable accuracy and training speed when compared with full-precision federated learning. Tien-Ju Yang, Yonghui Xiao, Giovanni Motta, Françoise Beaufays, Rajiv Mathews, Mingqing Chen |
ICASSP | 3 |
| 2022 | Enabling On-Device Training of Speech Recognition Models With Federated DropoutabstractFederated learning can be used to train machine learning models on the edge on local data that never leave devices, providing privacy by default. This presents a challenge pertaining to the communication and computation costs associated with clients’ devices. These costs are strongly correlated with the size of the model being trained, and are significant for state-of-the-art automatic speech recognition models.We propose using federated dropout to reduce the size of client models while training a full-size model server-side. We provide empirical evidence of the effectiveness of federated dropout, and propose a novel approach to vary the dropout rate applied at each layer. Furthermore, we find that federated dropout enables a set of smaller sub-models within the larger model to independently have low word error rates, making it easier to dynamically adjust the size of the model deployed for inference. Dhruv Guliani, Lillian Zhou, Changwan Ryu, Tien-Ju Yang, Harry Zhang, Yonghui Xiao, Françoise Beaufays, Giovanni Motta |
ICASSP | 8 |
| 2022 | Partial Variable Training for Efficient on-Device Federated LearningabstractThis paper aims to address the major challenges of Federated Learning (FL) on edge devices: limited memory and expensive communication. We propose a novel method, called Partial Variable Training (PVT), that only trains a small subset of variables on edge devices to reduce memory usage and communication cost. With PVT, we show that network accuracy can be maintained by utilizing more local training steps and devices, which is favorable for FL involving a large population of devices. According to our experiments on two state-of-the-art neural networks for speech recognition and two different datasets, PVT can reduce memory usage by up to 1.9× and communication cost by up to 593× while attaining comparable accuracy when compared with full network training. Tien-Ju Yang, Dhruv Guliani, Françoise Beaufays, Giovanni Motta |
ICASSP | 4 |
| 2022 | Exploring Heterogeneous Characteristics of Layers in ASR Models for More Efficient TrainingabstractTransformer-based architectures have been the subject of research aimed at understanding their overparameterization and the non-uniform importance of their layers. Applying these approaches to Automatic Speech Recognition, we demonstrate that the state-of-the-art Conformer models generally have multiple ambient layers. We study the stability of these layers across runs and model sizes, propose that group normalization may be used without disrupting their formation, and examine their correlation with model weight updates in each layer. Finally, we apply these findings to Federated Learning in order to improve the training procedure, by targeting Federated Dropout to layers by importance. This allows us to reduce the model size optimized by clients without quality degradation, and shows potential for future exploration. Lillian Zhou, Dhruv Guliani, Andreas Kabel, Giovanni Motta, Françoise Beaufays |
ICASSP | 4 |
| 2022 | Federated Pruning: Improving Neural Network Efficiency with Federated LearningabstractAutomatic Speech Recognition models require large amount of speech data for training, and the collection of such data often leads to privacy concerns.Federated learning has been widely used and is considered to be an effective decentralized technique by collaboratively learning a shared prediction model while keeping the data local on different clients devices.However, the limited computation and communication resources on clients devices present practical difficulties for large models.To overcome such challenges, we propose Federated Pruning to train a reduced model under the federated setting, while maintaining similar performance compared to the full model.Moreover, the vast amount of clients data can also be leveraged to improve the pruning results compared to centralized training.We explore different pruning schemes and provide empirical evidence of the effectiveness of our methods. Rongmei Lin, Yonghui Xiao, Tien-Ju Yang, Ding Zhao, Li Xiong 0001, Giovanni Motta, Françoise Beaufays |
INTERSPEECH | 6 |
| 2021 | Training Speech Recognition Models with Federated Learning: A Quality/Cost FrameworkabstractWe propose using federated learning, a decentralized on-device learning paradigm, to train speech recognition models. By performing epochs of training on a per-user basis, federated learning must incur the cost of dealing with non-IID data distributions, which are expected to negatively affect the quality of the trained model. We propose a framework by which the degree of non-IID-ness can be varied, consequently illustrating a trade-off between model quality and the computational cost of federated training, which we capture through a novel metric. Finally, we demonstrate that hyper-parameter optimization and appropriate use of variational noise are sufficient to compensate for the quality impact of non-IID distributions, while decreasing the cost. Dhruv Guliani, Françoise Beaufays, Giovanni Motta |
ICASSP | 3 |
| 2020 | Low-Rank Gradient Approximation for Memory-Efficient on-Device Training of Deep Neural NetworkabstractTraining machine learning models on mobile devices has the potential of improving both privacy and accuracy of the models. However, one of the major obstacles to achieving this goal is the memory limitation of mobile devices. Reducing training memory enables models with high-dimensional weight matrices, like automatic speech recognition (ASR) models, to be trained on-device. In this paper, we propose approximating the gradient matrices of deep neural networks using a low-rank parameterization as an avenue to save training memory. The low-rank gradient approximation enables more advanced, memory-intensive optimization techniques to be run on device. Our experimental results show that we can reduce the training memory by about 33.0% for Adam optimization. It uses comparable memory to momentum optimization and achieves a 4.5% relative lower word error rate on an ASR personalization task. Mary Gooneratne, Khe Chai Sim, Petr Zadrazil, Andreas Kabel, Françoise Beaufays, Giovanni Motta |
ICASSP | 6 |
| 2019 | Personalization of End-to-End Speech Recognition on Mobile Devices for Named EntitiesabstractWe study the effectiveness of several techniques to personalize end-to-end speech models and improve the recognition of proper names relevant to the user. These techniques differ in the amounts of user effort required to provide supervision, and are evaluated on how they impact speech recognition performance. We propose using keyword-dependent precision and recall metrics to measure vocabulary acquisition performance. We evaluate the algorithms on a dataset that we designed to contain names of persons that are difficult to recognize. Therefore, the baseline recall rate for proper names in this dataset is very low: 2.4%. A data synthesis approach we developed brings it to 48.6%, with no need for speech input from the user. With speech input, if the user corrects only the names, the name recall rate improves to 64.4%. If the user corrects all the recognition errors, we achieve the best recall of 73.5%. To eliminate the need to upload user data and store personalized models on a server, we focus on performing the entire personalization workflow on a mobile device. Khe Chai Sim, Leif Johnson, Giovanni Motta, Lillian Zhou, Françoise Beaufays, Arnaud Benard, Dhruv Guliani, Andreas Kabel, Nikhil Khare, Tamar Lucassen, Petr Zadrazil, Harry Zhang |
ASRU | 3 |
| 2011 | The iDUDE Framework for Grayscale Image DenoisingabstractWe present an extension of the discrete universal denoiser DUDE, specialized for the denoising of grayscale images. The original DUDE is a low-complexity algorithm aimed at recovering discrete sequences corrupted by discrete memoryless noise of known statistical characteristics. It is universal, in the sense of asymptotically achieving, without access to any information on the statistics of the clean sequence, the same performance as the best denoiser that does have access to such information. The DUDE, however, is not effective on grayscale images of practical size. The difficulty lies in the fact that one of the DUDE's key components is the determination of conditional empirical probability distributions of image samples, given the sample values in their neighborhood. When the alphabet is relatively large (as is the case with grayscale images), even for a small-sized neighborhood, the required distributions would be estimated from a large collection of sparse statistics, resulting in poor estimates that would not enable effective denoising. The present work enhances the basic DUDE scheme by incorporating statistical modeling tools that have proven successful in addressing similar issues in lossless image compression. Instantiations of the enhanced framework, which is referred to as iDUDE, are described for examples of additive and nonadditive noise. The resulting denoisers significantly surpass the state of the art in the case of salt and pepper (S&P) and M -ary symmetric noise, and perform well for Gaussian noise. Giovanni Motta, Erik Ordentlich, Ignacio Ramírez, Gadiel Seroussi, Marcelo J. Weinberger |
IEEE Trans. Image Process. | 1 |
| 2010 | Shape Recognition Using Vector QuantizationabstractWe present a framework to recognize objects in images based on their silhouettes. In previous work we developed translation and rotation invariant classification algorithms for textures based on Fourier transforms in the polar space followed by dimensionality reduction. Here we present a new approach to recognizing shapes by following a similar classification step with a "soft" retrieval algorithm where the search of a shape database is based on the VQ centroids found by the classification step. Experiments presented on the MPEG-7 CE-Shape 1 database show significant gains in retrieval accuracy over previous work. An interesting aspect of this recognition algorithm is that the first phase of classification seems to be a powerful tool for both texture and shape recognition. Antonella Di Lillo, Giovanni Motta, James A. Storer |
DCC | 2 |
| 2010 | Enhanced Adaptive Interpolation Filters for Video CodingabstractH.264/AVC uses motion compensated prediction with fractional-pixel precision to reduce temporal redundancy of the input video signal. It has been shown that the Adaptive Interpolation Filter (AIF) framework [3] can significantly improve accuracy of the motion compensated prediction. In this paper, we present the Enhanced Adaptive Interpolation Filters (E-AIF) scheme, which enhances the AIF framework with a number of useful features, aimed at both improving performance and reducing complexity. These features include the full-pixel position filter and the filter offset, the radial-shaped 12-position filter support, and a RD-based filter selection. Simulations show that E-AIF can achieve up to 20% bit rate reduction compared to H.264/AVC. Compared to all other AIF schemes, E-AIF further reduces the bit rate by up to 6%, and demonstrates the highest performance consistently. Giovanni Motta, Marta Karczewicz |
DCC | 2 |
| 2010 | A rotation and scale invariant descriptor for shape recognitionabstractWe address the problem of retrieving the silhouettes of objects from a database of shapes with a translation and rotation invariant feature extractor. We retrieve silhouettes by using a “soft” classification based on the Euclidean distance. Experiments show significant gains in retrieval accuracy over the existing literature. This work extends the use of our previously employed feature extractor and shows that the same descriptor can be used for both texture and shape recognition. Antonella Di Lillo, Giovanni Motta, James A. Storer |
ICIP | 2 |
| 2008 | Multiresolution Rotation-Invariant Texture Classification Using Feature Extraction in the Frequency Domain and Vector QuantizationabstractTexture identification can be a key component in content based image retrieval systems. Although formal definitions of texture vary in the literature, it is commonly accepted that textures are naturally extracted and recognized as such by the human visual system, and that this analysis is performed in the frequency domain. The vast majority of the methods proposed in the literature provide good characterization of texture in controlled environments. In order to better describe textures, features must capture the nature of the texture, invariant to rotational, shift, and scale transformations. In this work, a rotation-invariant feature extraction technique is presented, extending our previous work (A. Di Lillo et al., 2007), which was not rotation-invariant. The technique demonstrated here similarly employs a discrete Fourier transform in the polar space followed by a dimensionality reduction, but achieves rotational invariance by incorporating an additional transform into the process. Selected features are then processed with vector quantization for the classification of textures. Experiments performed on a standard test suite show that this method improves over previous methods. Antonella Di Lillo, Giovanni Motta, James A. Storer |
DCC | 2 |
| 2008 | Defect List CompressionabstractWe consider a setting relevant to the design of storage systems in which a list of defective storage blocks, as determined at manufacture time, is stored in a high-speed system controller memory to enable the efficient bypassing of defective blocks in a transparent manner, external to the system. Conventionally, such lists have been compressed losslessly to save on the cost of the controller memory. Under the assumption of a total system cost that is a linear combination of the number of storage blocks and the controller memory size, we study the potential benefits of compressing the defect list using lossy algorithms. The only restriction is that the reconstructed defect list not label any defective storage blocks as being non-defective. Giovanni Motta, Erik Ordentlich, Marcelo J. Weinberger |
DCC | 1 |
| 2008 | Defect list compressionabstractWe consider a setting relevant to the design of storage systems in which a list of defective storage blocks, as determined at manufacture time, is stored in a high-speed system controller memory to enable the efficient bypassing of defective blocks in a transparent manner, external to the system. Conventionally, such lists have been compressed losslessly to save on the cost of the controller memory. Under the assumption of a total system cost that is a linear combination of the number of storage blocks and the controller memory size, we study the potential benefits of compressing the defect list using lossy algorithms, with the restriction that the reconstructed defect list not label any defective storage blocks as being non-defective. Giovanni Motta, Erik Ordentlich, Marcelo J. Weinberger |
ISIT | 1 |
| 2007 | Texture Classification Using VQ with Feature Extraction based on Transforms Motivated by the Human Visual SystemabstractTexture identification can be a key component in CBIR (Content Based Image Recognition) systems. It can also be a tool for separation of video object planes in MPEG4 video compression systems. Although formal definitions of texture vary in the literature, it is commonly accepted that textures are naturally extracted and recognized as such by the human visual system, and that this analysis is performed in the frequency domain. Antonella Di Lillo, James A. Storer, Giovanni Motta |
DCC | 3 |
| 2007 | Differential Compression of Executable CodeabstractA platform-independent algorithm to compress file differences is presented here. Since most file updates consist of software updates and security patches, particular attention is dedicated to making this algorithm suitable to efficient compression of differences between executable files. This algorithm is designed so that its low-complexity decoder can be used in mobile and embedded devices. Compression is compared with several existing methods on a common test suite Giovanni Motta, James Gustafson, Samson Chen |
DCC | 1 |
| 2007 | Texture Classification Based on Discriminative Features Extracted in the Frequency DomainabstractTexture identification can be a key component in Content Based Image Retrieval systems. Although formal definitions of texture vary in the literature, it is commonly accepted that textures are naturally extracted and recognized as such by the human visual system, and that this analysis is performed in the frequency domain. In this work, a feature extraction method is presented which employs a discrete Fourier transform in the polar space, followed by a dimensionality reduction. Selected features are then processed with vector quantization for the supervised segmentation of images into uniformly textured regions. Experiments performed on a standard test suite show that this method compares favorably to the state-of-the-art and improves over previously studied frequency-domain based methods. Antonella Di Lillo, Giovanni Motta, James A. Storer |
ICIP (2) | 2 |
| 2005 | The DUDE framework for continuous tone image denoisingabstractThis paper discusses the challenges of applying the DUDE framework to continuous tone images and the tools used to address these challenges. As in lossless image compression, a key component of the DUDE framework is the determination of a probability distribution for samples of the input (noisy) image, conditioned on their contexts. Thus, we can leverage from tools developed and tested in the context of lossless compression for determining such distributions, together with tools that are specific to the assumptions of the denoising application. These tools combine with the DUDE principles into a framework that yields powerful and practical denoisers for continuous tone images corrupted by a variety of noise processes. Gadiel Seroussi, Giovanni Motta, Erik Ordentlich, Ignacio Ramírez, Marcelo J. Weinberger |
ICIP (3) | 2 |
| 2005 | Low-complexity lossless compression of hyperspectral imagery via linear predictionabstractWe present a new low-complexity algorithm for hyperspectral image compression that uses linear prediction in the spectral domain. We introduce a simple heuristic to estimate the performance of the linear predictor from a pixel spatial context and a context modeling mechanism with one-band look-ahead capability, which improves the overall compression with marginal usage of additional memory. The proposed method is suitable to spacecraft on-board implementation, where limited hardware and low power consumption are key requirements. Finally, we present a least-squares optimized linear prediction technique that achieves better compression on data cubes acquired by the NASA JPL Airborne Visible/Infrared Imaging Spectrometer (AVIRIS). Francesco Rizzo, Bruno Carpentieri, Giovanni Motta, James A. Storer |
IEEE Signal Process. Lett. | 3 |
| 2004 | High Performance Compression of Hyperspectral Imagery with Reduced Search Complexity in the Compressed DomainabstractIn previous work we considered LPVQ, a compression algorithm based on locally optimal partitioned vector quantization that can be used to compress hyperspectral images by applying partitioned VQ to the spectral signatures (e.g., to the 224 16-bit values of a NASA AVIRIS pixel) and then encoding error information with a threshold that can be varied from high quality lossy to near lossless to lossless (e.g., 50-to-1 lossy, 10-to-1 near lossless, or 3-to-1 lossless). An advantage of LPVQ is extremely fast decoding (table lookup followed by entropy decoding), but it is at the cost of more complex encoding. Here we present a new low complexity algorithm for hyperspectral image compression, called SLSQ, that employs linear prediction targeted at spectral correlation followed by entropy coding of the prediction error. We then consider how SLSQ can be combined with LPVQ in a scenario commonly arising in practice. In this scenario, a low-complexity lossless encoder on the remote acquisition platform compresses the data for transmission to a central computing facility, where it is processed and re-coded using LPVQ, so that the compressed data can be distributed to the final users at various quality levels. The VQ indices of the LPVQ form a lossy compressed image of only about 2% of the original size; this small image can be employed to greatly reduce the time for browsing and classification. Francesco Rizzo, Bruno Carpentieri, Giovanni Motta, James A. Storer |
Data Compression Conference | 3 |
| 2003 | Compression of Hyperspectral ImageryabstractHigh dimensional source vectors, such as those that occur in hyperspectral imagery, are partitioned into a number of subvectors of different length and then each subvector is vector quantized (VQ) individually with an appropriate codebook. A locally adaptive partitioning algorithm is introduced that performs comparably in this application to a more expensive globally optimal one that employs dynamic programming. The VQ indices are entropy coded and used to condition the lossless or near-lossless coding of the residual error. Motivated by the need for maintaining uniform quality across all vector components, a percentage maximum absolute error distortion measure is employed. Experiments on the lossless and near-lossless compression of NASA AVIRIS images are presented. A key advantage of the approach is the use of independent small VQ codebooks that allow fast encoding and decoding. Giovanni Motta, Francesco Rizzo, James A. Storer |
DCC | 1 |
| 2003 | Partitioned vector quantization: application to lossless compression of hyperspectral imagesabstractA novel design for a vector quantizer that uses multiple codebooks of variable dimensionality is proposed. High dimensional source vectors are first partitioned into two or more subvectors of (possibly) different length and then, each subvector is individually encoded with an appropriate codebook. Further redundancy is exploited by conditional entropy coding of the subvectors indices. This scheme allows practical quantization of high dimensional vectors in which each vector component is allowed to have different alphabet and distribution. This is typically the case of the pixels representing a hyperspectral image. We present experimental results in the lossless and near-lossless encoding of such images. The method can be easily adapted to lossy coding. Giovanni Motta, Francesco Rizzo, James A. Storer |
ICASSP (3) | 1 |
| 2003 | Partitioned vector quantization: application to lossless compression of hyperspectral imagesabstractA novel design for a vector quantizer that uses multiple codebooks of variable dimensionality is proposed. High dimensional source vectors are first partitioned into two or more subvectors of (possibly) different length and then, each subvector is individually encoded with an appropriate codebook. Further redundancy is exploited by conditional entropy coding of the subvectors indices. This scheme allows practical quantization of high dimensional vectors in which each vector component is allowed to have different alphabet and distribution. This is typically the case of the pixels representing a hyperspectral image. We present experimental results in the lossless and near-lossless encoding of such images. The method can be easily adapted to lossy coding. Giovanni Motta, Francesco Rizzo, James A. Storer |
ICME | 1 |
| 2000 | Improving Scene Cut Quality for Real-Time Video DecodingabstractWe address the problem of improving the scene cut quality in fixed bit-rate real-time video decoding such as is used in the H.263 and MPEG standards. In low bandwidth applications, scene cuts can cause the bits required to encode a single frame to greatly exceed the target average bits per frame, and necessitate the skipping of other frames to provide sufficient time to transmit the scene cut frame. We present an optimal algorithm for minimizing the number of skipped frames and keep the decoding synchronized. Although the algorithm requires additional encoding complexity, there is no change in decoding complexity (in fact, no change to the decoder at all). Experimental results, obtained with a simplified strategy within the framework of H.263+ video encoding, confirm that the method provides an effective alternative to current frame skipping strategies. The overall quality in the presence of scene cuts is improved with respect the TMN-8 rate control. Although the overall bit rate benefits from our method, our focus is to improve the quality of the video where scene cuts occur (by reducing skipped frames and improving decoder synchronization). The approach here can be combined with more sophisticated rate controls, as, for example, the newer rate-distortion optimized TMN-10 and TMN-11. Giovanni Motta, James A. Storer, Bruno Carpentieri |
Data Compression Conference | 1 |
| 2000 | Lossless image coding via adaptive linear prediction and classificationabstractIn past years, there have been several improvements in lossless image compression. All the recently proposed state-of-the-art lossless image compressors can be roughly divided into two categories: single and double-pass compressors. Linear prediction is rarely used in the first category, while TMW, a state-of-the-art double-pass image compressor, relies on linear prediction for its performance. We propose a single-pass adaptive algorithm that uses context classification and multiple linear predictors, locally optimized on a pixel-by-pixel basis. Locality is also exploited in the entropy coding of the prediction error. The results we obtained on a test set of several standard images are encouraging. On the average, our ALPC obtains a compression ratio comparable to CALIC while improving on some images. Giovanni Motta, James A. Storer, Bruno Carpentieri |
Proc. IEEE | 1 |
| 1999 | Adaptive Linear Prediction Lossless Image CodingabstractThe practical lossless digital image compressors that achieve the best results in terms of compression ratio are also simple and fast algorithms with low complexity both in terms of memory usage and running time. Surprisingly, the compression ratio achieved by these systems cannot be substantially improved even by using image-by-image optimization techniques or more sophisticate and complex algorithms. Meyer and Tischer (1998) were able, with their TMW, to improve some current best results (they do not report results for all test images) by using global optimization techniques and multiple blended linear predictors. Our investigation is directed to determine the effectiveness of an algorithm that uses multiple adaptive linear predictors, locally optimized on a pixel-by-pixel basis. The results we obtained on a test set of nine standard images are encouraging, where we improve over CALIC on some images. Giovanni Motta, James A. Storer, Bruno Carpentieri |
Data Compression Conference | 1 |
| 1997 | A new trellis vector residual quantizer: applications to image codingabstractWe present a new trellis coded vector residual quantizer (TCVRQ) that combines trellis coding and vector residual quantization. We propose new methods for computing quantization levels and experimentally analyze the performances of our TCVRQ in the case of still image coding. Experimental comparisons show that our quantizer performs better than the standard tree and exhaustive search quantizers based on the generalized Lloyd algorithm (GLA). Giovanni Motta, Bruno Carpentieri |
ICASSP | 1 |