VLDB 2026 Research / reviewers in the wild / expert
Carlo Luschi
dblp:72/10621
· DBLP profile ↗
13ranked-venue papers
3as first author
6since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 5 since 2021Computer networks · 4 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
Deep learning architectures and training · 48% Learning theory · 22% Efficient and distributed learning · 18% | |
| Interdisciplinary, comprehensive, and emerging computing
2 papers |
Bioinformatics and computational biology · 57% Computational science and engineering · 43% |
Topics — the 21 heaviest of 25, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Learning theory › neural network theory › neural network parameterization
maximal update parameterization |
0.9 | 1 | 2025 | u-μP: The Unit-Scaled Maximal Update Parametrization · ICLR 2025 |
Machine learning › Learning theory › neural network theory
neural network parameterization |
0.9 | 1 | 2025 | u-μP: The Unit-Scaled Maximal Update Parametrization · ICLR 2025 |
Machine learning › Deep learning architectures and training
attention mechanism |
0.8 | 1 | 2024 | SparQ Attention: Bandwidth-Efficient LLM Inference · ICML 2024 |
Machine learning › Efficient and distributed learning
model compression |
0.8 | 1 | 2024 | SparQ Attention: Bandwidth-Efficient LLM Inference · ICML 2024 |
Machine learning › Deep learning architectures and training › attention mechanism
sparse attention |
0.8 | 1 | 2024 | SparQ Attention: Bandwidth-Efficient LLM Inference · ICML 2024 |
Machine learning › Efficient and distributed learning
low-precision training |
0.7 | 1 | 2023 | Unit Scaling: Out-of-the-Box Low-Precision Training · ICML 2023 |
Machine learning › Deep learning architectures and training
weight initialization |
0.7 | 1 | 2023 | Unit Scaling: Out-of-the-Box Low-Precision Training · ICML 2023 |
Computational science and engineering › computational chemistry
quantum chemistry |
0.7 | 1 | 2023 | Generating QM1B with PySCFIPU · NeurIPS 2023 |
Machine learning › Deep learning architectures and training
normalization |
0.5 | 1 | 2021 | Proxy-Normalizing Activations to Match Batch Normalization while Removing Batch Dependence · NeurIPS 2021 |
Machine learning › Optimization for machine learning
large-scale optimization |
0.4 | 1 | 2020 | Improving Neural Network Training in Low Dimensional Random Bases · NeurIPS 2020 |
Machine learning › Optimization for machine learning
stochastic gradient descent |
0.4 | 1 | 2020 | Improving Neural Network Training in Low Dimensional Random Bases · NeurIPS 2020 |
Physical-layer communications
equalization |
0.0 | 1 | 2003 | Nonparametric trellis equalization in the presence of non-Gaussian interference · IEEE Trans. Commun. 2003 |
Physical-layer communications › equalization
trellis-based equalization |
0.0 | 1 | 2003 | Nonparametric trellis equalization in the presence of non-Gaussian interference · IEEE Trans. Commun. 2003 |
Physical-layer communications › equalization
blind equalization |
0.0 | 2 | 1996 | Joint clock recovery and baseband combining for the diversity radio channel · IEEE Trans. Commun. 1996 Joint Clock Recovery and Baseband Combining for the Diversity Radio Channel · IEEE Trans. Commun. 1995 |
Physical-layer communications
signal processing for communications |
0.0 | 2 | 1996 | Joint clock recovery and baseband combining for the diversity radio channel · IEEE Trans. Commun. 1996 Joint Clock Recovery and Baseband Combining for the Diversity Radio Channel · IEEE Trans. Commun. 1995 |
Physical-layer communications
channel coding |
0.0 | 1 | 2000 | Advanced signal-processing algorithms for energy-efficient wireless communications · Proc. IEEE 2000 |
Physical-layer communications › channel coding › decoding algorithms › iterative decoding
turbo decoding |
0.0 | 1 | 2000 | Advanced signal-processing algorithms for energy-efficient wireless communications · Proc. IEEE 2000 |
Physical-layer communications › interference suppression
cochannel interference |
0.0 | 1 | 2003 | Nonparametric trellis equalization in the presence of non-Gaussian interference · IEEE Trans. Commun. 2003 |
Physical-layer communications › interference › interference analysis
non-gaussian interference |
0.0 | 1 | 2003 | Nonparametric trellis equalization in the presence of non-Gaussian interference · IEEE Trans. Commun. 2003 |
Physical-layer communications › channel modeling
channel impairment |
0.0 | 2 | 1996 | Joint clock recovery and baseband combining for the diversity radio channel · IEEE Trans. Commun. 1996 Joint Clock Recovery and Baseband Combining for the Diversity Radio Channel · IEEE Trans. Commun. 1995 |
Physical-layer communications › fading channels
multipath fading |
0.0 | 2 | 1996 | Joint clock recovery and baseband combining for the diversity radio channel · IEEE Trans. Commun. 1996 Joint Clock Recovery and Baseband Combining for the Diversity Radio Channel · IEEE Trans. Commun. 1995 |
Methods — techniques the papers use, named apart from their topics
PySCF · 1.3IPU · 1.3knowledge graph embedding · 0.9selective fetching · 0.8KV cache · 0.8unit scaling · 0.7FP8 training · 0.7layer normalization · 0.5instance normalization · 0.5group normalization · 0.5pseudo-random number generation · 0.4whitening filter · 0.0probability density estimation · 0.0kernel smoothing · 0.0constant modulus algorithm · 0.0clock recovery · 0.0turbo processing · 0.0maximum a posteriori estimation · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | u-μP: The Unit-Scaled Maximal Update Parametrization
Charlie Blake, Constantin Eichenberg, Josef Dean, Lukas Balles, Luke Y. Prince, Björn Deiseroth, Andrés Felipe Cruz-Salinas, Carlo Luschi, Samuel Weinbach, Douglas Orr |
ICLR | 8 |
| 2025 | The role of graph topology in the performance of biomedical Knowledge Graph Completion modelsabstractMOTIVATION: Knowledge Graph Completion has been increasingly adopted as a useful method for helping address several tasks in biomedical research, such as drug repurposing or drug-target identification. To that end, a variety of datasets and Knowledge Graph Embedding models have been proposed over the years. However, little is known about the properties that render a dataset, and associated modelling choices, useful for a given task. Moreover, even though theoretical properties of Knowledge Graph Embedding models are well understood, their practical utility in this field remains controversial. RESULTS: In this work, we conduct a comprehensive investigation into the topological properties of publicly available biomedical Knowledge Graphs and establish links to the accuracy observed in real-world tasks. By releasing all model predictions and a new suite of analysis tools we invite the community to build upon our work and continue improving the understanding of these crucial applications. AVAILABILITY AND IMPLEMENTATION: The code used to perform experiments and analyze results in this article as well as all experimental data is available at https://github.com/graphcore-research/kg-topology-toolbox/tree/main/the_role_of_graph_topology_paper and archived on Zenodo, at https://doi.org/10.5281/zenodo.12097376. Alberto Cattaneo, Stephen Bonner, Thomas Martynec, Edward R. Morrissey, Carlo Luschi, Ian P. Barrett, Daniel Justus |
Bioinform. | 5 |
| 2024 | SparQ Attention: Bandwidth-Efficient LLM InferenceabstractThe computational difficulties of large language model (LLM) inference remain a significant obstacle to their widespread deployment. The need for many applications to support long input sequences and process them in large batches typically causes token-generation to be bottlenecked by data transfer. For this reason, we introduce **SparQ Attention**, a technique for increasing the inference throughput of LLMs by utilising memory bandwidth more efficiently within the attention layers, through selective fetching of the cached history. Our proposed technique can be applied directly to off-the-shelf LLMs during inference, without requiring any modification to the pre-training setup or additional fine-tuning. We show that SparQ Attention brings up to 8x savings in attention data transfers without substantial drops in accuracy, by evaluating Llama 2 and 3, Mistral, Gemma and Pythia models on a wide range of downstream tasks. Luka Ribar, Ivan Chelombiev, Luke Hudlass-Galley, Charlie Blake, Carlo Luschi, Douglas Orr |
ICML | 5 |
| 2023 | Unit Scaling: Out-of-the-Box Low-Precision TrainingabstractWe present unit scaling, a paradigm for designing deep learning models that simplifies the use of low-precision number formats. Training in FP16 or the recently proposed FP8 formats offers substantial efficiency gains, but can lack sufficient range for out-of-the-box training. Unit scaling addresses this by introducing a principled approach to model numerics: seeking unit variance of all weights, activations and gradients at initialisation. Unlike alternative methods, this approach neither requires multiple training runs to find a suitable scale nor has significant computational overhead. We demonstrate the efficacy of unit scaling across a range of models and optimisers. We further show that existing models can be adapted to be unit-scaled, training BERT-Large in FP16 and then FP8 with no degradation in accuracy. Charlie Blake, Douglas Orr, Carlo Luschi |
ICML | 3 |
| 2023 | Generating QM1B with PySCFIPU
Alexander Mathiasen, Hatem Helal, Kerstin Kläser 0001, Paul Balanca, Josef Dean, Carlo Luschi, Dominique Beaini, Andrew W. Fitzgibbon, Dominic Masters |
NeurIPS | 6 |
| 2021 | Proxy-Normalizing Activations to Match Batch Normalization while Removing Batch DependenceabstractWe investigate the reasons for the performance degradation incurred with batch-independent normalization. We find that the prototypical techniques of layer normalization and instance normalization both induce the appearance of failure modes in the neural network's pre-activations: (i) layer normalization induces a collapse towards channel-wise constant functions; (ii) instance normalization induces a lack of variability in instance statistics, symptomatic of an alteration of the expressivity. To alleviate failure mode (i) without aggravating failure mode (ii), we introduce the technique "Proxy Normalization" that normalizes post-activations using a proxy distribution. When combined with layer normalization or group normalization, this batch-independent normalization emulates batch normalization's behavior and consistently matches or exceeds its performance. Antoine Labatie, Dominic Masters, Zach Eaton-Rosen, Carlo Luschi |
NeurIPS | 4 |
| 2020 | Improving Neural Network Training in Low Dimensional Random BasesabstractStochastic Gradient Descent (SGD) has proven to be remarkably effective in optimizing deep neural networks that employ ever-larger numbers of parameters. Yet, improving the efficiency of large-scale optimization remains a vital and highly active area of research. Recent work has shown that deep neural networks can be optimized in randomly-projected subspaces of much smaller dimensionality than their native parameter space. While such training is promising for more efficient and scalable optimization schemes, its practical application is limited by inferior optimization performance. Here, we improve on recent random subspace approaches as follows. We show that keeping the random projection fixed throughout training is detrimental to optimization. We propose re-drawing the random subspace at each step, which yields significantly better performance. We realize further improvements by applying independent projections to different parts of the network, making the approximation more efficient as network dimensionality grows. To implement these experiments, we leverage hardware-accelerated pseudo-random number generation to construct the random projections on-demand at every optimization step, allowing us to distribute the computation of independent random directions across multiple workers with shared random seeds. This yields significant reductions in memory and is up to 10x faster for the workloads in question. Frithjof Gressmann, Zach Eaton-Rosen, Carlo Luschi |
NeurIPS | 3 |
| 2005 | Estimation of the output error statistics of space-time equalization in an antenna array EGPRS receiver with soft-decision decodingabstractThe use of antenna arrays can help combat cochannel interference (CCI) in wireless cellular systems. In this paper, we consider an enhanced general packet radio service diversity receiver based on least squares spatio-temporal equalization and soft-decision decoding in the presence of decision feedback and/or asynchronous CCI. We compare known and novel estimators of the error mean and variance at the output of the deterministic space-time equalizer. The collected simulation data indicate that the estimation of the error mean and variance is critical to the performance of soft-in/hard-out Viterbi decoding in the presence of nonstationary input disturbance. Moreover, the use of short-term error statistics provides receiver performance gains of up to 15-20 dB in terms of signal-to-interference ratio, with respect to the use of burst statistics based on the training sequence midamble and tentative decisions on the payload symbols. Carlo Luschi, Bernard Mulgrew |
IEEE Trans. Wirel. Commun. | 1 |
| 2003 | Nonparametric trellis equalization in the presence of non-Gaussian interferenceabstractWe consider the problem of trellis equalization of the intersymbol interference channel in the presence of thermal noise and cochannel interference (CCI). Conventional maximum-likelihood sequence estimation (MLSE) and maximum a posteriori probability (MAP) trellis equalizers treat the sum of noise and interference as additive white Gaussian noise, while CCI is generally a colored non-Gaussian process. We propose a novel nonparametric approach based on the estimation of the probability density function of the noise-plus-interference. Given the availability of a limited volume of data, the density is estimated by kernel-smoothing techniques. The use of a whitening filter in the presence of temporally colored disturbance is also addressed. Simulation results are provided for the global system for mobile communications (GSM), showing a significant performance improvement with respect to the equalizer based on the Gaussian assumption. Major advantages of the proposed strategy are its intrinsic robustness and general applicability to those cases where accurate modeling of the interference is difficult or a model is not available. Carlo Luschi, Bernard Mulgrew |
IEEE Trans. Commun. | 1 |
| 2000 | Advanced signal-processing algorithms for energy-efficient wireless communicationsabstractSubstantial progress has been made in the receiver signal-processing algorithms for wireless communications to minimize the requirements on signal-to-noise (and/or interference) power ratio and computational complexities for the same quality of service. In cellular infrastructure systems, one of the key system design objectives in the base stations is to maximize the receiver sensitivity, so that the required signal level from the mobile stations can be minimized. The use of advance signal-processing algorithms, based on maximum a posteriori (MAP) estimation, iterative (turbo) channel estimation, equalization, and decoding, allows for a reduction of the required transmitter power by one-third to one-half. Lower computational complexities in the terminals, which implies a reduced power drain on the digital circuits, can be achieved by using techniques that adapt the state complexity of the receiver to the propagation channel. We give an in-depth review of these algorithms, and discuss their performance and implementation requirements. Carlo Luschi, Magnus Sandell, Paul Strauch, Jian-Jun Wu, Costel Ilas, Ping-Wen Ong, Romain Baeriswyl, Frederic Battaglia, Spyros Karageorgis |
Proc. IEEE | 1 |
| 1999 | Low complexity source controlled channel decoding in a GSM systemabstractIn this paper we investigate source controlled channel decoding with a hard output channel decoder. Various methods have been devised in the past for source controlled channel decoding, but most of them assume that a soft output channel decoder is used. Most receivers in mobile wireless communications have a standard Viterbi channel decoder which produces only hard outputs. It is shown that a simple sliding histogram is capable of improving the speech quality significantly. The ideas and methods in this paper are applied for the full-rate and enhanced full-rate speech codecs in the GSM system. Paul Strauch, Carlo Luschi, Magnus Sandell |
ICASSP | 2 |
| 1996 | Joint clock recovery and baseband combining for the diversity radio channelabstractMultipath fading is one of the major impairments encountered in terrestrial digital radio. A common countermeasure to limit the outage time due to multipath is the space diversity technique, which takes its effectiveness from the low correlation between the field samples of two well separated antennas. A very simple and effective blind receiver is proposed for the diversity radio channel. The novelty of the present study is the application of the constant modulus algorithm to joint clock recovery and baseband combining. The effectiveness of our proposal relies upon the synergic action of clock recovery and adaptive baseband combining, which allows optimal equalization of the two-ray diversity channel. Franco Guglielmi, Carlo Luschi, Arnaldo Spalvieri |
IEEE Trans. Commun. | 2 |
| 1995 | Joint Clock Recovery and Baseband Combining for the Diversity Radio ChannelabstractMultipath fading is one of the major impairments encountered in terrestrial digital radio. A common countermeasure to limit the outage time due to multipath is the space diversity technique, which takes its effectiveness from the low correlation between the field samples of two well separated antennas. A very simple and effective blind receiver is proposed for the diversity radio channel. The novelty of the present study is the application of the constant modulus algorithm to joint clock recovery and baseband combining. The effectiveness of our proposal relies upon the synergic action of clock recovery and adaptive baseband combining, which allows optimal equalization of the two-ray diversity channel. Franco Guglielmi, Carlo Luschi, Arnaldo Spalvieri |
IEEE Trans. Commun. | 2 |