Cristiano Malossi

dblp:49/9752 · also A. Cristiano I. Malossi, Adelmo Cristiano Innocenza Malossi · DBLP profile ↗
← Back
13ranked-venue papers
1as first author
5since 2021 · last 2024
0000-0002-8201-1533ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 4 since 2021Systems, architecture and hardware · 5 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 5 · 4 since 2021Software engineering, systems software and programming languages · 1 · 1 first-authorTheory of computation · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Efficient and distributed learning · 44% Vision and language · 34% 3D vision · 10%
Computer architecture, parallel and distributed computing, and storage systems
2 papers
High-performance computing · 63% Hardware accelerators and domain-specific architectures · 37%
Theoretical computer science
1 paper
Mathematical optimization · 100%

Topics — the 12 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Vision and language
visual prompting
0.812024
Probabilistic Feature Matching for Fast Scalable Visual Prompting · IJCAI 2024
Machine learning › Efficient and distributed learning
automated machine learning
0.512021
AutoText: An End-to-End AutoAI Framework for Text · AAAI 2021
Machine learning › Efficient and distributed learning › automated machine learning
neural architecture search
0.412019
TAPAS: Train-Less Accuracy Predictor for Architecture Search · AAAI 2019
Hardware accelerators and domain-specific architectures › machine learning accelerator
neural network accelerator
0.412019
Constrained deep neural network architecture search for IoT devices accounting for hardware calibration · NeurIPS 2019
Computer vision › 3D vision
feature matching
0.212024
Probabilistic Feature Matching for Fast Scalable Visual Prompting · IJCAI 2024
High-performance computing › performance optimization at scale
extreme-scale scalability
0.212015
An extreme-scale implicit solver for complex PDEs: highly heterogeneous flow in earth's mantle · SC 2015
High-performance computing
performance optimization at scale
0.212015
An extreme-scale implicit solver for complex PDEs: highly heterogeneous flow in earth's mantle · SC 2015
High-performance computing
scientific computing systems
0.212015
An extreme-scale implicit solver for complex PDEs: highly heterogeneous flow in earth's mantle · SC 2015
Computer vision › Image recognition and object detection
image classification
0.112019
TAPAS: Train-Less Accuracy Predictor for Architecture Search · AAAI 2019
Machine learning › Efficient and distributed learning
model compression
0.112019
Constrained deep neural network architecture search for IoT devices accounting for hardware calibration · NeurIPS 2019
Mathematical optimization › numerical analysis › multigrid methods
algebraic multigrid
0.112015
An extreme-scale implicit solver for complex PDEs: highly heterogeneous flow in earth's mantle · SC 2015
Mathematical optimization › numerical analysis
multigrid methods
0.112015
An extreme-scale implicit solver for complex PDEs: highly heterogeneous flow in earth's mantle · SC 2015

Methods — techniques the papers use, named apart from their topics

neural architecture search · 1.3quantization · 0.8probabilistic feature matching · 0.8hyperparameter optimization · 0.5schur-complement preconditioning · 0.4multi-octree adaptivity · 0.4mixed continuous-discontinuous discretization · 0.4deep neural network · 0.4architecture search · 0.4
YearPublicationVenuePosition
2024 Beyond Static and Dynamic Quantization - Hybrid Quantization of Vision Transformers
Piotr Kluska, Florian Scheidegger, Cristiano Malossi, Enrique S. Quintana-Ortí
BMVC3
2024 Combining Data Generation and Active Learning for Low-Resource Question Answering
Maximilian Kimmich, Andrea Bartezzaghi, Jasmina Bogojeska, Cristiano Malossi, Ngoc Thang Vu
ICANN (7)4
2024 Probabilistic Feature Matching for Fast Scalable Visual Prompting
Thomas Frick, Cezary Skura, Filip Janicki, Roy Assaf, Niccolò Avogaro, Daniel Caraballo, Yagmur Gizem Cinar, Brown Ebouky, Ioana Giurgiu, Takayuki Katsuki, Piotr Kluska, Cristiano Malossi, Haoxiang Qiu, Florian Scheidegger, Andrej Simeski, Daniel Yang, Andrea Bartezzaghi, Mattia Rigotti
IJCAI12
2021 AutoText: An End-to-End AutoAI Framework for Text
abstract
Building models for natural language processing (NLP) tasks remains a daunting task for many, requiring significant technical expertise, efforts, and resources. In this demonstration, we present AutoText, an end-to-end AutoAI framework for text, to lower the barrier of entry in building NLP models. AutoText combines state-of-the-art AutoAI optimization techniques and learning algorithms for NLP tasks into a single extensible framework. Through its simple, yet powerful UI, non-AI experts (e.g., domain experts) can quickly generate performant NLP models with support to both control (e.g., via specifying constraints) and understand learned models.
Arunima Chaudhary, Alayt Issak, Kiran Kate, Yannis Katsis, Abel N. Valente, Dakuo Wang, Alexandre V. Evfimievski, Sairam Gurajada, Ban Kawas, Cristiano Malossi, Lucian Popa 0001, Tejaswini Pedapati, Horst Samulowitz, Martin Wistuba, Yunyao Li 0001
AAAI10
2021 Efficient image dataset classification difficulty estimation for predicting deep-learning accuracy
abstract
Abstract In the deep-learning community, new algorithms are published at a very fast pace. Therefore, solving an image classification problem for new datasets becomes a challenging task, as it requires to re-evaluate published algorithms and their different configurations in order to find a close to optimal classifier. To facilitate this process, before biasing our decision toward a class of neural networks or running an expensive search over the network space, we propose to estimate the classification difficulty of the dataset. Our method computes a single number that characterizes the dataset difficulty $$97\times $$ 97 × faster than training state-of-the-art networks. The proposed method can be used in combination with network topology and hyper-parameter search optimizers to efficiently drive the search toward promising neural network configurations.
Florian Scheidegger, Roxana Istrate, Giovanni Mariani, Luca Benini, Costas Bekas, Cristiano Malossi
Vis. Comput.6
2019 TAPAS: Train-Less Accuracy Predictor for Architecture Search
abstract
In recent years an increasing number of researchers and practitioners have been suggesting algorithms for large-scale neural network architecture search: genetic algorithms, reinforcement learning, learning curve extrapolation, and accuracy predictors. None of them, however, demonstrated highperformance without training new experiments in the presence of unseen datasets. We propose a new deep neural network accuracy predictor, that estimates in fractions of a second classification performance for unseen input datasets, without training. In contrast to previously proposed approaches, our prediction is not only calibrated on the topological network information, but also on the characterization of the dataset-difficulty which allows us to re-tune the prediction without any training. Our predictor achieves a performance which exceeds 100 networks per second on a single GPU, thus creating the opportunity to perform large-scale architecture search within a few minutes. We present results of two searches performed in 400 seconds on a single GPU. Our best discovered networks reach 93.67% accuracy for CIFAR-10 and 81.01% for CIFAR-100, verified by training. These networks are performance competitive with other automatically discovered state-of-the-art networks however we only needed a small fraction of the time to solution and computational resources.
Roxana Istrate, Florian Scheidegger, Giovanni Mariani, Dimitrios S. Nikolopoulos, Costas Bekas, Cristiano Malossi
AAAI6
2019 Constrained deep neural network architecture search for IoT devices accounting for hardware calibration
abstract
Deep neural networks achieve outstanding results for challenging image classification tasks. However, the design of network topologies is a complex task, and the research community is conducting ongoing efforts to discover top-accuracy topologies, either manually or by employing expensive architecture searches. We propose a unique narrow-space architecture search that focuses on delivering low-cost and rapidly executing networks that respect strict memory and time requirements typical of Internet-of-Things (IoT) near-sensor computing platforms. Our approach provides solutions with classification latencies below 10~ms running on a low-cost device with 1~GB RAM and a peak performance of 5.6~GFLOPS. The narrow-space search of floating-point models improves the accuracy on CIFAR10 of an established IoT model from 70.64% to 74.87% within the same memory constraints. We further improve the accuracy to 82.07% by including 16-bit half types and obtain the highest accuracy of 83.45% by extending the search with model-optimized IEEE 754 reduced types. To the best of our knowledge, this is the first empirical demonstration of more than 3000 trained models that run with reduced precision and push the Pareto optimal front by a wide margin. Within a given memory constraint, accuracy is improved by more than 7% points for half and more than 1% points for the best individual model format.
Florian Scheidegger, Luca Benini, Costas Bekas, Cristiano Malossi
NeurIPS4
2019 FloatX: A C++ Library for Customized Floating-Point Arithmetic
abstract
We present FloatX (Float eXtended), a C ++ framework to investigate the effect of leveraging customized floating-point formats in numerical applications. FloatX formats are based on binary IEEE 754 with smaller significand and exponent bit counts specified by the user. Among other properties, FloatX facilitates an incremental transformation of the code, relies on hardware-supported floating-point types as back-end to preserve efficiency, and incurs no storage overhead. The article discusses in detail the design principles, programming interface, and datatype casting rules behind FloatX. Furthermore, it demonstrates FloatX’s usage and benefits via several case studies from well-known numerical dense linear algebra libraries, such as BLAS and LAPACK; the Ginkgo library for sparse linear systems; and two neural network applications related with image processing and text recognition.
Goran Flegar, Florian Scheidegger, Vedran Novakovic, Giovanni Mariani, Andrés Tomás, Cristiano Malossi, Enrique S. Quintana-Ortí
ACM Trans. Math. Softw.6
2018 The transprecision computing paradigm: Concept, design, and applications
abstract
Guaranteed numerical precision of each elementary step in a complex computation has been the mainstay of traditional computing systems for many years. This era, fueled by Moore's law and the constant exponential improvement in computing efficiency, is at its twilight: from tiny nodes of the Internet-of-Things, to large HPC computing centers, sub-picoJoule/operation energy efficiency is essential for practical realizations. To overcome the power wall, a shift from traditional computing paradigms is now mandatory. In this paper we present the driving motivations, roadmap, and expected impact of the European project OPRECOMP. OPRECOMP aims to (i) develop the first complete transprecision computing framework, (ii) apply it to a wide range of hardware platforms, from the sub-milliWatt up to the MegaWatt range, and (iii) demonstrate impact in a wide range of computational domains, spanning IoT, Big Data Analytics, Deep Learning, and HPC simulations. By combining together into a seamless design transprecision advances in devices, circuits, software tools, and algorithms, we expect to achieve major energy efficiency improvements, even when there is no freedom to relax end-to-end application quality of results. Indeed, OPRECOMP aims at demolishing the ultra-conservative “precise” computing abstraction, replacing it with a more flexible and efficient one, namely transprecision computing.
Cristiano Malossi, Michael Schaffner, Anca Mariana Molnos, Luca Gammaitoni, Giuseppe Tagliavini, Andrew P. J. Emerson, Andrés Tomás, Dimitrios S. Nikolopoulos, Eric Flamand, Norbert Wehn
DATE1
2018 A scalable iterative dense linear system solver for multiple right-hand sides in data analytics
Vassilis Kalantzis, Cristiano Malossi, Costas Bekas, Alessandro Curioni, Efstratios Gallopoulos, Yousef Saad
Parallel Comput.2
2016 The Impact of Voltage-Frequency Scaling for the Matrix-Vector Product on the IBM POWER8
Sandra Catalán, Cristiano Malossi, Costas Bekas, Enrique S. Quintana-Ortí
Euro-Par2
2016 Stochastic Matrix-Function Estimators: Scalable Big-Data Kernels with High Performance
abstract
In this era of Big Data, large graphs appear in many scientific domains. To extract the hidden knowledge/correlations in these graphs, novel methods need to be developed to analyse these graphs fast. In this paper, we present a unified framework of stochastic matrix-function estimators, which allows one to compute a subset of elements of the matrix f(A), where f is an arbitrary function and A is the adjacency matrix of the graph. The new framework has a computational cost proportional to the size of the subset, i.e. to obtain the diagonal of f(A) with matrix-size N, the computational cost is proportional to N contrary to the traditional N^3 from diagonalization. Furthermore, we will show that the new framework allows us to write implementations of the algorithm that scale naturally with the number of compute nodes and is easily ported to accelerators where the kernels perform very well.
Peter W. J. Staar, Panagiotis Kl. Barkoutsos, Roxana Istrate, Cristiano Malossi, Ivano Tavernelli, Nikolaj Moll, Heiner Giefers, Christoph Hagleitner, Costas Bekas, Alessandro Curioni
IPDPS4
2015 An extreme-scale implicit solver for complex PDEs: highly heterogeneous flow in earth's mantle
abstract
Mantle convection is the fundamental physical process within earth's interior responsible for the thermal and geological evolution of the planet, including plate tectonics. The mantle is modeled as a viscous, incompressible, non-Newtonian fluid. The wide range of spatial scales, extreme variability and anisotropy in material properties, and severely nonlinear rheology have made global mantle convection modeling with realistic parameters prohibitive. Here we present a new implicit solver that exhibits optimal algorithmic performance and is capable of extreme scaling for hard PDE problems, such as mantle convection. To maximize accuracy and minimize runtime, the solver incorporates a number of advances, including aggressive multi-octree adaptivity, mixed continuous-discontinuous discretization, arbitrarily-high-order accuracy, hybrid spectral/geometric/algebraic multigrid, and novel Schur-complement preconditioning. These features present enormous challenges for extreme scalability. We demonstrate that---contrary to conventional wisdom---algorithmically optimal implicit solvers can be designed that scale out to 1.5 million cores for severely nonlinear, ill-conditioned, heterogeneous, and anisotropic PDEs.
Johann Rudi, Cristiano Malossi, Tobin Isaac, Georg Stadler, Michael Gurnis, Peter W. J. Staar, Yves Ineichen, Costas Bekas, Alessandro Curioni, Omar Ghattas
SC2