Daniele Moro

dblp:204/5584 · DBLP profile ↗
← Back
8ranked-venue papers
0as first author
5since 2021 · last 2024
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer networks
1 paper
Network management and operations · 46% Software-defined and programmable networks · 46% Cellular and mobile networks · 7%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Hardware accelerators and domain-specific architectures · 67% Performance modeling and evaluation · 33%
Artificial intelligence
2 papers
Efficient and distributed learning · 84% Image recognition and object detection · 16%

Topics — the 8 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning
model compression
1.022024
MobileNetV4: Universal Models for the Mobile Ecosystem · ECCV (40) 2024
PikeLPN: Mitigating Overlooked Inefficiencies of Low-Precision Neural Networks · CVPR 2024
Hardware accelerators and domain-specific architectures › machine learning accelerator › DNN inference
low-precision DNN inference
0.812024
PikeLPN: Mitigating Overlooked Inefficiencies of Low-Precision Neural Networks · CVPR 2024
Hardware accelerators and domain-specific architectures
machine learning accelerator
0.812024
PikeLPN: Mitigating Overlooked Inefficiencies of Low-Precision Neural Networks · CVPR 2024
Network management and operations
network verification
0.712023
Hydra: Effective Runtime Network Verification · SIGCOMM 2023
Software-defined and programmable networks › programmable data plane
p4
0.712023
Hydra: Effective Runtime Network Verification · SIGCOMM 2023
Software-defined and programmable networks
programmable data plane
0.712023
Hydra: Effective Runtime Network Verification · SIGCOMM 2023
Computer vision › Image recognition and object detection › efficient visual recognition
mobile image recognition
0.212024
MobileNetV4: Universal Models for the Mobile Ecosystem · ECCV (40) 2024
Machine learning › Efficient and distributed learning › model compression
quantization
0.212024
PikeLPN: Mitigating Overlooked Inefficiencies of Low-Precision Neural Networks · CVPR 2024

Methods — techniques the papers use, named apart from their topics

quantization · 1.5double quantization · 1.5distribution-heterogeneous quantization · 1.5neural architecture search · 0.8runtime verification · 0.7domain-specific language · 0.7
YearPublicationVenuePosition
2024 PikeLPN: Mitigating Overlooked Inefficiencies of Low-Precision Neural Networks
abstract
Low-precision quantization is recognized for its efficacy in neural network optimization. Our analysis reveals that non-quantized elementwise operations which are prevalent in layers such as parameterized activation functions, batch normalization, and quantization scaling dominate the inference cost of low-precision models. These non-quantized elementwise operations are commonly overlooked in SOTA efficiency metrics such as Arithmetic Computation Effort (ACE) [46]. In this paper, we propose ACEv2- an extended version of ACE which offers a better alignment with the inference cost of quantized models and their energy consumption on ML hardware. Moreover, we introduce PikeLPN11Pike is a slim fast fish, LPN stands for Low-Precision Network., a model that addresses these efficiency issues by applying quantization to both elementwise operations and multiply-accumulate operations. In particular, we present a novel quantization technique for batch normalization layers named QuantNorm which allows for quantizing the batch normalization parameters without compromising the model performance. Additionally, we propose applying Double Quantization where the quantization scaling parameters are quantized. Furthermore, we recognize and resolve the issue of distribution mismatch in Separable Convolution layers by introducing Distribution-Heterogeneous Quantization which enables quantizing them to low-precision. PikeLPN achieves Pareto-optimality in efficiency-accuracy trade-off with up to 3× efficiency improvement compared to SOTA low-precision models.
Marina Neseem, Conor McCullough, Randy Hsin, Chas Leichner, Shan Li 0001, In Suk Chong, Andrew G. Howard, Lukasz Lew, Sherief Reda, Ville-Mikko Rautio, Daniele Moro
CVPR11
2024 MobileNetV4: Universal Models for the Mobile Ecosystem
Danfeng Qin, Chas Leichner, Manolis Delakis, Marco Fornoni, Shixin Luo, Colby R. Banbury, Chengxi Ye, Berkin Akin, Vaibhav Aggarwal, Tenghui Zhu, Daniele Moro, Andrew G. Howard
ECCV (40)13
2023 Hydra: Effective Runtime Network Verification
abstract
It is notoriously difficult to verify that a network is behaving as intended, especially at scale. This paper presents Hydra, a system that uses ideas from runtime verification to check that every packet is correctly processed with respect to a specification in real time. We propose a domain-specific language for writing properties, called Indus, and we develop a compiler that turns properties thus specified into executable P4 code that runs alongside the forwarding code at line rate. To evaluate our approach, we used Indus to model a range of properties, showing that it is expressive enough to capture examples studied in prior work. We also deployed Hydra checkers for validating paths in source routing and for enforcing slice isolation in Aether, an open-source cellular platform. We confirmed a subtle bug in Aether's 5G mobile core that would have been hard to detect using static techniques. We also evaluated the overheads of Hydra on hardware, finding that it does not significantly increase latency and often does not require additional pipeline stages.
Sundararajan Renganathan, Benny Rubin, Hyojoon Kim, Pier Luigi Ventre, Carmelo Cascone, Daniele Moro, Charles Chan, Nick McKeown, Nate Foster
SIGCOMM6
2023 Enabling Binary Neural Network Training on the Edge
abstract
The ever-growing computational demands of increasingly complex machine learning models frequently necessitate the use of powerful cloud-based infrastructure for their training. Binary neural networks are known to be promising candidates for on-device inference due to their extreme compute and memory savings over higher-precision alternatives. However, their existing training methods require the concurrent storage of high-precision activations for all layers, generally making learning on memory-constrained devices infeasible. In this article, we demonstrate that the backward propagation operations needed for binary neural network training are strongly robust to quantization, thereby making on-the-edge learning with modern models a practical proposition. We introduce a low-cost binary neural network training strategy exhibiting sizable memory footprint reductions while inducing little to no accuracy loss vs Courbariaux & Bengio’s standard approach. These decreases are primarily enabled through the retention of activations exclusively in binary format. Against the latter algorithm, our drop-in replacement sees memory requirement reductions of 3–5×, while reaching similar test accuracy (± 2 pp) in comparable time, across a range of small-scale models trained to classify popular datasets. We also demonstrate from-scratch ImageNet training of binarized ResNet-18, achieving a 3.78× memory reduction. Our work is open-source, and includes the Raspberry Pi-targeted prototype we used to verify our modeled memory decreases and capture the associated energy drops. Such savings will allow for unnecessary cloud offloading to be avoided, reducing latency, increasing energy efficiency, and safeguarding end-user privacy.
Erwei Wang, James J. Davis 0001, Daniele Moro, Jia Jie Lim, Claudionor José Nunes Coelho Jr., Satrajit Chatterjee, Peter Y. K. Cheung, George A. Constantinides
ACM Trans. Embed. Comput. Syst.3
2022 CHIMA: a Framework for Network Services Deployment and Performance Assurance
abstract
Network Function Virtualization has dramatically increased the flexibility in the deployment of network services, however the execution of virtual functions on compute nodes equipped with general purpose hardware can result in worse performance compared to the middleboxes they aim to replace. The use of programmable network hardware to perform part of the processing at line rate can drastically increase the throughput while retaining the flexibility.This work presents a new framework, called CHIMA, which extends the capabilities of other frameworks proposed in the literature for the deployment of heterogeneous Service Function Chains (SFCs). Heterogeneous SFCs comprise a combination of virtual functions meant to be executed in containers running on general purpose hardware and of functions for programmable switches written using the P4 language. CHIMA exploits programmable data planes to perform real time monitoring of the services through In-band Network Telemetry and uses the collected information to guarantee the requested levels of performance by redeploying and rerouting sections that are affected by adverse conditions, allowing applications with critical requirements to be deployed as SFCs.The solution has been tested by emulating various topologies and services on the FOP4 platform with bmv2 switches. The analysis shows that the system is capable of detecting faults in the order of hundreds of milliseconds, and the overhead it causes in the process of redeployment is negligible compared to the startup time of functions. Measurements also reveal that the current bottleneck for the runtime relocation of heterogeneous functions is the redeployment and reconfiguration of P4 programs.
Elia Battiston, Daniele Moro, Giacomo Verticale, Antonio Capone
NetSoft2
2020 rrSDS: Towards a Robot-ready Spoken Dialogue System
abstract
Spoken interaction with a physical robot requires a dialogue system that is modular, multimodal, distributive, incremental and temporally aligned.In this demo paper, we make significant contributions towards fulfilling these requirements by expanding upon the ReTiCo incremental framework.We outline the incremental and multimodal modules and how their computation can be distributed.We demonstrate the power and flexibility of our robotready spoken dialogue system to be integrated with almost any robot.
Casey Kennington, Daniele Moro, Lucas Marchand, Jake Carns, David McNeill
SIGdial2
2019 "Hello? Who Am I Talking to?" A Shallow CNN Approach for Human vs. Bot Speech Classification
abstract
Automatic speech generation algorithms, enhanced by deep learning techniques, enable an increasingly seamless and immediate machine-to-human interaction. As a result, the latest generation of phone-calling bots sounds more convincingly human than previous generations. The application of this technology has a strong social impact in terms of privacy issues (e.g., in customer-care services), fraudulent actions (e.g., social hacking) and erosion of trust (e.g., generation of fake conversation). For these reasons, it is crucial to identify the nature of a speaker, as either a human or a bot. In this paper, we propose a speech classification algorithm based on Convolutional Neural Networks (CNNs), which enables the automatic classification of human vs non-human speakers from the analysis of short audio excerpts. We evaluate the effectiveness of the proposed solution by exploiting a real human speech database populated with audio recordings from various sources, and automatically generated speeches using state-of-the-art text-to-speech generators based on deep learning (e.g., Google WaveNet).
Alessandro Lieto, Daniele Moro, Francesco Devoti, Claudia Parera, Vincenzo Lipari, Paolo Bestagini, Stefano Tubaro
ICASSP2
2017 Towards traffic classification offloading to stateful SDN data planes
abstract
Traffic classification allows network operators to gain important insights to better characterize packet flows, enabling fundamental applications such as traffic engineering, network analytics and Quality of Service (QoS) enforcing. A common approach adopted for flow classification is based on Deep Packet Inspection (DPI): all the traffic is processed by a middlebox whose task is the association of a network flow to the application-level information by inspecting the entire content of the packets. The increased volume of encrypted traffic limits the type of analysis performed by network middleboxes. However, an important amount of information can still be extracted from packets belonging to the very initial phase of a connection which are transmitted in clear (e.g. DNS and TLS handshake). Furthermore, recent research work has shown that it is possible to reduce the burden on the DPI without a significant loss in classification accuracy, by limiting the amount of data processed per flow. In this paper, we propose to exploit the programmability of new stateful SDN data planes to offload down to the network the process of filtering traffic to the DPI. We show that it is jointly possible to reduce the required computing power of the DPI, as well as the network bandwidth between the switches and the DPI. By taking advantage of the flexibility of stateful data planes we also manage to delegate to switches the computation of useful network analytics metrics (such as number of packets, number of bytes and duration) which would otherwise require the DPI to inspect the entire traffic flow.
Davide Sanvito, Daniele Moro, Antonio Capone
NetSoft2