EDBT 2026 Demo / reviewers in the wild / expert
Mário Almeida
dblp:122/3298
· DBLP profile ↗
13ranked-venue papers
5as first author
5since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 6 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-authorSecurity and privacy · 2 · 1 first-authorArtificial intelligence and machine learning · 1 · 1 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Efficient and distributed learning · 96% Deep learning architectures and training · 4% | |
| Computer architecture, parallel and distributed computing, and storage systems
4 papers |
Hardware accelerators and domain-specific architectures · 84% Performance modeling and evaluation · 8% Cloud and datacenter computing · 6% | |
| Computer networks
3 papers |
Network measurement and analytics · 54% Internet architecture and protocols · 13% Cellular and mobile networks · 12% | |
| Human-computer interaction and pervasive computing
1 paper |
Ubiquitous computing and smart environments · 100% | |
| Software engineering, system software, and programming languages
1 paper |
Software testing · 100% | |
| Computer graphics and multimedia
1 paper |
Image and video processing · 100% |
Topics — the 26 heaviest of 28, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Efficient and distributed learning
federated learning |
1.0 | 2 | 2021 | FjORD: Fair and Accurate Federated Learning under heterogeneous targets with Ordered Dropout · NeurIPS 2021 Federated mobile sensing for activity recognition · MobiCom 2021 |
Hardware accelerators and domain-specific architectures › quantization
mixed-precision quantization |
0.8 | 1 | 2024 | NAWQ-SR: A Hybrid-Precision NPU Engine for Efficient On-Device Super-Resolution · IEEE Trans. Mob. Comput. 2024 |
Hardware accelerators and domain-specific architectures › machine learning accelerator › neural network accelerator
neural processing unit |
0.8 | 1 | 2024 | NAWQ-SR: A Hybrid-Precision NPU Engine for Efficient On-Device Super-Resolution · IEEE Trans. Mob. Comput. 2024 |
Machine learning › Efficient and distributed learning › federated learning › heterogeneous federated learning
device heterogeneity |
0.5 | 1 | 2021 | FjORD: Fair and Accurate Federated Learning under heterogeneous targets with Ordered Dropout · NeurIPS 2021 |
Machine learning › Efficient and distributed learning
on-device inference |
0.5 | 1 | 2021 | Smart at what cost?: characterising mobile deep neural networks in the wild · Internet Measurement Conference 2021 |
Ubiquitous computing and smart environments › context recognition › activity recognition
accelerometer-based activity recognition |
0.5 | 1 | 2021 | Federated mobile sensing for activity recognition · MobiCom 2021 |
Ubiquitous computing and smart environments › context recognition
activity recognition |
0.5 | 1 | 2021 | Federated mobile sensing for activity recognition · MobiCom 2021 |
Machine learning › Efficient and distributed learning › distributed inference › collaborative inference
cloud-edge collaborative inference |
0.4 | 1 | 2020 | SPINN: synergistic progressive inference of neural networks over device and cloud · MobiCom 2020 |
Machine learning › Efficient and distributed learning
distributed inference |
0.4 | 1 | 2020 | SPINN: synergistic progressive inference of neural networks over device and cloud · MobiCom 2020 |
Software testing › dynamic testing › manual testing
crowdsourced testing |
0.3 | 1 | 2018 | CHIMP: Crowdsourcing Human Inputs for Mobile Phones · WWW 2018 |
Software testing
mobile application testing |
0.3 | 1 | 2018 | CHIMP: Crowdsourcing Human Inputs for Mobile Phones · WWW 2018 |
Software testing
test input generation |
0.3 | 1 | 2018 | CHIMP: Crowdsourcing Human Inputs for Mobile Phones · WWW 2018 |
Network measurement and analytics › traffic analysis
DNS traffic analysis |
0.3 | 1 | 2017 | Dissecting DNS Stakeholders in Mobile Networks · CoNEXT 2017 |
Network measurement and analytics › mobile network measurement
mobile traffic analysis |
0.3 | 1 | 2017 | Dissecting DNS Stakeholders in Mobile Networks · CoNEXT 2017 |
Image and video processing › super-resolution
learning-based super-resolution |
0.2 | 1 | 2024 | NAWQ-SR: A Hybrid-Precision NPU Engine for Efficient On-Device Super-Resolution · IEEE Trans. Mob. Comput. 2024 |
Image and video processing
super-resolution |
0.2 | 1 | 2024 | NAWQ-SR: A Hybrid-Precision NPU Engine for Efficient On-Device Super-Resolution · IEEE Trans. Mob. Comput. 2024 |
Wireless networking
cross-layer interaction |
0.2 | 1 | 2013 | RILAnalyzer: a comprehensive 3G monitor on your phone · Internet Measurement Conference 2013 |
Network measurement and analytics › mobile network measurement
mobile network monitoring |
0.2 | 1 | 2013 | RILAnalyzer: a comprehensive 3G monitor on your phone · Internet Measurement Conference 2013 |
Machine learning › Efficient and distributed learning
model compression |
0.1 | 1 | 2021 | FjORD: Fair and Accurate Federated Learning under heterogeneous targets with Ordered Dropout · NeurIPS 2021 |
Privacy and data protection
privacy-preserving machine learning |
0.1 | 1 | 2021 | Federated mobile sensing for activity recognition · MobiCom 2021 |
Performance modeling and evaluation
workload characterization |
0.1 | 1 | 2021 | Smart at what cost?: characterising mobile deep neural networks in the wild · Internet Measurement Conference 2021 |
Machine learning › Deep learning architectures and training
convolutional neural network |
0.1 | 1 | 2020 | SPINN: synergistic progressive inference of neural networks over device and cloud · MobiCom 2020 |
Edge and fog computing › mobile edge computing
computation offloading |
0.1 | 1 | 2020 | SPINN: synergistic progressive inference of neural networks over device and cloud · MobiCom 2020 |
Cloud and datacenter computing
virtualization |
0.1 | 1 | 2018 | CHIMP: Crowdsourcing Human Inputs for Mobile Phones · WWW 2018 |
Internet architecture and protocols › domain name system
DNS caching |
0.1 | 1 | 2017 | Dissecting DNS Stakeholders in Mobile Networks · CoNEXT 2017 |
Internet architecture and protocols
domain name system |
0.1 | 1 | 2017 | Dissecting DNS Stakeholders in Mobile Networks · CoNEXT 2017 |
Methods — techniques the papers use, named apart from their topics
runtime neural image codec · 1.5hybrid-precision quantization · 1.5federated learning · 1.5deep neural network · 1.5measurement · 1.0scheduling · 0.9early exit · 0.9CNN splitting · 0.9network traffic classification · 0.7crowdsourcing · 0.7singular value decomposition · 0.5self-distillation · 0.5user study · 0.3on-device radio logging · 0.3passive traffic analysis · 0.3active measurement · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | NAWQ-SR: A Hybrid-Precision NPU Engine for Efficient On-Device Super-ResolutionabstractIn recent years, image and video delivery systems have begun integrating deep learning super-resolution (SR) approaches, leveraging their unprecedented visual enhancement capabilities while reducing reliance on networking conditions. Nevertheless, deploying these solutions on mobile devices still remains an active challenge as SR models are excessively demanding with respect to workload and memory footprint. Despite recent progress on on-device SR frameworks, existing systems either penalize visual quality, lead to excessive energy consumption or make inefficient use of the available resources. This work presents NAWQ-SR, a novel framework for the efficient on-device execution of SR models. Through a novel hybrid-precision quantization technique and a runtime neural image codec, NAWQ-SR exploits the multi-precision capabilities of modern mobile NPUs in order to minimize latency, while meeting user-specified quality constraints. Moreover, NAWQ-SR selectively adapts the arithmetic precision at run time to equip the SR DNN's layers with wider representational power, improving visual quality beyond what was previously possible on NPUs.Altogether, NAWQ-SR achieves an average speedup of 7.9×, 3× and 1.91× over the state-of-the-art on-device SR systems that use heterogeneous processors (MobiSR), CPU (SplitSR) and NPU (XLSR), respectively.Furthermore, NAWQ-SR delivers an average of 3.2× speedup and 0.39 dB higher PSNR over status-quo INT8 NPU designs, but most importantly mitigates the negative effects of quantization on visual quality, setting a new state-of-the-art in the attainable quality of NPU-based SR. Stylianos I. Venieris, Mário Almeida, Royson Lee, Nicholas D. Lane |
IEEE Trans. Mob. Comput. | 2 |
| 2022 | DynO: Dynamic Onloading of Deep Neural Networks from Cloud to DeviceabstractRecently, there has been an explosive growth of mobile and embedded applications using convolutional neural networks (CNNs). To alleviate their excessive computational demands, developers have traditionally resorted to cloud offloading, inducing high infrastructure costs and a strong dependence on networking conditions. On the other end, the emergence of powerful SoCs is gradually enabling on-device execution. Nonetheless, low- and mid-tier platforms still struggle to run state-of-the-art CNNs sufficiently. In this article, we present DynO, a distributed inference framework that combines the best of both worlds to address several challenges, such as device heterogeneity, varying bandwidth, and multi-objective requirements. Key components that enable this are its novel CNN-specific data packing method, which exploits the variability of precision needs in different parts of the CNN when onloading computation, and its novel scheduler, which jointly tunes the partition point and transferred data precision at runtime to adapt inference to its execution environment. Quantitative evaluation shows that DynO outperforms the current state of the art, improving throughput by over an order of magnitude over device-only execution and up to 7.9× over competing CNN offloading systems, with up to 60× less data transferred. Mário Almeida, Stefanos Laskaridis, Stylianos I. Venieris, Ilias Leontiadis, Nicholas D. Lane |
ACM Trans. Embed. Comput. Syst. | 1 |
| 2021 | Smart at what cost?: characterising mobile deep neural networks in the wildabstractWith smartphones' omnipresence in people's pockets, Machine Learning (ML) on mobile is gaining traction as devices become more powerful. With applications ranging from visual filters to voice assistants, intelligence on mobile comes in many forms and facets. However, Deep Neural Network (DNN) inference remains a compute intensive workload, with devices struggling to support intelligence at the cost of responsiveness. On the one hand, there is significant research on reducing model runtime requirements and supporting deployment on embedded devices. On the other hand, the strive to maximise the accuracy of a task is supported by deeper and wider neural networks, making mobile deployment of state-of-the-art DNNs a moving target. Mário Almeida, Stefanos Laskaridis, Abhinav Mehrotra, Lukasz Dudziak, Ilias Leontiadis, Nicholas D. Lane |
Internet Measurement Conference | 1 |
| 2021 | Federated mobile sensing for activity recognitionabstractDespite advances in hardware and software enabling faster on-device inference, training Deep Neural Networks (DNN) models has largely been a long-running task over TBs of collected user data in centralised repositories. Federated Learning has emerged as an alternative, privacy-preserving paradigm to train models without accessing directly on-device data, by leveraging device resources to create per client updates and aggregate centrally. This has been applied to various tasks, ranging from next-word prediction to automatic speech recognition (ASR). In this tutorial, we recognise on-device sensing as a privacy-sensitive task and build a federated learning system from scratch to showcase how to train a model for accelerometer-based activity recognition in a federated manner. In addition, we present the current landscape and challenges in the realm of federated learning and mobile sensing and provide guidelines on how to build such systems in a privacy-preserving and scalable manner. Stefanos Laskaridis, Dimitris Spathis, Mário Almeida |
MobiCom | 3 |
| 2021 | FjORD: Fair and Accurate Federated Learning under heterogeneous targets with Ordered DropoutabstractFederated Learning (FL) has been gaining significant traction across different ML tasks, ranging from vision to keyboard predictions. In large-scale deployments, client heterogeneity is a fact and constitutes a primary problem for fairness, training performance and accuracy. Although significant efforts have been made into tackling statistical data heterogeneity, the diversity in the processing capabilities and network bandwidth of clients, termed system heterogeneity, has remained largely unexplored. Current solutions either disregard a large portion of available devices or set a uniform limit on the model's capacity, restricted by the least capable participants.In this work, we introduce Ordered Dropout, a mechanism that achieves an ordered, nested representation of knowledge in Neural Networks and enables the extraction of lower footprint submodels without the need for retraining. We further show that for linear maps our Ordered Dropout is equivalent to SVD. We employ this technique, along with a self-distillation methodology, in the realm of FL in a framework called FjORD. FjORD alleviates the problem of client system heterogeneity by tailoring the model width to the client's capabilities. Extensive evaluation on both CNNs and RNNs across diverse modalities shows that FjORD consistently leads to significant performance gains over state-of-the-art baselines while maintaining its nested structure. Samuel Horváth, Stefanos Laskaridis, Mário Almeida, Ilias Leontiadis, Stylianos I. Venieris, Nicholas D. Lane |
NeurIPS | 3 |
| 2020 | SPINN: synergistic progressive inference of neural networks over device and cloudabstractDespite the soaring use of convolutional neural networks (CNNs) in mobile applications, uniformly sustaining high-performance inference on mobile has been elusive due to the excessive computational demands of modern CNNs and the increasing diversity of deployed devices. A popular alternative comprises offloading CNN processing to powerful cloud-based servers. Nevertheless, by relying on the cloud to produce outputs, emerging mission-critical and high-mobility applications, such as drone obstacle avoidance or interactive applications, can suffer from the dynamic connectivity conditions and the uncertain availability of the cloud. In this paper, we propose SPINN, a distributed inference system that employs synergistic device-cloud computation together with a progressive inference method to deliver fast and robust CNN inference across diverse settings. The proposed system introduces a novel scheduler that co-optimises the early-exit policy and the CNN splitting at run time, in order to adapt to dynamic conditions and meet user-defined service-level requirements. Quantitative evaluation illustrates that SPINN outperforms its state-of-the-art collaborative inference counterparts by up to 2× in achieved throughput under varying network conditions, reduces the server cost by up to 6.8× and improves accuracy by 20.7% under latency constraints, while providing robust operation under uncertain connectivity conditions and significant energy savings compared to cloud-centric execution. Stefanos Laskaridis, Stylianos I. Venieris, Mário Almeida, Ilias Leontiadis, Nicholas D. Lane |
MobiCom | 3 |
| 2018 | A Family of Droids-Android Malware Detection via Behavioral Modeling: Static vs Dynamic AnalysisabstractFollowing the increasing popularity of the mobile ecosystem, cybercriminals have increasingly targeted mobile ecosystems, designing and distributing malicious apps that steal information or cause harm to the device's owner. Aiming to counter them, detection techniques based on either static or dynamic analysis that model Android malware, have been proposed. While the pros and cons of these analysis techniques are known, they are usually compared in the context of their limitations e.g., static analysis is not able to capture runtime behaviors, full code coverage is usually not achieved during dynamic analysis, etc. Whereas, in this paper, we analyze the performance of static and dynamic analysis methods in the detection of Android malware and attempt to compare them in terms of their detection performance, using the same modeling approach.To this end, we build on MAMADROID, a state-of-the-art detection system that relies on static analysis to create a behavioral model from the sequences of abstracted API calls. Then, aiming to apply the same technique in a dynamic analysis setting, we modify CHIMP, a platform recently proposed to crowdsource human inputs for app testing, in order to extract API calls' sequences from the traces produced while executing the app on a CHIMP virtual device. We call this system AUNTIEDROID and instantiate it by using both automated (Monkey) and usergenerated inputs. We find that combining both static and dynamic analysis yields the best performance, with $F -$measure reaching 0.92. We also show that static analysis is at least as effective as dynamic analysis, depending on how apps are stimulated during execution, and investigate the reasons for inconsistent misclassifications across methods. Lucky Onwuzurike, Mário Almeida, Enrico Mariconti, Jeremy Blackburn, Gianluca Stringhini, Emiliano De Cristofaro |
PST | 2 |
| 2018 | CHIMP: Crowdsourcing Human Inputs for Mobile PhonesabstractWhile developing mobile apps is becoming easier, testing and characterizing their behavior is still hard. On the one hand, the de facto testing tool, called "Monkey," scales well due to being based on random inputs, but fails to gather inputs useful in understanding things like user engagement and attention. On the other hand, gathering inputs and data from real users requires distributing instrumented apps, or even phones with pre-installed apps, an expensive and inherently unscaleable task. To address these limitations we present CHIMP, a system that integrates automated tools and large-scale crowdsourced inputs. CHIMP is different from previous approaches in that it runs apps in a virtualized mobile environment that thousands of users all over the world can access via a standard Web browser. CHIMP is thus able to gather the full range of real-user inputs, detailed run-time traces of apps, and network traffic. We thus describe CHIMP»s design and demonstrate the efficiency of our approach by testing thousands of apps via thousands of crowdsourced users. We calibrate CHIMP with a large-scale campaign to understand how users approach app testing tasks. Finally, we show how CHIMP can be used to improve both traditional app testing tasks, as well as more novel tasks such as building a traffic classifier on encrypted network flows. Mário Almeida, Muhammad Bilal 0007, Alessandro Finamore, Ilias Leontiadis, Yan Grunenberger, Matteo Varvello, Jeremy Blackburn |
WWW | 1 |
| 2017 | Dissecting DNS Stakeholders in Mobile NetworksabstractThe functioning of mobile apps involves a large number of protocols and entities, with the Domain Name System (DNS) acting as a predominant one. Despite being one of the oldest Internet systems, DNS still operates with semi-obscure interactions among its stakeholders: domain owners, network operators, operating systems, and app developers. The goal of this work is to holistically understand the dynamics of DNS in mobile traffic along with the role of each of its stakeholders. We use two complementary (anonymized) datasets: traffic logs provided by a European mobile network operator (MNO) with 19M customers, and traffic logs from 5,000 users of Lumen, a traffic monitoring app for Android. We complement such passive traffic analysis with active measurements at four European MNOs. Our study reveals that 10k domains (out of 198M) account for 87% of total network flows. The time to live (TTL) values for such domains are mostly short (< 1min), despite domain-to-IPs mapping tends to change on a longer time-scale. Further, depending on the operators recursive resolver architecture, end-user devices receive even smaller TTL values leading to suboptimal effectiveness of the on-device DNS cache. Despite a number of on-device and in-network optimizations available to minimize DNS overhead, which we find corresponding to 10% of page load time (PLT) on average, we have not found wide evidence of their adoption in the wild. Mário Almeida, Alessandro Finamore, Diego Perino, Narseo Vallina-Rodriguez, Matteo Varvello |
CoNEXT | 1 |
| 2016 | An Empirical Study of Android Alarm Usage for Application Scheduling
Mário Almeida, Muhammad Bilal 0007, Jeremy Blackburn, Konstantina Papagiannaki |
PAM | 1 |
| 2013 | RILAnalyzer: a comprehensive 3G monitor on your phoneabstractThe popularity of smartphones, cloud computing, and the app store model have led to cellular networks being used in a completely different way than what they were designed for. As a consequence, mobile applications impose new challenges in the design and efficient configuration of constrained networks to maximize application's performance. Such difficulties are largely caused by the lack of cross-layer under- standing of interactions between different entities -applications, devices, the network and its management plane. In this paper, we describe RILAnalyzer, an open-source tool that provides mechanisms to perform network analysis from within a mobile device. RILAnalyzer is capable of recording low-level radio information and accurate cellular net- work control-plane data, as well as user-plane data. We demonstrate how such data can be used to identify previously overlooked issues. Through a small user study across four cellular network providers in two European countries we infer how different network configurations are in reality and explore how such configurations interact with application logic, causing network and energy overheads. Narseo Vallina-Rodriguez, Andrius Aucinas, Mário Almeida, Yan Grunenberger, Konstantina Papagiannaki, Jon Crowcroft |
Internet Measurement Conference | 3 |
| 2012 | Using DEMO-based SLAs for Improving City Council Services
Carlos Mendes, Mário Almeida, Nuno Salvador, Miguel Mira da Silva |
KEOD | 2 |
| 2012 | Modelling Services with DEMO
Carlos Mendes, Mário Almeida, Nuno Salvador, Miguel Mira da Silva |
IC3K | 2 |