EDBT 2026 Demo / reviewers in the wild / expert
Michal Malka
dblp:221/1620
· DBLP profile ↗
6ranked-venue papers
0as first author
5since 2021 · last 2025
0000-0001-8215-7682ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | ZipNN: Lossless Compression for AI ModelsabstractWith the growth of model sizes and the scale of their deployment, their sheer size burdens the infrastructure requiring more network and more storage to accommodate these. While there is a vast model compression literature deleting parts of the model weights for faster inference, we investigate a more traditional type of compression - one that represents the model in a compact form and is coupled with a decompression algorithm that returns it to its original form and size - namely lossless compression. We present ZipNN, a lossless compression tailored to neural networks. Somewhat surprisingly, we show that specific lossless compression can gain significant network and storage reduction on popular models, often saving 33% and at times reducing over 50% of the model size. We investigate the source of model compressibility and introduce specialized compression variants tailored for models that further increase the effectiveness of compression. On popular models (e.g. Llama 3) ZipNN shows space savings that are over 17% better than vanilla compression while also improving compression and decompression speeds by 62%. Using multiple workers and threads, ZipNN can achieve decompression speeds of up to 80GB/s and compression speed of up to 13GB/s. We estimate that these methods could save over an ExaByte per year of network traffic downloaded from a large model hub like Hugging Face. Moshe Hershcovitch, Andrew Wood, Leshem Choshen, Guy Girmonsky, Roy Leibovitz, Or Ozeri, Ilias Ennmouri, Michal Malka, Sang (Peter) Chin, Swaminathan Sundararaman, Danny Harnik |
CLOUD | 8 |
| 2025 | ClusterLink: Redefining Application Connectivity for the Multi-cloud EraabstractModern software development abstracts applications from the underlying infrastructure, enabling global-scale deployment with minimal concern about low-level networking details. However, when these infrastructure-agnostic software components need to communicate, they encounter significant networking limitations. This forces developers to either navigate complex, low-level networking constructs to achieve the desired connectivity or give up on truly flexible connectivity and limit their software to static connectivity patterns. In this paper, we focus on the evolving challenges of application connectivity in today's hyper-distributed reality. We propose to model connectivity around the notion of application services and have realized this proposal as ClusterLink, which exposes the app-level APIs for specifying communication policies at a very granular level and implements them efficiently. This paper shares details on ClusterLink design principles, APIs, architecture, and implementation, and shows that ClusterLink outperforms its closest competitor by 2.5x in throughput in a cloud-based experimental setting. Kfir Toledo, Pravein G. Kannan, Michal Malka, Etai Lev-Ran, Or Ozeri, Vita Bortnikov, Ziv Nevo, Katherine Barabash |
CLOUD | 3 |
| 2023 | ClusterLink: A Multi-Cluster Application InterconnectabstractEnterprises often deploy their business applications in multiple clouds as well as in multiple traditional environments. This work focuses on the connectivity aspects of this new way of operating and consuming digital services. We define the related requirements, analyze the challenges, and present ClusterLink, our solution for interconnecting today's and future multi-cloud applications. Kfir Toledo, Pravein G. Kannan, Michal Malka, Etai Lev-Ran, Katherine Barabash, Vita Bortnikov |
SYSTOR | 3 |
| 2022 | Hybrid anomaly detection and prioritization for network logs at cloud scaleabstractMonitoring the health of large-scale systems requires significant manual effort, usually through the continuous curation of alerting rules based on keywords, thresholds and regular expressions, which might generate a flood of mostly irrelevant alerts and obscure the actual information operators would like to see. Existing approaches try to improve the observability of systems by intelligently detecting anomalous situations. Such solutions surface anomalies that are statistically significant, but may not represent events that reliability engineers consider relevant. We propose ADEPTUS, a practical approach for detection of relevant health issues in an established system. ADEPTUS combines statistics and unsupervised learning to detect anomalies with supervised learning and heuristics to determine which of the detected anomalies are likely to be relevant to the Site Reliability Engineers (SREs). ADEPTUS overcomes the labor-intensive prerequisite of obtaining anomaly labels for supervised learning by automatically extracting information from historic alerts and incident tickets. We leverage ADEPTUS for observability in the network infrastructure of IBM Cloud. We perform an extensive real-world evaluation on 10 months of logs generated by tens of thousands of network devices across 11 data centers and demonstrate that ADEPTUS achieves higher alerting accuracy than the rule-based log alerting solution, curated by domain experts, used by SREs daily. David Ohana, Bruno Wassermann, Nicolas Dupuis, Elliot K. Kolodner, Eran Raichstein, Michal Malka |
EuroSys | 6 |
| 2021 | DeCorus-NSA: detection and correlation of unusual signals for network syslog analyticsabstractThe management of large data centre (DC) network infrastructure confronts Network Reliability Engineers (NRE) with challenges. A single DC at a modern cloud services provider can host thousands of network devices. The syslog messages generated by these devices are an important type of monitoring data to detect and diagnose failures. Devices in a single DC produce millions of syslog messages per day in a variety of formats. David Ohana, Bruno Wassermann, Moshe Hershcovitch, Elliot K. Kolodner, Michal Malka, Eran Raichstein, Ronen Schaffer, Robert Shahla |
SYSTOR | 5 |
| 2018 | Shared Cloud Object Store, governed by permissioned blockchainabstractNo abstract available. Artem Barger, Yacov Manevich, Vita Bortnikov, Yoav Tock, Michael Factor, Michal Malka |
SYSTOR | 6 |