EDBT 2026 Demo / reviewers in the wild / expert
Eran Raichstein
dblp:32/8210
· DBLP profile ↗
15ranked-venue papers
1as first author
7since 2021 · last 2024
0000-0001-7962-4876ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 12 · 1 first-author · 6 since 2021Computer networks · 1Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Designing a Lightweight Network Observability Agent for Cloud Applications
Pravein G. Kannan, Shachee Mishra Gupta, Dushyant Behl, Eran Raichstein, Joel Takvorian |
PAM (1) | 4 |
| 2024 | Observability Volume ManagementabstractObservability Volume Management (OVM) presents a lightweight, automated processing system to help manage the large amounts of Observability data. The focus is on automating, analyzing, and making recommendations for volume management in multi-cloud, edge, and distributed systems. Eran Raichstein, Kalman Meth, Seep Goel, Priyanka Naik, Kavya Govindarajan |
SYSTOR | 1 |
| 2024 | ARISE: AI Right Sizing Engine for AI workload configurationsabstractData scientists and platform engineers who maintain AI stacks are required to continuously run AI workloads. When executing any part of the AI pipeline, whether data preprocessing, training, fine-tuning or inference, a frequent question is how to optimally configure the environment to meet Service Level Objectives (SLOs), such as desired throughput, runtime deadlines, and avoid memory and CPU exhaustion. We present ARISE, a tool that enables making data-driven decisions about AI workload configuration questions. ARISE trains performance prediction machine-learning regression models on historical workloads and performance benchmark metadata, and then predicts the performance of future workloads based on their input metadata, using the best performing regression models. Initial evaluation of ARISE on real-world workloads shows high prediction accuracy. Rachel Tzoref, Bruno Wassermann, Eran Raichstein, Dean H. Lorenz |
SYSTOR | 3 |
| 2023 | Smart Network Observability - Connection TrackingabstractFlow Logs Pipeline (a.k.a. FLP) is an observability tool that consumes flow logs from various inputs, transforms them and exports logs to Loki and / or time series metrics to Prometheus. While flow logs encompass a lot of valuable data, observing the network from the level of flow logs is often too low. In many cases, we are interested in observing it from a higher level, the level of connections. In this work, we introduce a new processing stage in FLP that allows aggregating flow logs from the same connection - connection tracking. Ronen Schaffer, Eran Raichstein, Kalman Meth, Joel Takvorian, Julien Pinsonneau |
SYSTOR | 2 |
| 2022 | Hybrid anomaly detection and prioritization for network logs at cloud scaleabstractMonitoring the health of large-scale systems requires significant manual effort, usually through the continuous curation of alerting rules based on keywords, thresholds and regular expressions, which might generate a flood of mostly irrelevant alerts and obscure the actual information operators would like to see. Existing approaches try to improve the observability of systems by intelligently detecting anomalous situations. Such solutions surface anomalies that are statistically significant, but may not represent events that reliability engineers consider relevant. We propose ADEPTUS, a practical approach for detection of relevant health issues in an established system. ADEPTUS combines statistics and unsupervised learning to detect anomalies with supervised learning and heuristics to determine which of the detected anomalies are likely to be relevant to the Site Reliability Engineers (SREs). ADEPTUS overcomes the labor-intensive prerequisite of obtaining anomaly labels for supervised learning by automatically extracting information from historic alerts and incident tickets. We leverage ADEPTUS for observability in the network infrastructure of IBM Cloud. We perform an extensive real-world evaluation on 10 months of logs generated by tens of thousands of network devices across 11 data centers and demonstrate that ADEPTUS achieves higher alerting accuracy than the rule-based log alerting solution, curated by domain experts, used by SREs daily. David Ohana, Bruno Wassermann, Nicolas Dupuis, Elliot K. Kolodner, Eran Raichstein, Michal Malka |
EuroSys | 5 |
| 2022 | Smart network metrics derivation from flow logsabstractFlow-Logs Pipeline is an observability tool that consumes raw network flow-logs and transforms them from their original format (e.g. NetFlow or IPFIX) into numeric metrics format. FLP allows to define transformations of data to generate condensed metrics that encapsulate network domain knowledge. Kalman Meth, Eran Raichstein, Katherine Barabash, Ronen Schaffer, Joel Takvorian, Mario Macías |
SYSTOR | 2 |
| 2021 | DeCorus-NSA: detection and correlation of unusual signals for network syslog analyticsabstractThe management of large data centre (DC) network infrastructure confronts Network Reliability Engineers (NRE) with challenges. A single DC at a modern cloud services provider can host thousands of network devices. The syslog messages generated by these devices are an important type of monitoring data to detect and diagnose failures. Devices in a single DC produce millions of syslog messages per day in a variety of formats. David Ohana, Bruno Wassermann, Moshe Hershcovitch, Elliot K. Kolodner, Michal Malka, Eran Raichstein, Ronen Schaffer, Robert Shahla |
SYSTOR | 6 |
| 2019 | Exploring and troubleshooting istio issuesabstractCloud computing gave rise to a Cloud-native[1] approach for operating application software in the cloud, whereby applications are segmented into micro-services that can be designed and deployed independently of each other. This significantly increases application maintainability, reduces time to market, and helps leveraging cloud computing model. On the other hand, this approach increases the system level complexity of the application and poses new challenges, such as how services discover each other, and how application handles individual service upgrades. To support cloud-native paradigm, new development, deployment, and orchestration tools are created. One of such tools is Istio [2] service mesh, built to connect, secure, control, and observe services. While immensely useful to application developers, Istio is an additional layer in cloud compute platform software stack and is thus prone to failure or misuse. Tomer Lange, Aidan Shribman, Eran Raichstein, Katherine Barabash |
SYSTOR | 3 |
| 2019 | Estimating client QoE from measured network QoSabstractThis research is done in the context of the SliceNet project [4] that aims to extend 5G infrastructure with cognitive management of cross-domain, cross-layer network slices [1], with emphasis on Quality of Experience (QoE) for vertical industries. The provisioning of network slices with proper QoE guarantees is seen as one of the key enablers of future 5G-enabled networks. The challenge is to assess the QoE experienced by the vertical application and its users without requiring the applications or the users to measure and report QoE related metrics back to the provider. To address this challenge, we propose a method for deriving application-level QoE from network-level Quality of Service (QoS) measurements, easily accessible by the provider. In particular, we describe a PoC where QoE, perceived by application users, is estimated from low level network monitoring data, by applying cognitive methods. Our main goal is enabling the cloud provider to support the desired E2E QoE-based Service Level Agreements (SLAs), e.g. by monitoring QoS metrics within the provider's domain to optimize resource allocation through provider's actuators. Additional benefit can be achieved by applying the same technique to troubleshoot issues in the provider's infrastructure. In this work, we employed classical statistical methods to assess the relationship between the application-level QoE and the network-level QoS. Kenneth Nagin, Andre Kassis, Dean H. Lorenz, Katherine Barabash, Eran Raichstein |
SYSTOR | 5 |
| 2019 | Noisy neighbor detection using skydiveabstractCloud computing technology enables uniform access to shared pools of configurable system resources and higher-level services, rapidly provisioned with minimal management effort. Cloud computing relies on sharing the resources to achieve coherence and economies of scale, through virtualizion. Cloud network, in particular, is virtualized through multiple logical constructs and SW layers, making cloud connectivity complex to configure, debug, and visualize. In this work, we show how to detect cloud network operational issues through monitoring and analytics, using and enhancing open source network analyzer, Skydive [2]. In particular, we focus on Noisy Neighbor Effect, a situation in which a common resource is monopolized by a noisy tenant, resulting in performance degradation experienced by other tenants. Ofir Zeilig, Noa Bratman, Itzik Ashkenazi, Eran Raichstein, Anna Levin, Katherine Barabash |
SYSTOR | 4 |
| 2018 | Heterogeneous Resource ReservationabstractGiven a large variety of resources and billing contracts offered by today’s cloud providers, customers face a nontrivial optimization challenge for their application workloads. A number of works are dealing with either billing contracts selection optimization or resource types selection. We argue that the largest cost savings to elastic workloads result from jointly optimizing heterogeneous resources and billing contracts selection. To this end, we introduce a novel cloud control and management framework and formulate a novel optimization problem called Heterogeneous Resource Reservation (HRR). We evaluate our solution through a thorough simulation study using publicly available cloud workload data as well as internal anonymous customer data. For these data our approach attain dramatic cost savings compared to the current state of the art. Ofer Biran, David Breitgand, Dean H. Lorenz, Michael Masin, Eran Raichstein, Avi Weit, Ilyas Iyoob |
IC2E | 5 |
| 2018 | SliceNet: Cognitive Slice Management Framework for Virtual Multi-Domain 5G NetworksabstractNo abstract available. Dean H. Lorenz, V. Perelman, Eran Raichstein, Katherine Barabash, Aidan Shribman |
SYSTOR | 3 |
| 2017 | CogNETive: insights and visualization for operations@scaleabstractOperating a cloud-scale service is a huge challenge. There are millions of users worldwide and millions of requests per seconds. For example, Amazon's Simple Storage Service (S3) in 2013 contained two trillion objects and its logs contained 1.1 million log lines per second, which are approximately 10 PB of log records per year (see [1]). Cloud scale implies thousands of servers and network elements, and hundreds of services from multiple cross-regional data centers. Cloud service operation data is scattered over various types of semi-structured and unstructured logs (e.g., application, error, debug), telemetry and network data, as well as customer service records. It is therefore extremely difficult for the multiple owners and administrators in such systems, coming from different units of the organization, to follow the possible paths and system alternatives in order to detect problems, solve issues and understand the service operation. Dean H. Lorenz, Eran Raichstein, Katherine Barabash, Hillel Kolodner, Liran Schour, Shelly Garion |
SYSTOR | 2 |
| 2015 | EnforSDN: Network policies enforcement with SDNabstractNetwork services, such as security, load-balancing, and monitoring, are an indisputable part of modern networking infrastructure and are traditionally realized as specialized appliances or middleboxes. Middleboxes complicate the management, the deployment, and the operations of the entire network. Moreover, they induce network performance issues and scalability limitations by requiring huge amounts of traffic to be, often sub-optimally redirected, and sometimes redundantly processed. Recent trends of server virtualization and Network Function Vir-tualization (NFV) exacerbate these scalability and performance issues. In this paper, we present EnforSDN — a new management approach that exploits SDN principles to decouple the policy resolution layer from the policy enforcement layer in network service appliances. Our approach improves the enforcement management, network utilization and communication latency, without compromising the policy and the functionality of the network. Using emulated SDN-based data center environment, we demonstrate higher throughput and lower latency achieved with EnforSDN, as compared to a baseline SDN network. In addition, we show that EnforSDN reduces the overall network appliances load, as well as the forwarding tables size. Yaniv Ben-Itzhak, Katherine Barabash, Rami Cohen, Anna Levin, Eran Raichstein |
IM | 5 |
| 2010 | Using machine learning techniques to enhance the performance of an automatic backup and recovery systemabstractA typical disaster recovery system will have mirrored storage at a site that is geographically separate from the main operational site. In many cases, communication between the local site and the backup repository site is performed over a network which is inherently slow, such as a WAN, or is highly strained, for example due to a whole-site disaster recovery operation.The goal of this work is to alleviate the performance impact of the network in such a scenario, and to do so using machine learning techniques. We focus on two main areas, prefetching and read-ahead size determination. In both cases we significantly improve the performance of the system.Our main contributions are as follows: We introduce a theoretical model of the system and the problem we are trying to solve and bound the gain from prefetching techniques. We construct two frequent pattern mining algorithms and use them for prefetching. A framework for controlling and combining multiple prefetch algorithms is presented as well. These algorithms, as well as various simple prefetch algorithms, are compared on a simulation environment. We introduce a novel algorithm for determining the amount of read ahead on such a system that is based on intuition from online competitive analysis and on regression techniques. The significant positive impact of this algorithm is demonstrated on IBM's FastBack system.Much of our improvements have been applied with little or no modification of the current implementation's internals. We therefore feel confident in stating that the techniques are general and are likely to have applications elsewhere. Dan Pelleg, Eran Raichstein, Amir Ronen |
SYSTOR | 2 |