Bruno Wassermann

dblp:87/1429 · DBLP profile ↗
← Back
9ranked-venue papers
3as first author
3since 2021 · last 2024
0000-0003-2584-2629ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 3 since 2021Software engineering, systems software and programming languages · 4 · 3 first-authorDatabases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2024 ARISE: AI Right Sizing Engine for AI workload configurations
abstract
Data scientists and platform engineers who maintain AI stacks are required to continuously run AI workloads. When executing any part of the AI pipeline, whether data preprocessing, training, fine-tuning or inference, a frequent question is how to optimally configure the environment to meet Service Level Objectives (SLOs), such as desired throughput, runtime deadlines, and avoid memory and CPU exhaustion. We present ARISE, a tool that enables making data-driven decisions about AI workload configuration questions. ARISE trains performance prediction machine-learning regression models on historical workloads and performance benchmark metadata, and then predicts the performance of future workloads based on their input metadata, using the best performing regression models. Initial evaluation of ARISE on real-world workloads shows high prediction accuracy.
Rachel Tzoref, Bruno Wassermann, Eran Raichstein, Dean H. Lorenz
SYSTOR2
2022 Hybrid anomaly detection and prioritization for network logs at cloud scale
abstract
Monitoring the health of large-scale systems requires significant manual effort, usually through the continuous curation of alerting rules based on keywords, thresholds and regular expressions, which might generate a flood of mostly irrelevant alerts and obscure the actual information operators would like to see. Existing approaches try to improve the observability of systems by intelligently detecting anomalous situations. Such solutions surface anomalies that are statistically significant, but may not represent events that reliability engineers consider relevant. We propose ADEPTUS, a practical approach for detection of relevant health issues in an established system. ADEPTUS combines statistics and unsupervised learning to detect anomalies with supervised learning and heuristics to determine which of the detected anomalies are likely to be relevant to the Site Reliability Engineers (SREs). ADEPTUS overcomes the labor-intensive prerequisite of obtaining anomaly labels for supervised learning by automatically extracting information from historic alerts and incident tickets. We leverage ADEPTUS for observability in the network infrastructure of IBM Cloud. We perform an extensive real-world evaluation on 10 months of logs generated by tens of thousands of network devices across 11 data centers and demonstrate that ADEPTUS achieves higher alerting accuracy than the rule-based log alerting solution, curated by domain experts, used by SREs daily.
David Ohana, Bruno Wassermann, Nicolas Dupuis, Elliot K. Kolodner, Eran Raichstein, Michal Malka
EuroSys2
2021 DeCorus-NSA: detection and correlation of unusual signals for network syslog analytics
abstract
The management of large data centre (DC) network infrastructure confronts Network Reliability Engineers (NRE) with challenges. A single DC at a modern cloud services provider can host thousands of network devices. The syslog messages generated by these devices are an important type of monitoring data to detect and diagnose failures. Devices in a single DC produce millions of syslog messages per day in a variety of formats.
David Ohana, Bruno Wassermann, Moshe Hershcovitch, Elliot K. Kolodner, Michal Malka, Eran Raichstein, Ronen Schaffer, Robert Shahla
SYSTOR2
2011 Monere: Monitoring of Service Compositions for Failure Diagnosis
Bruno Wassermann, Wolfgang Emmerich
ICSOC1
2010 Improving wide-area distributed system availability
abstract
The Software-as-a-Service (SaaS) paradigm and corresponding service-oriented technologies have simplified the development of larger, more complex software systems that routinely span administrative and organisational boundaries. These systems inhabit a complex operating environment with numerous threats to the dependability of service compositions. These threats include many system-level failures whose causes are difficult and time-consuming to determine. It is difficult to detect vulnerabilities to these failures prior to deployment of an application into production and applications are currently not well-equipped to handle them effectively. This results in lengthy downtimes of production systems and hence low availability. The goal of this PhD is to increase the availability of such systems by eliminating as many failures as possible before deployment and by assisting administrators to diagnose their causes more efficiently. We propose a novel monitoring technique and apply failure injection techniques that target these difficult failures and enable separate administrative domains to cooperate in handling them. Furthermore, we investigate the extent to which we can equip these systems to be self-diagnosing.
Bruno Wassermann
ICSE (2)1
2009 Distributed Cross-Domain Change Management
abstract
Distributed systems increasingly span organizational boundaries and, with this, system and service management domains. Web services are the primary means of exposing services to clients, be it in electronic commerce, Software-as-a-Service (SaaS) or on cloud platforms and are being used and integrated with customer-managed applications as well as in complex mashups. Maturing cross-domain relationships and an increase in loose coupling and ad-hocness makes managing configuration changes, e.g., changes in interfaces or endpoints, increasingly relevant. Traditional service management processes within organizations, in particular change management, relies on a central configuration management database (CMDB) to assess the impact a change has on other components of the system. However, this approach does not work in a cross-domain environment, due to the lack of a central CMDB, centralized management processes, and knowledge by service providers which clients depends on their respective services. This paper proposes the Change 2.0 approach to cross-domain change management based on an inversion of responsibility for impact assessment and the facilitation of cross-domain service process integration. We present the requirements imposed by cross-domain change management, the Change 2.0 architecture, and a brief evaluation of its benefits.
Bruno Wassermann, Heiko Ludwig, Jim Laredo, Kamal Bhattacharya, Liliana Pasquale
ICWS1
2009 REST-based management of loosely coupled services
abstract
Applications increasingly make use of the distributed platform that the World Wide Web provides - be it as a Software-as-a-Service such as salesforce.com, an application infrastructure such as facebook.com, or a computing infrastructure such as a "cloud". A common characteristic of applications of this kind is that they are deployed on infrastructure or make use of components that reside in different management domains. Current service management approaches and systems, however, often rely on a centrally managed configuration management database (CMDB), which is the basis for centrally orchestrated service management processes, in particular change management and incident management. The distribution of management responsibility of WWW based applications requires a decentralized approach to service management. This paper proposes an approach of decentralized service management based on distributed configuration management and service process co-ordination, making use RESTful access to configuration information and ATOM-based distribution of updates as a novel foundation for service management processes.
Heiko Ludwig, Jim Laredo, Kamal Bhattacharya, Liliana Pasquale, Bruno Wassermann
WWW5
2006 Web service orchestration with BPEL
abstract
No abstract available.
Bruno Wassermann, Wolfgang Emmerich, Howard Foster
ICSE2
2005 Grid Service Orchestration Using the Business Process Execution Language (BPEL)
Wolfgang Emmerich, Ben Butchart, Bruno Wassermann, Sarah L. Price
J. Grid Comput.4