Liran Schour

dblp:21/9709 · DBLP profile ↗
← Back
9ranked-venue papers
0as first author
4since 2021 · last 2025
0000-0002-6163-0060ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 2 since 2021Computer networks · 3 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
High-performance computing · 26% GPUs and heterogeneous computing · 26% Cloud and datacenter computing · 20%
Computer networks
1 paper
Software-defined and programmable networks · 100%
Network and information security
1 paper
Network security · 100%

Topics — the 6 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
GPUs and heterogeneous computing › GPU communication
GPU networking
0.912025
Vela: A Virtualized LLM Training System with GPU Direct RoCE · ASPLOS (2) 2025
High-performance computing › large-scale training
large language model training
0.912025
Vela: A Virtualized LLM Training System with GPU Direct RoCE · ASPLOS (2) 2025
Software-defined and programmable networks
programmable data plane
0.712023
Janus: An Experimental Reconfigurable SmartNIC with P4 Programmability and SDN Isolation · FPGA 2023
Cloud and datacenter computing
cloud infrastructure
0.712023
Janus: An Experimental Reconfigurable SmartNIC with P4 Programmability and SDN Isolation · FPGA 2023
Electronic design automation
hardware/software co-design
0.712023
Janus: An Experimental Reconfigurable SmartNIC with P4 Programmability and SDN Isolation · FPGA 2023
Parallel and multicore computing › parallelization strategies
model parallelism
0.312025
Vela: A Virtualized LLM Training System with GPU Direct RoCE · ASPLOS (2) 2025

Methods — techniques the papers use, named apart from their topics

p4 · 2.0hardware-enforced isolation · 2.0hardware offloading · 2.0peer-to-peer DMA · 0.9SR-IOV · 0.9KVM virtualization · 0.9
YearPublicationVenuePosition
2025 Routing Strategies for RoCE Networks in AI Clouds
abstract
The rapid explosion of Artificial Intelligence (AI) workloads utilizing a growing number of accelerators has placed unprecedented demand on the network. These workloads typically leverage Remote Direct Memory Access (RDMA) and require a high-performance network fabric. While many purpose-built cloud networking solutions can provide high performance, efficiently utilizing these costly infrastructures require a fabric that is multi-tenant for ease of consumption. Furthermore, the fabric must be resilient to faults and for operational manageability. Resilient cloud networks typically employ mature Ethernet segmentation techniques over Clos topologies with Equal Cost Multi Path (ECMP) routing. ECMP hashes flows to paths, which in case of collisions can significantly degrade performance for large RDMA over Converged Ethernet (RoCE) flows. To mitigate ECMP penalties, we evaluate routing strategies with varying levels of operational complexity. We explore load balancing and path pinning solutions that leverage non-proprietary, mature technologies over commodity Ethernet. Our evaluation follows a three-fold strategy, focusing on the key dimensions of performance, resiliency, and operational complexity. By applying this methodology to representative implementations, we highlight the trade-offs. While all techniques are resilient, path pinning-based solutions excel at performance but introduce greater complexity. Specifically, path pinning achieves up to 1.6× improvement over ECMP for RoCE test traffic and up to 2.5× for NCCL AllReduce. These results validate the promising performance benefits of path pinning and highlight the need to explore less complex implementations for broader adoption. Our methodology can be used to rigorously evaluate future implementations in support of AI network design.
Abdul Alim, Ali Sydney, Liran Schour, Abdullah Kayi, Laurent Schares, Pavlos Maniotis, Bengi Karaçali
CLOUD3
2025 Vela: A Virtualized LLM Training System with GPU Direct RoCE
abstract
Vela is a cloud-native system designed for LLM training workloads built using off-the-shelf hardware, Linux KVM-based virtualization, and a virtualized RDMA over Converged Ethernet (RoCE) network. Vela virtual machines (VMs) support peer-to-peer DMA between the GPUs and SRIOV-based network interface. In this paper, we share Vela's key architectural aspects with details from an NVIDIA A100 GPU-based deployment in one of the IBM Cloud data centers. Throughout the paper, we share insights and experiences from designing, building, and operating the system over a ~2.5 year timeframe to highlight the capabilities of readily available software and hardware technologies and the improvement opportunities for future AI systems, thereby making AI infrastructure more accessible to a broader community. As we evaluated the system for performance at ~1500 GPU scale, we achieved ~80% of the ideal throughput while training a 50 billion parameter decoder model using model parallelism, and ~70% per GPU FLOPS compared to a single VM with the High-Performance Linpack benchmark.
Apoorve Mohan, Robert Walkup, Bengi Karaçali, Ming-Hung Chen, Abdullah Kayi, Liran Schour, Shweta Salaria, Sophia Wen, I-Hsin Chung, Abdul Alim, Constantinos Evangelinos, Lixiang Luo, Marc Dombrowa, Laurent Schares, Ali Sydney, Pavlos Maniotis, Sandhya Koteshwara, Brent Tang, Joel Belog, Rei Odaira, Vasily Tarasov, Eran Gampel, Drew Thorstensen, Talia Gershon, Seetharami Seelam
ASPLOS (2)6
2025 FlowTracer: A Tool for Uncovering Network Path Usage Imbalance in AI Training Clusters
Hasibul Jamil, Abdul Alim, Laurent Schares, Pavlos Maniotis, Liran Schour, Ali Sydney, Abdullah Kayi, Tevfik Kosar, Bengi Karaçali
ICC5
2023 Janus: An Experimental Reconfigurable SmartNIC with P4 Programmability and SDN Isolation
abstract
Disparate deployment models of cloud computing pose varying requirements on cloud infrastructure components such as networking, storage, provisioning, and security. Infrastructure providers need to study these and often create custom infrastructure components to satisfy these requirements. A major challenge in the research and development of these cloud infrastructure solutions, however, is the availability of customizable platforms for experimentation and trade-off analysis of the various hardware and software components. Most platforms are either general purpose or bespoke solutions created to assist a particular task, too rigid to allow meaningful customization. In this work, we present a 100G reconfigurable smartNIC prototyping platform called Janus that enables cloud infrastructure research and hardware-software co-design of infrastructure components such as hypervisor, secure boot, software defined networking and distributed storage. The platform provides a path to optimize the stack by offloading the functionalities from the host x86 to the embedded processor on the smartNIC and optimize performance by moving pieces to hardware using P4. Further, our platform provides hardware-enforced isolation of cloud network control plane, thereby securing the control plane from the tenants even for bare-metal deployments.
Bharat Sukhwani, Mohit Kapur, Alda Ohmacht, Liran Schour, Martin Ohmacht, Chris Ward, Chuck Haymes, Sameh W. Asaad
FPGA4
2017 CogNETive: insights and visualization for operations@scale
abstract
Operating a cloud-scale service is a huge challenge. There are millions of users worldwide and millions of requests per seconds. For example, Amazon's Simple Storage Service (S3) in 2013 contained two trillion objects and its logs contained 1.1 million log lines per second, which are approximately 10 PB of log records per year (see [1]). Cloud scale implies thousands of servers and network elements, and hundreds of services from multiple cross-regional data centers. Cloud service operation data is scattered over various types of semi-structured and unstructured logs (e.g., application, error, debug), telemetry and network data, as well as customer service records. It is therefore extremely difficult for the multiple owners and administrators in such systems, coming from different units of the organization, to follow the possible paths and system alternatives in order to detect problems, solve issues and understand the service operation.
Dean H. Lorenz, Eran Raichstein, Katherine Barabash, Hillel Kolodner, Liran Schour, Shelly Garion
SYSTOR5
2015 Networking Architecture for Seamless Cloud Interoperability
abstract
Seamless cloud interoperability is highly desired but not yet easily attainable in the current cloud solutions market. This work tackles one aspect of achieving cloud interoperability, namely, inter-cloud networking. We list the requirements and propose an inter-cloud networking architecture for a case of independent clouds owned by different entities and powered by different cloud management and network virtualization technologies. Then we validate the proposed architecture by describing an example of working implementation for Open Stack cloud powered by Open Daylight Open DOVE SDN solution. Finally, we compare our architecture to the existing solutions.
Anna Levin, Katherine Barabash, Yaniv Ben-Itzhak, Sergey Guenender, Liran Schour
CLOUD5
2013 An intent-based approach for network virtualization
Rami Cohen, Katherine Barabash, Benny Rochwerger, Liran Schour, Daniel Crisan, Robert Birke, Cyriel Minkenberg, Mitchell Gusat, Renato Recio, Vinit Jain
IM4
2013 Distributed Overlay Virtual Ethernet (DOVE) integration with Openstack
Rami Cohen, Katherine Barabash, Liran Schour
IM3
2011 Inter-cloud mobility of virtual machines
abstract
Cloud computing is increasingly gaining inroads among a variety of organizational users. As clouds are introduced for use by enterprises, service providers, and governmental and educational entities, new challenges related to the interconnection between such clouds emerge. Cloud administrators seek to maintain acceptable levels of autonomy and control over their cloud infrastructure, while ensuring the integrity of the cloud services. At the same time, they are expected to enable cross-cloud services, including mobility of workloads between clouds. We present the design and implementation of a technology that enables live mobility of virtual machines between clouds, while enforcing the cloud insularity requirements of autonomy, privacy, and security. We also provide an empirical evaluation of our solution, demonstrating its viability and compliance with requirements.
Kenneth Nagin, David Hadas, Zvi Dubitzky, Alex Glikson, Irit Loy, Benny Rochwerger, Liran Schour
SYSTOR7