VLDB 2026 Research / reviewers in the wild / expert
Liran Schour
dblp:21/9709
· DBLP profile ↗
9ranked-venue papers
0as first author
4since 2021 · last 2025
0000-0002-6163-0060ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 2 since 2021Computer networks · 3 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
High-performance computing · 26% GPUs and heterogeneous computing · 26% Cloud and datacenter computing · 20% | |
| Computer networks
1 paper |
Software-defined and programmable networks · 100% | |
| Network and information security
1 paper |
Network security · 100% |
Topics — the 6 heaviest of 7, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
GPUs and heterogeneous computing › GPU communication
GPU networking |
0.9 | 1 | 2025 | Vela: A Virtualized LLM Training System with GPU Direct RoCE · ASPLOS (2) 2025 |
High-performance computing › large-scale training
large language model training |
0.9 | 1 | 2025 | Vela: A Virtualized LLM Training System with GPU Direct RoCE · ASPLOS (2) 2025 |
Software-defined and programmable networks
programmable data plane |
0.7 | 1 | 2023 | Janus: An Experimental Reconfigurable SmartNIC with P4 Programmability and SDN Isolation · FPGA 2023 |
Cloud and datacenter computing
cloud infrastructure |
0.7 | 1 | 2023 | Janus: An Experimental Reconfigurable SmartNIC with P4 Programmability and SDN Isolation · FPGA 2023 |
Electronic design automation
hardware/software co-design |
0.7 | 1 | 2023 | Janus: An Experimental Reconfigurable SmartNIC with P4 Programmability and SDN Isolation · FPGA 2023 |
Parallel and multicore computing › parallelization strategies
model parallelism |
0.3 | 1 | 2025 | Vela: A Virtualized LLM Training System with GPU Direct RoCE · ASPLOS (2) 2025 |
Methods — techniques the papers use, named apart from their topics
p4 · 2.0hardware-enforced isolation · 2.0hardware offloading · 2.0peer-to-peer DMA · 0.9SR-IOV · 0.9KVM virtualization · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Routing Strategies for RoCE Networks in AI CloudsabstractThe rapid explosion of Artificial Intelligence (AI) workloads utilizing a growing number of accelerators has placed unprecedented demand on the network. These workloads typically leverage Remote Direct Memory Access (RDMA) and require a high-performance network fabric. While many purpose-built cloud networking solutions can provide high performance, efficiently utilizing these costly infrastructures require a fabric that is multi-tenant for ease of consumption. Furthermore, the fabric must be resilient to faults and for operational manageability. Resilient cloud networks typically employ mature Ethernet segmentation techniques over Clos topologies with Equal Cost Multi Path (ECMP) routing. ECMP hashes flows to paths, which in case of collisions can significantly degrade performance for large RDMA over Converged Ethernet (RoCE) flows. To mitigate ECMP penalties, we evaluate routing strategies with varying levels of operational complexity. We explore load balancing and path pinning solutions that leverage non-proprietary, mature technologies over commodity Ethernet. Our evaluation follows a three-fold strategy, focusing on the key dimensions of performance, resiliency, and operational complexity. By applying this methodology to representative implementations, we highlight the trade-offs. While all techniques are resilient, path pinning-based solutions excel at performance but introduce greater complexity. Specifically, path pinning achieves up to 1.6× improvement over ECMP for RoCE test traffic and up to 2.5× for NCCL AllReduce. These results validate the promising performance benefits of path pinning and highlight the need to explore less complex implementations for broader adoption. Our methodology can be used to rigorously evaluate future implementations in support of AI network design. Abdul Alim, Ali Sydney, Liran Schour, Abdullah Kayi, Laurent Schares, Pavlos Maniotis, Bengi Karaçali |
CLOUD | 3 |
| 2025 | Vela: A Virtualized LLM Training System with GPU Direct RoCEabstractVela is a cloud-native system designed for LLM training workloads built using off-the-shelf hardware, Linux KVM-based virtualization, and a virtualized RDMA over Converged Ethernet (RoCE) network. Vela virtual machines (VMs) support peer-to-peer DMA between the GPUs and SRIOV-based network interface. In this paper, we share Vela's key architectural aspects with details from an NVIDIA A100 GPU-based deployment in one of the IBM Cloud data centers. Throughout the paper, we share insights and experiences from designing, building, and operating the system over a ~2.5 year timeframe to highlight the capabilities of readily available software and hardware technologies and the improvement opportunities for future AI systems, thereby making AI infrastructure more accessible to a broader community. As we evaluated the system for performance at ~1500 GPU scale, we achieved ~80% of the ideal throughput while training a 50 billion parameter decoder model using model parallelism, and ~70% per GPU FLOPS compared to a single VM with the High-Performance Linpack benchmark. Apoorve Mohan, Robert Walkup, Bengi Karaçali, Ming-Hung Chen, Abdullah Kayi, Liran Schour, Shweta Salaria, Sophia Wen, I-Hsin Chung, Abdul Alim, Constantinos Evangelinos, Lixiang Luo, Marc Dombrowa, Laurent Schares, Ali Sydney, Pavlos Maniotis, Sandhya Koteshwara, Brent Tang, Joel Belog, Rei Odaira, Vasily Tarasov, Eran Gampel, Drew Thorstensen, Talia Gershon, Seetharami Seelam |
ASPLOS (2) | 6 |
| 2025 | FlowTracer: A Tool for Uncovering Network Path Usage Imbalance in AI Training Clusters
Hasibul Jamil, Abdul Alim, Laurent Schares, Pavlos Maniotis, Liran Schour, Ali Sydney, Abdullah Kayi, Tevfik Kosar, Bengi Karaçali |
ICC | 5 |
| 2023 | Janus: An Experimental Reconfigurable SmartNIC with P4 Programmability and SDN IsolationabstractDisparate deployment models of cloud computing pose varying requirements on cloud infrastructure components such as networking, storage, provisioning, and security. Infrastructure providers need to study these and often create custom infrastructure components to satisfy these requirements. A major challenge in the research and development of these cloud infrastructure solutions, however, is the availability of customizable platforms for experimentation and trade-off analysis of the various hardware and software components. Most platforms are either general purpose or bespoke solutions created to assist a particular task, too rigid to allow meaningful customization. In this work, we present a 100G reconfigurable smartNIC prototyping platform called Janus that enables cloud infrastructure research and hardware-software co-design of infrastructure components such as hypervisor, secure boot, software defined networking and distributed storage. The platform provides a path to optimize the stack by offloading the functionalities from the host x86 to the embedded processor on the smartNIC and optimize performance by moving pieces to hardware using P4. Further, our platform provides hardware-enforced isolation of cloud network control plane, thereby securing the control plane from the tenants even for bare-metal deployments. Bharat Sukhwani, Mohit Kapur, Alda Ohmacht, Liran Schour, Martin Ohmacht, Chris Ward, Chuck Haymes, Sameh W. Asaad |
FPGA | 4 |
| 2017 | CogNETive: insights and visualization for operations@scaleabstractOperating a cloud-scale service is a huge challenge. There are millions of users worldwide and millions of requests per seconds. For example, Amazon's Simple Storage Service (S3) in 2013 contained two trillion objects and its logs contained 1.1 million log lines per second, which are approximately 10 PB of log records per year (see [1]). Cloud scale implies thousands of servers and network elements, and hundreds of services from multiple cross-regional data centers. Cloud service operation data is scattered over various types of semi-structured and unstructured logs (e.g., application, error, debug), telemetry and network data, as well as customer service records. It is therefore extremely difficult for the multiple owners and administrators in such systems, coming from different units of the organization, to follow the possible paths and system alternatives in order to detect problems, solve issues and understand the service operation. Dean H. Lorenz, Eran Raichstein, Katherine Barabash, Hillel Kolodner, Liran Schour, Shelly Garion |
SYSTOR | 5 |
| 2015 | Networking Architecture for Seamless Cloud InteroperabilityabstractSeamless cloud interoperability is highly desired but not yet easily attainable in the current cloud solutions market. This work tackles one aspect of achieving cloud interoperability, namely, inter-cloud networking. We list the requirements and propose an inter-cloud networking architecture for a case of independent clouds owned by different entities and powered by different cloud management and network virtualization technologies. Then we validate the proposed architecture by describing an example of working implementation for Open Stack cloud powered by Open Daylight Open DOVE SDN solution. Finally, we compare our architecture to the existing solutions. Anna Levin, Katherine Barabash, Yaniv Ben-Itzhak, Sergey Guenender, Liran Schour |
CLOUD | 5 |
| 2013 | An intent-based approach for network virtualization
Rami Cohen, Katherine Barabash, Benny Rochwerger, Liran Schour, Daniel Crisan, Robert Birke, Cyriel Minkenberg, Mitchell Gusat, Renato Recio, Vinit Jain |
IM | 4 |
| 2013 | Distributed Overlay Virtual Ethernet (DOVE) integration with Openstack
Rami Cohen, Katherine Barabash, Liran Schour |
IM | 3 |
| 2011 | Inter-cloud mobility of virtual machinesabstractCloud computing is increasingly gaining inroads among a variety of organizational users. As clouds are introduced for use by enterprises, service providers, and governmental and educational entities, new challenges related to the interconnection between such clouds emerge. Cloud administrators seek to maintain acceptable levels of autonomy and control over their cloud infrastructure, while ensuring the integrity of the cloud services. At the same time, they are expected to enable cross-cloud services, including mobility of workloads between clouds. We present the design and implementation of a technology that enables live mobility of virtual machines between clouds, while enforcing the cloud insularity requirements of autonomy, privacy, and security. We also provide an empirical evaluation of our solution, demonstrating its viability and compliance with requirements. Kenneth Nagin, David Hadas, Zvi Dubitzky, Alex Glikson, Irit Loy, Benny Rochwerger, Liran Schour |
SYSTOR | 7 |