VLDB 2026 Research / reviewers in the wild / expert
Ryota Kawashima
dblp:94/1132
· DBLP profile ↗
13ranked-venue papers
7as first author
4since 2021 · last 2024
0000-0002-4025-6970ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 3 · 3 first-author · 1 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Meeting Latency and Jitter Demands of Beyond 5G Networking Era: Are CNFs Up to the Challenge?abstractThe introduction of Network Function Virtualization (NFV) has shifted network processing from specialized hardware to more flexible commodity servers, and this transition is still evolving. New industrial applications, the Internet of Things (IoT), and technologies like augmented, virtual, and mixed reality (AR/VR/MR) require networks that can handle event-based operations and middleware with very low and predictable latency. These requirements pose performance optimization challenges for packet processing in a layered infrastructure. In this study, we look deep into the challenges of implementing such network infrastructures using general-purpose hardware, a strategy motivated by its flexibility to realize telco-cloud and the potential to reduce electronic waste. Focusing on NFV with an emphasis on containerized network functions (CNFs), we investigate the performance limitations, particularly the high jitter and throughput variation observed in packet forwarding. We used a network function (NF) implemented using defacto industry standard user-space I/O architecture DPDK in bare-metal and containerized environments for performance evaluation. We conducted ten experiments in a 40 GbE environment to measure throughput, latency, and jitter across various packet sizes, traffic rates, and system configurations. The results indicate that adjusting CPU settings can significantly enhance throughput for CNFs despite a potential increase in jitter. We found that CNFs are feasible for latency-sensitive tasks, particularly under conditions of low traffic and specific packet sizes. With careful system-level configuration, CNFs can be used in beyond 5G cloud-native networking, offering promising potential for latency-sensitive applications. Adil Bin Bhutto, Ryota Kawashima, Yuzo Taenaka, Youki Kadobayashi |
COMPSAC | 2 |
| 2023 | Understanding Roadblocks in Virtual Network I/O: A Comprehensive Analysis of CPU Cache UsageabstractOne big challenge in cloud-native network functions (CNFs) is the poor performance of packet forwarding. In this paper, we comprehensively analyze CPU cache usage in regard to the interprocess communication (vhost-user) as well as the packet I/O framework (DPDK). We cover four high-end CPUs, and examine 141 implementation/configuration patterns on an EIVU (Essential Implementation of Vhost-User) platform to clarify how the patterns affect the cache behaviors. The result pinpoints the roadblocks in a generalized form; cache invalidations stemming from three design/implementation factors are the major bottle-necks in vhost-user/DPDK, and shows potential of 100+ Mpps container networking for future software-centric environment. Daichi Takeya, Ryota Kawashima, Hiroki Nakayama, Tsunemasa Hayashi, Hiroshi Matsuo |
NetSoft | 2 |
| 2021 | A Vision to Software-Centric Cloud Native Network Functions: Achievements and ChallengesabstractNetwork slicing qualitatively transforms network infrastructures such that they have maximum flexibility in the context of ever-changing service requirements. While the agility of cloud native network functions (CNFs) demonstrates significant promise, virtualization and softwarization severely degrade the performance of such network functions. Considerable efforts were expended to improve the performance of virtualized systems, and at this stage 10 Gbps throughput is a real target even for container/VM-based applications. Nonetheless, the current performance of CNFs with state-of-the-art enhancements does not meet the performance requirements of next-generation 6G networks that aim for terabit-class throughput. The present pace of performance enhancements in hardware indicates that straightforward optimization of existing system components has limited possibility of filling the performance gap. As it would be reasonable to expect a single silver-bullet technology to dramatically enhance the ability of CNFs, an organic integration of various data-plane technologies with a comprehensive vision is a potential approach. In this paper, we show a future vision of system architecture for terabit-class CNFs based on effective harmonization of the technologies within the wide-range of network systems consisting of commodity hardware devices. We focus not only on the performance aspect of CNFs but also other pragmatic aspects such as interoperability with the current environment (not clean slate). We also highlight the remaining missing-link technologies revealed by the goal-oriented approach. Ryota Kawashima |
HPSR | 1 |
| 2021 | Software Physical/Virtual Rx Queue Mapping Toward High-Performance Containerized NetworkingabstractSoftwarization of Network Functions (NFs) accelerates automated deployment and management of services on next-gen networks. Combining flexibility and high-performance is a vital requirement for Network Functions Virtualisation (NFV); however, many studies have demonstrated that containerization or virtualization of NFs severely degrades the fundamental efficiency of packet forwarding. Virtual network I/O, a mechanism of packet transferring between a guest and the host, has been seen as the performance bottleneck in the PVP (Physical-Virtual-Physical) datapath, and one of the main causes of this deterioration is packet copy between them. Various techniques, such as zero-copy, pass-through, and hardware offloading, have been examined to alleviate the performance overhead. However, existing designs and implementations incur pragmatic issues, such as compatibility, manageability, and insufficient quality of performance. We propose yet another design and implementation of zero-copy/pass-through acceleration (named IOVTee) to resolve real-world problems as well as to enhance the forwarding efficiency. IOVTee takes advantage of pre-processing of virtual switches with achieving zero-copy on the receive (Rx) path. The pluggable style of IOVTee for vhost-user (the de-facto virtual network I/O) enables our approach to be transparent to both containers/VMs and virtual switches. In this article, we explain the heart of IOVTee, a fully software-based Rx queue mapping mechanism (between physical and virtual) that enables a concept of Virtual DMA Write-through (to the NF). Our evaluation results showed that applying IOVTee to vhost-user drastically increased efficiency of packet forwarding in the PVP datapath (by 45% and 98% for traffic of 64-byte and 1514-byte packets respectively). Ryota Kawashima |
IEEE Trans. Netw. Serv. Manag. | 1 |
| 2019 | NFV-VIPP: Catching Internal Figures of Packet Processing for Accelerating Development and Operations of NFV-nodesabstractServer-based NFV-nodes have disparate internals, such as simultaneous deployment of Virtual Network Functions (VNFs) and layered software abstractions including a virtual switch. The traditional operations tailored for function-hardware-coupled devices cannot cope with the increase of related components as well as complicated packet forwarding paths inside. Besides, self-development of VNFs attracting Telcos is still highly complicated work, due to lack of exact troubleshooting of internal NFV-nodes caused by exclusive resource management by Data-Plane Development Kit (DPDK). OPNFV Barometer provides means of stats acquisition, but internal figures of packet processing are still unveiled. In this paper, we propose an integrated metrics collection framework (NFV-VIPP) specialized to NFV-nodes. NFV-VIPP provides seamless understandings of system components in a node, and reveals the inside by transparently exposing implementation-related metrics. NFV-VIPP can be incorporated into Barometer/collected via RESTful APIs to reinforce system visibility, meaning that our framework bridges NFV-node internals to existing management frameworks. We explore NFV-node management using intra-VNF metrics obtained by NFVVIPP. Specifically, we prove that CPU-cycle consumption of inter-receive-polling is a driving force to estimate system load. Masahiro Dodare, Yuki Taguchi, Ryota Kawashima, Hiroki Nakayama, Tsunemasa Hayashi, Hiroshi Matsuo |
CNSM | 3 |
| 2017 | PA-Flow: Gradual Packet Aggregation at Virtual Network I/O for Efficient Service ChainingabstractA logical chaining of Virtual Network Functions (VNFs) enables toy-blocking style composition of various network services. VNFs are generally based on Virtual Machines (VMs) running on commodity servers. However, such VM-based VNFs have poor packet forwarding performance because of the virtual network I/O overhead, and therefore virtualization of the core network functions can degrade overall network performance. In this paper, we propose a novel packet aggregation method (PA-Flow) that aggregates packets having a same next-hop VNF with no further latency. Unlike existing packet aggregation approaches, our method is effective for the virtual network I/O to reduce packet processing cost inside the host. Moreover, aggregated packets are further aggregated as going through the service chain, which drastically reduces traffic amount at the core network. We have implemented a PA-Flow module that works between a virtual switch and QEMU under the DPDK/vhost-user framework. PA-Flow showed 170% higher throughput than that of the default DPDK/vhost-user based system for a 2-VNF service chain. In addition, latency was also improved by about 2 us because of the packet amount reduction. Yuki Taguchi, Ryota Kawashima, Hiroki Nakayama, Tsunemasa Hayashi, Hiroshi Matsuo |
CloudCom | 2 |
| 2017 | Evaluation of Forwarding Efficiency in NFV-Nodes Toward Predictable Service Chain PerformanceabstractThe concept of network functions virtualization (NFV) has been embodied in commercial networks over the past years. Software-based virtual network functions have forwarding performance concerns in general, and various acceleration technologies have been developed so far, such as DPDK and vhost-user. Existence of several alternatives requires network engineers or operators to select appropriate technologies; however, no pragmatic criterion exists for constructing high-performance NFV-nodes. From their points of view, a lack of common benchmark and understanding of performance characteristics makes it difficult to predict hop-by-hop performance in a service chain, which results in prevention of NFV deployment in mission-critical networks. In this paper, we clarify performance characteristics of packet forwarding in NFV nodes focusing on three types of acceleration technologies; packet I/O architecture, virtual network I/O, and forwarding engine in a practical stage. We examined three packet I/O architectures (NAPI, netmap, and DPDK), three virtual I/O mechanisms (vhost-net, vhost-user, and SR-IOV), and four practical forwarding programs (Open vSwitch, OVS-DPDK, xDPd-DPDK, and Lagopus) with three referential programs (Linux Bridge, VALE, and L2FWD-DPDK). The experiment was conducted on a 40 GbE environment and we examined two device-under-test machines having different CPU performance. We argue performance characteristics of each technology and give quantitative analyses of the result. The key findings are: 1) CPU core speed has impact on both throughput and latency/jitter; 2) DPDK can allow performance prediction; 3) vhost-user is appropriate for real environment; and 4) OVS-DPDK provides a good combination of performance and functionality. Ryota Kawashima, Hiroki Nakayama, Tsunemasa Hayashi, Hiroshi Matsuo |
IEEE Trans. Netw. Serv. Manag. | 1 |
| 2015 | SCLP: Segment-oriented Connection-less Protocol for high-performance software tunneling in datacenter networksabstractThe notion of Software-Defined Networking (SDN) has already been introduced into cloud datacenter networks for provisioning virtual network environment. Network virtualization of today is generally achieved by L2-in-L3 tunneling protocols like VXLAN (Virtual eXtensible LAN) and NVGRE (Network Virtualization using Generic Routing Encapsulation) in public cloud datacenters. Some leading production packages for network virtualization have adopted an Edge-Overlay model that performs tunnel encapsulation and decapsulation processes at high-functional virtual switches to utilize existing network equipment. However, a severe performance problem arises because of the software-based tunneling processes. Alternatively, the STT (Stateless Transport Tunneling) protocol overcomes the problem by modifying the semantics of the TCP header, but such changes in semantics raises pragmatic issues in that network middleboxes can discard STT packets as an anomaly. In this paper, we propose a novel layer 4 protocol (Segment-oriented Connection-less Protocol, SCLP) for existing tunneling protocols such as VXLAN and NVGRE. SCLP is designed to not only accelerate the throughput of tunneling protocols, but prevent the packet discarding problem by providing a single-semantic header. Specifically, SCLP can exploit GRO (Generic Receive Offload) feature supported by the Linux kernel to reduce the number of packets to be software-interrupted. We implemented the SCLP protocol and applied it to the VXLAN protocol instead of UDP. As a result, the throughput of the VXLAN over SCLP protocol was almost doubled to the original UDP-based one at maximum. Ryota Kawashima, Shin Muramatsu, Hiroki Nakayama, Tsunemasa Hayashi, Hiroshi Matsuo |
NetSoft | 1 |
| 2014 | Implementation and Performance Analysis of STT Tunneling Using vNIC Offloading Framework (CVSW)abstractNetwork Virtualization Overlays (NVO3) provides multi-tenancy services in cloud data centers with existing networking equipment. IP tunneling is an essential technology to logically separate each virtual traffic, in particular, Stateless Transport Tunneling (STT) is considered to achieve better performance using TCP Segmentation Offload (TSO) feature. Currently, there is no openly available implementation of STT, and its implementation and performance characteristics have not been studied in academic field so far. We have therefore implemented STT protocol and conducted performance evaluation by comparing with VXLAN protocol. In practice, the STT implementation has been done using a virtual NIC offloading framework, co-virtual switch (CVSW). CVSW is a software component that extends virtual NICs and provides high-level packet processing framework such as Open Flow Match-Action. In this paper, we describe implementation details of STT and performance evaluation results from various perspectives. The results showed that the actual performance of STT was almost equal to non-tunneling VM-to-VM communication and was two-times higher than that of VXLAN. Furthermore, we clarify the high-performance nature of STT is brought from both byte-stream characteristic of TCP and Generic Receive Offload (GRO) feature rather than widely believed TSO. Ryota Kawashima, Hiroshi Matsuo |
CloudCom | 1 |
| 2014 | VSE: Virtual Switch Extension for Adaptive CPU Core Assignment in SoftirqabstractAn Edge-Overlay model constructing virtual networks using both virtual switches and IP tunnels is promising in cloud datacenter networks. But software-implemented virtual switches can cause performance problems because the packet processing load is concentrated on a particular CPU core. Although multi queue functions like Receive Side Scaling (RSS) can distribute the load onto multiple CPU cores, there are still problems to be solved such as IRQ core collision of heavy traffic flows as well as competitive resource use between physical and virtual for packet processing. In this paper, we propose a software packet processing unit named VSE (Virtual Switch Extension) to address these problems by adaptively determining softirq cores based on both CPU load and VM-running information. Furthermore, the behavior of VSE can be managed by Open Flow controllers. Our performance evaluation results showed that throughput of our approach was higher than an existing RSSbased model as packet processing load increased. In addition, we show that our method prevented performance of high-loaded flows from being degraded by priority-based CPU core selection. Shin Muramatsu, Ryota Kawashima, Shoichi Saito, Hiroshi Matsuo |
CloudCom | 2 |
| 2013 | Non-tunneling Edge-Overlay Model Using OpenFlow for Cloud Datacenter NetworksabstractIn current SDN paradigm, an edge-overlay (distributed tunneling) model using L2-in-L3 tunneling protocols, such as VXLAN, has attracted attentions for multi-tenant data center networks. The edge-overlay model can establish rapid-deployment of virtual networks onto existing traditional network facilities, ensure flexible IP/MAC address allocation to VMs, and extend the number of virtual networks regardless of the VLAN ID limitation. However, such model has performance and incompatibility problems on the traditional network environment. For L2 data center networks, this paper proposes a pure software approach that uses Open Flow virtual switches to realize yet another edge-overlay without IP tunneling. Our model leverages a header rewriting method as well as a host-based VLAN ID usage to ensure address space isolation and scalability of the number of virtual networks. In our model, any special hardware equipments like Open Flow hardware switch are not required and only software-based virtual switches and the controller are used. In this paper, we evaluate the performance of the proposed model comparing with the tunneling model using GRE or VXLAN protocol. Our model showed better performance and less CPU usage. In addition, qualitative evaluations of the model are also conducted from a broader perspective. Ryota Kawashima, Hiroshi Matsuo |
CloudCom (2) | 1 |
| 2008 | Design and Implementation of Multi-Platform Infrastructure of Extensible Network FunctionsabstractDynamic and flexible composition of higher-level network services, such as security, QoS, or adaptive services are required by future network applications. However, the development of such extensible applications makes them rather complex. In addition, many old applications, which do not support such services, would stick to be used. To solve these problems, we propose a generic and multi-platform infrastructure called FreeNA1 that extends existing applications by transparently incorporating the services to them. FreeNA offers abstract interfaces such that users can insert the services into each packet flow based on a configuration file. In this paper, we describe the design and implementation of FreeNA including a functionality comparison with relevant systems, and our performance evaluation results. The result shows that FreeNA offers finer configurability, composability, and usability and can be used widely than other similar systems. We also show that overhead of transparent service insertion is about 1-2% at a maximum compared to a method of inserting such services into applications directly. Ryota Kawashima, Yusheng Ji, Katsumi Maruyama |
GLOBECOM | 1 |
| 2008 | Design and implementation of real-time acoustic steganographyabstractRecently, steganography using multiple media contents has been proposed as an information security technology. In this paper, model and algorithm of real-time steganography scheme is proposed. This system is implemented as embedding a piece of secret audio data stream which is recorded as synchronous as it is embedded into another piece of audio cover data stream. In addition, embedding positions in cover data with sampling size of 16-bit can be arbitrarily-designated. SNR analysis function is presented to analyze the embedding robustness and capacity of first signal ldquo1rdquo in every sampling point of cover data. We derive a conclusion through empirical test that the stego data can avoid drawing suspicion even when we embed two bits of secret data into the 7thand 8thbit of cover data from LSB in each sampling point. Xuping Huang, Ryota Kawashima, Norihisa Segawa, Yoshihiko Abe |
ICME | 2 |