EDBT 2026 Demo / reviewers in the wild / expert
Shinpei Kato
dblp:87/7044
· DBLP profile ↗
69ranked-venue papers
18as first author
12since 2021 · last 2026
0000-0003-1782-5319ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 32 · 7 first-author · 4 since 2021Artificial intelligence and machine learning · 10 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 3 first-author · 1 since 2021Software engineering, systems software and programming languages · 2Security and privacy · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ipc_shared_ptr: A Publish/Subscribe-Aware Smart Pointer for Cross-Process Object Lifetime Management
Takahiro Ishikawa-Aso, Atsushi Yano, Koichi Imai, Takuya Azumi, Shinpei Kato |
ISORC | 5 |
| 2025 | ROS 2 Agnocast: Supporting Unsized Message Types for True Zero-Copy Publish/Subscribe IPC
Takahiro Ishikawa-Aso, Shinpei Kato |
ISORC | 2 |
| 2025 | Work in Progress: Middleware-Transparent Callback Enforcement in Commoditized Component-Oriented Real-Time SystemsabstractReal-time scheduling in commoditized componentoriented real-time systems, such as ROS 2 systems on Linux, has been studied under nested scheduling: OS thread scheduling and middleware layer scheduling (e.g., ROS 2 Executor). However, by establishing a persistent one-to-one correspondence between callbacks and OS threads, we can ignore the middleware layer and directly apply OS scheduling parameters (e.g., scheduling policy, priority, and affinity) to individual callbacks. We propose a middleware model that enables this idea and implements CallbackIsolatedExecutor as a novel ROS 2 Executor. We demonstrate that the costs (user-kernel switches, context switches, and memory usage) of CallbackIsolatedExecutor remain lower than those of the MultiThreadedExecutor, regardless of the number of callbacks. Additionally, the cost of CallbackIsolatedExecutor relative to SingleThreadedExecutor stays within a fixed ratio (1.4x for inter-process and 5x for intra-process communication). Future ROS 2 real-time scheduling research can avoid nested scheduling, ignoring the existence of the middleware layer. Takahiro Ishikawa-Aso, Atsushi Yano, Takuya Azumi, Shinpei Kato |
RTAS | 4 |
| 2025 | Work-in-Progress: Function-as-Subtask API Replacing Publish/Subscribe for OS-Native DAG SchedulingabstractThe Directed Acyclic Graph (DAG) task model for real-time scheduling finds its primary practical target in Robot Operating System 2 (ROS 2). However, ROS 2's publish/subscribe API leaves DAG precedence constraints unenforced: a callback may publish mid-execution, and multi-input callbacks let developers choose topic-matching policies. Thus preserving DAG semantics relies on conventions; once violated, the model collapses. We propose the Function-as-Subtask (FasS) API, which expresses each subtask as a function whose arguments/return values are the subtask's incoming/outgoing edges. By minimizing description freedom, DAG semantics is guaranteed at the API rather than by programmer discipline. We implement a DAGnative scheduler using FasS on a Rust-based experimental kernel and evaluate its semantic fidelity, and we outline design guidelines for applying FasS to Linux sched_ext. Takahiro Ishikawa-Aso, Atsushi Yano, Yutaro Kobayashi, Takumi Jin, Yuuki Takano, Shinpei Kato |
RTSS | 6 |
| 2023 | Moment-Based Kalman Filter: Nonlinear Kalman Filtering with Exact Moment PropagationabstractThis paper develops a new nonlinear filter, called Moment-based Kalman Filter (MKF), using the exact moment propagation method. Existing state estimation methods use linearization techniques or sampling points to compute approximate values of moments. However, moment propagation of probability distributions of random variables through nonlinear process and measurement models play a key role in the development of state estimation and directly affects their performance. The proposed moment propagation procedure can compute exact moments for non-Gaussian as well as non-independent Gaussian random variables. Thus, MKF can propagate exact moments of uncertain state variables up to any desired order. MKF is derivative-free and does not require tuning parameters. Moreover, MKF has the same computation time complexity as the extended or unscented Kalman filters, i.e., EKF and UKF. The experimental evaluations show that MKF is the preferred filter in comparison to EKF and UKF and outperforms both filters in non-Gaussian noise regimes. Yutaka Shimizu, Ashkan Jasour, Maani Ghaffari Jadidi, Shinpei Kato |
ICRA | 4 |
| 2022 | CARET: Chain-Aware ROS 2 Evaluation ToolabstractThis paper presents a tool to evaluate the latency of Robot Operating System (ROS) 2 applications called Chain-Aware ROS 2 Evaluation Tool (CARET). ROS 2 is designed to enhance the modularity of real-time robotic applications, including self-driving software such as Autoware. To analyze the performance of ROS 2 applications, CARET supports measurement functionalities for the callback latency, node latency, communication time between nodes, and end-to-end latency. To calculate each latency, CARET provides tracepoints, an architecture file, information on tracepoint connections, and a message tracking functionality. Furthermore, CARET can visualize different types of latency, such as bottlenecks and lost message, to analyze ROS 2 applications. The experimental results demonstrate that CARET can successfully measure the end-to-end latency of Autoware.Universe, which is ROS 2-based self-driving software. Takahisa Kuboichi, Atsushi Hasegawa, Keita Miura, Kenji Funaoka, Shinpei Kato, Takuya Azumi |
EUC | 6 |
| 2022 | Jerk Constrained Velocity Planning for an Autonomous Vehicle: Linear Programming ApproachabstractVelocity Planning for self-driving vehicles in a complex environment is one of the most challenging tasks. It must satisfy the following three requirements: safety with regards to collisions; respect of the maximum velocity limits defined by the traffic rules; comfort of the passengers. In order to achieve these goals, the jerk and dynamic objects should be considered, however, it makes the problem as complex as a non-convex optimization problem. In this paper, we propose a linear programming (LP) based velocity planning method with jerk limit and obstacle avoidance constraints for an autonomous driving system. To confirm the efficiency of the proposed method, a comparison is made with several optimization-based approaches, and we show that our method can generate a velocity profile which satisfies the aforementioned requirements more efficiently than the compared methods. In addition, we tested our algorithm on a real vehicle at a test field to validate the effectiveness of the proposed method. Yutaka Shimizu, Takamasa Horibe, Fumiya Watanabe, Shinpei Kato |
ICRA | 4 |
| 2022 | VMVG-Loc: Visual Localization for Autonomous Driving using Vector Map and Voxel Grid MapabstractThis study proposes a visual localization method using a vector map and voxel grid map with a stereo camera. The two maps provide different modality advantages and are integrated using a particle filter. In contrast to other vector map-based methods, our method does not use road markings because creating and maintaining vector maps that include high-accuracy road markings is laborious. Furthermore, it limits the regions where they are available. This method uses only lane center-lines from vector maps, which are easier to create than road markings. The method performs ray casting and computes the reprojection error to evaluate the vehicle position for voxel grid maps. Although this makes the method environmentally sensitive, the constraints by lanes make the estimation stable. Experiments confirmed that the method could perform localization stably and accurately without failure even over long distances. In addition, an ablation study showed the benefits of combining both maps. Kento Yabuuchi, Shinpei Kato |
IROS | 2 |
| 2022 | LaneFusion: 3D Object Detection with Rasterized Lane Mapabstract3D object detection is the task of locating and classifying objects in a 3D space. The task of 3d object detection is accompanied by the problem that objects are often detected while they are reversing, introducing a directional error of 180 degrees. This has a negative impact on following tasks of tracking and motion forecasting. One approach to solving this problem is by using a high definition (HD) map based on the information that cars drive along lanes; however, in its current form, this method does not fully utilize the given lane information and remains incapable of solving the problem of objects reversing. We propose a 3D object detection framework (”LaneFusion”) employing LiDAR and HD map fusion, using a vector map. LaneFusion overcomes the problem that the vector map format is difficult to input into current mainstream convolutional neural networks (CNNs), through a two-step rasterization process that incorporates vector map features into existing LiDAR-based detection methods. Our experiments confirmed that the proposed method increased the 3D average precision (AP) and average orientation similarity (AOS) of the vehicle class by up to 6.56 and 10.65 points, respectively. In addition, we analyzed the performance degradation caused by map input errors due to self-localization estimation and deviations from real road conditions. The proposed method was found to be more sensitive to orientation errors than to translation errors in self-localization, yet robust to the unavailability of map information by dropout during training. Taisei Fujimoto, Shinpei Kato |
IV | 3 |
| 2022 | A Conditional Confidence Calibration Method for 3D Point Cloud Object DetectionabstractWhen we apply neural networks to safety-critical systems such as self-driving cars, the reliability of their predictions must be considered. However, recent deep neural networks have tended to output biased confidence. Additionally, the extent of confidence bias estimated by object detectors varies depending on factors such as the detected object’s position and size. To address this problem, many researchers have proposed methods for calibrating confidences estimated by object detectors. In this study, we investigate the factors that may cause bias in the confidence of LiDAR-based 3D object detectors and show that our calibration method compensates for the effect of these factors to provide reliable confidence estimations, regardless of the neural network model used or the situations in which objects are detected. Yoshio Kato, Shinpei Kato |
IV | 2 |
| 2021 | Visual Localization for Autonomous Driving using Pre-built Point Cloud MapsabstractThis paper presents a vision-based metric localization method using pre-built point cloud maps. Matching the 3D structures reconstructed by visual SLAM to the point cloud map resolves the accumulative errors and scale ambiguity. In addition to the accuracy improvement, the proposed method achieves localization within given maps while ordinary visual SLAM constructs an on-line map and can only localize within this. Localization within a given map is crucial for autonomous driving, where various map types are employed. Point cloud maps are robust to appearance changes caused by illumination and seasonal changes. Once LiDAR sensors have built the point cloud maps, this paper demonstrates that localization is possible using solely low-cost and lightweight cameras. We verified the accuracy of the proposed method using real-world datasets. The results show that the accumulated error is suppressed even on extended vehicle trajectories. Also, we conducted experiments with various camera configurations and confirmed that the point cloud map improved the localization results for all configurations, Kento Yabuuchi, David Robert Wong, Takeshi Ishita, Yuki Kitsukawa, Shinpei Kato |
IV | 5 |
| 2021 | RAPLET: Demystifying Publish/Subscribe Latency for ROS ApplicationsabstractThe problem of real-time scheduling based on the Directed Acyclic Graph (DAG) task model has been extensively studied in the literature. Most of the studies are aimed at the development of efficient scheduling algorithms to reduce deadline misses and/or improve schedulability bounds on the given task system. In order to guarantee real-time performance of the DAG task model in practice, the latency imposed on the communication between the DAG nodes must be systematically taken into account. This paper aims at demystifying the latency of the Robot Operating System (ROS) as a practical DAG task model, which leverages the publish/subscribe mechanism to send and receive data between the nodes. To this end we present the ROS-Aware Publish/Subscribe Latency Evaluation Tool (RAPLET) which is designed to measure and visualize the details of the publish/subscribe latency in ROS. RAPLET consists of (i) the LD PRELOAD scheme that inserts function hooks in user-land and (ii) the extended Berkeley Packet Filter (eBPF) scheme that monitors the run-queue level and the network states in kernel-land. The performance analysis on ROS applications, including a real-world autonomous driving software, is performed using RAPLET to demonstrate that the publish/subscribe latency imposed on inter-node communication can be demystified and reasoned with respect to system issues including the message size and network bandwidth consumption. Keisuke Nishimura, Takahiro Ishikawa, Hiroshi Sasaki 0001, Shinpei Kato |
RTCSA | 4 |
| 2020 | ROS-lite: ROS Framework for NoC-Based Embedded Many-Core PlatformabstractThis paper proposes ROS-lite, a robot operating system (ROS) development framework for embedded many- core platforms based on network-on-chip (NoC) technology. Many-core platforms support the high processing capacity and low power consumption requirement of embedded systems. In this study, a self-driving software platform module is parallelized to run on many-core processors to demonstrate the practicality of embedded many-core platforms. The experimental results show that the proposed framework and the parallelized applications have met the deadline for low-speed self-driving systems. Takuya Azumi, Yuya Maruyama, Shinpei Kato |
IROS | 3 |
| 2020 | LIBRE: The Multiple 3D LiDAR DatasetabstractIn this work, we present LIBRE: LiDAR Benchmarking and Reference, a first-of-its-kind dataset featuring 10 different LiDAR sensors, covering a range of manufacturers, models, and laser configurations. Data captured independently from each sensor includes three different environments and configurations: static targets, where objects were placed at known distances and measured from a fixed position within a controlled environment; adverse weather, where static obstacles were measured from a moving vehicle, captured in a weather chamber where LiDARs were exposed to different conditions (fog, rain, strong light); and finally, dynamic traffic, where dynamic objects were captured from a vehicle driven on public urban roads, multiple times at different times of the day, and including supporting sensors such as cameras, infrared imaging, and odometry devices. LIBRE will contribute to the research community to (1) provide a means for a fair comparison of currently available LiDARs, and (2) facilitate the improvement of existing self-driving vehicles and robotics-related software, in terms of development and tuning of LiDAR-based perception algorithms. Alexander Carballo, Jacob Lambert 0001, Abraham Monrroy Cano, David Robert Wong, Patiphon Narksri, Yuki Kitsukawa, Eijiro Takeuchi, Shinpei Kato, Kazuya Takeda |
IV | 8 |
| 2019 | Resource Manager for Scalable Performance in ROS Distributed EnvironmentsabstractThis paper presents a resource manager to achieve scalable performance in Robot Operating System (ROS) for distributed environments. In robotics, using ROS in distributed environments via multiple host machines is trending for large-scale data processing, for example, cloud/edge computing and the data communication of point clouds and images in dynamic map composition. However, ROS is unable to manage the resources (e.g., the CPUs, memory, and disks) on each host machine. Therefore, it is difficult to use distributed environmental resources efficiently and achieve scalable performance. This paper proposes a resource management mechanism for ROS distributed environments using a master-slave model to execute ROS processes efficiently and smoothly. We manage the resource usage of each host machine and construct a mechanism to adaptively distribute the load to be balanced. Evaluations show that scalable performance can be achieved in ROS distributed environments comprising ten host machines using a real application (SLAM: simultaneous localization and mapping) processing large-scale point cloud data. Daisuke Fukutomi, Takuya Azumi, Shinpei Kato, Nobuhiko Nishio |
DATE | 3 |
| 2018 | GPUhd: Augmenting YARN with GPU Resource ManagementabstractThis paper presents GPUhd, a graphics processing unit (GPU) resource management approach that combines Hadoop and a GPU to obtain scale-out and scale-up functionality. There are several researches that combine Hadoop and GPU. However, there are no researches that can schedule tasks in consideration of GPU resource on Hadoop. Moreover, these researches cannot use multiple distributed frameworks. GPUhd extends the Yet Another Resource Negotiator (YARN) management mechanism and distributed processing frameworks for the coordinated use of GPU resources in Hadoop. We extend the YARN scheduling algorithm to consider GPU resources and incorporate a resources monitoring function. GPU resources can be managed on the basis of existing development methods because GPUhd simply handles GPU resources as host memory and CPU resources. In addition, GPUhd achieves high-speed processing, e.g., the computational time required to calculate 2048 x 2048 matrix multiplication is approximately 25 times less than that required when using only a CPU with Hadoop. GPUhd achieves high scalability and excellent response times in a heterogeneous distributed environment. Daisuke Fukutomi, Yuki Iida, Takuya Azumi, Shinpei Kato, Nobuhiko Nishio |
HPC Asia | 4 |
| 2018 | Scalable and Memory-Efficient Spin Locks for Embedded Tile-Based Many-Core ArchitecturesabstractEmbedded many-core System-on-Chip (SoC) architectures require scalability and memory constraints. However, communication between many cores, especially locking mechanisms of operating systems, is often the main obstacle to scalable and memory-efficient processing. Existing scalable spin locks consume non-negligible amounts of memory in many-core architectures, thus they are not suitable for memory constrained systems. This paper focuses on a combination of a global Mellor-Crummey and Scott (MCS) queue lock, and local ticket (TKT) locks. We refer to this lock as the C-MCS-TKT lock, which has much better memory efficiency than other scalable spin locks without degrading scalability. In addition, this paper also presents a memory-optimized version of the C-MCS-TKT lock, which slightly degrades scalability but reduces memory fragmentation, compared to the original C-MCS-TKT lock. Experimental results show that these locks have comparable performance to those of other highly scalable spin locks. Shinichi Awamoto, Hiroyuki Chishiro, Shinpei Kato |
ISORC | 3 |
| 2018 | Real-Time ROS Extension on Transparent CPU/GPU Coordination MechanismabstractRobot Operating System (ROS) promotes fault isolation, faster development, modularity, and core reusability and is therefore widely studied and used as the de facto standard for autonomous driving systems. Graphics processing units (GPUs) also facilitate high-performance computing and are therefore used for autonomous driving. As the requirements for real-time processing increase, methods for satisfying real-time constraints for ROS and GPUs are being developed. Unfortunately, scheduling algorithms specifying ROS's transportation (publish/subscribe) model, which can have execution order restrictions, are not being investigated, leading to the introduction of waiting time and degrading the responsiveness of the entire system. Furthermore, GPU tasks on ROS are also affected by the ROS transportation model, because central processing unit (CPU) time is occupied when GPU functions are launched. This paper proposes a loadable kernel module framework, called real-time ROS extension on transparent CPU/GPU coordination mechanism (ROSCH-G), for scheduling ROS in a heterogeneous environment without modifying the OS kernel and device drivers and then evaluates it experimentally. ROSCH-G provides a scheduling algorithm that considers ROS's execution order restrictions and a CPU/GPU coordination mechanism. Experimental results demonstrate that the proposed algorithm reduces the deadline miss rate and, compared with previous studies, makes effective use of the benefits of parallel processing. In addition, the results for the coordination mechanism demonstrate that ROSCH-G can schedule multiple GPU applications successfully. Yuhei Suzuki, Takuya Azumi, Shinpei Kato, Nobuhiko Nishio |
ISORC | 3 |
| 2018 | ROSCH: Real-Time Scheduling Framework for ROSabstractThis paper presents a real-time scheduling framework for the robot operating system (ROS) called ROSCH. ROS is an open-source software platform and a meta-operating system designed for robots. ROS provides extensive development libraries and has been widely used for autonomous driving systems. However, ROS does not guarantee real-time performance; hence, a ROS-based autonomous driving car could cause a traffic accident. Therefore, ROSCH comprises three functionalities that do not exist in the ROS to guarantee real-time performance: (1) a synchronization system, (2) fixed-priority based directed acyclic graph (DAG) scheduling framework, and (3) fail-safe function. In particular, the synchronization system guarantees that the timestamp gaps between sensor measurements will be less than or equal to the calculated value. The fixed-priority based DAG scheduling framework guarantees that end-to-end latency is less than or equal to an estimated value. Operating both mechanisms simultaneously guarantees the final output topic frequency. The fail-safe functionality provides a danger avoidance action at a deadline miss. In addition, these functionalities are designed to be compatible with legacy software. No previous method that guarantees real-time operation provides a comparable range of features while maintaining the compatibility with the existing software. Experimental results showed that the time differences between sensor timestamps remained below a calculated value. Moreover, correctly scheduled tasks incur an end-to-end latency that is within an estimated value. In addition, the throughput of the end node is achieved. Finally, the fail-safe functionality overhead is 45 μsec on an average, which is 0.004% of the average execution time for one node. Yukihiro Saito, Futoshi Sato, Takuya Azumi, Shinpei Kato, Nobuhiko Nishio |
RTCSA | 4 |
| 2018 | Deep Learning on Large-Scale Muticore ClustersabstractConvolutional neural networks (CNNs) have achieved outstanding accuracy among conventional machine learning algorithms. Recent works have shown that large and complicated models, which take significant cost for training are needed to get higher accuracy. To train these models efficiently in high performance computers (HPCs), many parallelization techniques for CNNs have been developed. However, most techniques are mainly targeting GPUs and parallelizations for CPUs are not fully investigated. This paper explores CNN training performance on large-scale multicore clusters by optimizing intra-node processing and applying techniques of inter-node parallelization for multiple GPUs. Detailed experiments conducted on state-of-the-art multi-core processors using the openMP API and MPI framework demonstrated that Caffe-based CNNs can be accelerated by using well-designed multithreaded programs. We achieved at most 1.64 times speedup in convolution operations with devised lowering strategy compared to conventional lowering and acquired 772 times speedup with 864 nodes compared to one node. Kazumasa Sakivama, Shinpei Kato, Yutaka Ishikawa, Atsushi Hori, Abraham Monrroy Cano |
SBAC-PAD | 2 |
| 2017 | GLoop: an event-driven runtime for consolidating GPGPU applicationsabstractGraphics processing units (GPUs) have become an attractive platform for general-purpose computing (GPGPU) in various domains. Making GPUs a time-multiplexing resource is a key to consolidating GPGPU applications (apps) in multi-tenant cloud platforms. However, advanced GPGPU apps pose a new challenge for consolidation. Such highly functional GPGPU apps, referred to as GPU eaters, can easily monopolize a shared GPU and starve collocated GPGPU apps. This paper presents GLoop, which is a software runtime that enables us to consolidate GPGPU apps including GPU eaters. GLoop offers an event-driven programming model, which allows GLoop-based apps to inherit the GPU eaters' high functionality while proportionally scheduling them on a shared GPU in an isolated manner. We implemented a prototype of GLoop and ported eight GPU eaters on it. The experimental results demonstrate that our prototype successfully schedules the consolidated GPGPU apps on the basis of its scheduling policy and isolates resources among them. Yusuke Suzuki, Shinpei Kato, Kenji Kono |
SoCC | 3 |
| 2017 | Exploring Scalable Data Allocation and Parallel Computing on NoC-Based Embedded Many CoresabstractIn embedded systems, high processing requirements and low power consumption need heterogeneous computing platforms. Considering embedded requirements, applications need to be designed based on scalable data allocation and parallel computing with non-uniform memory access (NUMA) many cores. In this paper, we use one of the embedded commercial off-the-shelf (COTS) multi/many-core components, the Massively Parallel Processor Arrays (MPPA) 256 developed by Kalray, and conduct evaluations of data transfer and parallelization of a practical application. We investigate currently achievable data transfer latencies between distributed memories on network-on-chip (NoC), memory access characteristics, and parallelization potential with many cores. Subsequently, we run a practical application, the core of the autonomous driving system, on many-core processors and acceleration by parallelization indicates practicality of many cores. By highlighting many-core computing capabilities, we explore the scalable data allocation and parallel computing on NoC-based embedded many cores. Yuya Maruyama, Shinpei Kato, Takuya Azumi |
ICCD | 2 |
| 2017 | Real-Time GPU Resource Management with Loadable Kernel ModulesabstractGraphics processing unit (GPU) programming environments have matured for general-purpose computing on GPUs. Significant challenges for GPUs include system software support for bounded response times and guaranteed throughput. In recent years, GPU technologies have been applied to real-time systems by extending the operating system modules to support real-time GPU resource management. Unfortunately, such a system extension makes it difficult to maintain the system with version updates because the OS kernel and device drivers must be modified at the source-code level, thereby preventing continuous research and development of GPU technologies for real-time systems. A loadable kernel module (LKM) framework, called Linux Real-Time eXtention with GPUs (Linux-RTXG), for managing real-time GPU resources with Linux without modifying the OS kernel and device drivers is proposed and evaluated experimentally. Linux-RTXG provides mechanisms for interrupt interception and independent synchronization to achieve real-time scheduling and resource reservation capabilities for GPU applications on top of existing device drivers and runtime libraries. Experimental results demonstrate that the overhead incurred by introducing the proposed Linux-RTXG is comparable to that of introducing existing kernel-dependent approaches. In addition, the results demonstrate that multiple GPU applications can be scheduled successfully by Linux-RTXG to meet their priority and quality-of-service requirements in real time. Yuhei Suzuki, Yusuke Fujii, Takuya Azumi, Nobuhiko Nishio, Shinpei Kato |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2017 | Relational Joins on GPUs: A Closer LookabstractThe problem of scaling out relational join performance for large data sets in the database management system (DBMS) has been studied for years. Although in-memory DBMS engines can reduce load times by storing data in the main memory, join queries still remain computationally expensive. Modern graphics processing units (GPUs) provide massively parallel computing and may enhance the performance of such join queries; however, it is not clearyet in what condition relational joins perform well on GPUs. In this paper, we identify the performance characteristics of GPU computing for relational joins by implementing several well-known GPU-based join algorithms under various configurations. Experimental results indicate that the speedup ratio of GPU-based relational joins to CPU-based counterparts depends on the number of compute cores, the size of data sets, join conditions, and join algorithms. In the best case, the speedup ratios are up to 6.67 times for non-index joins, 9.41 times for sort index joins, and 2.55 times for hash joins. The execution time of GPU-based implementation for index joins, on the other hand, is only about 0.696 times less than the execution time of the CPU's counterparts. Makoto Yabuta, Shinpei Kato, Masato Edahiro, Hideyuki Kawashima |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2016 | Exploring the performance of ROS2abstractMiddleware for robotics development must meet demanding requirements in real-time distributed embedded systems. The Robot Operating System (ROS), open-source middleware, has been widely used for robotics applications. However, the ROS is not suitable for real-time embedded systems because it does not satisfy real-time requirements and only runs on a few OSs. To address this problem, ROS1 will undergo a significant upgrade to ROS2 by utilizing the Data Distribution Service (DDS). DDS is suitable for real-time distributed embedded systems due to its various transport configurations (e.g., deadline and fault-tolerance) and scalability. ROS2 must convert data for DDS and abstract DDS from its users; however, this incurs additional overhead, which is examined in this study. Transport latencies between ROS2 nodes vary depending on the use cases, data size, configurations, and DDS vendors. We conduct proof of concept for DDS approach to ROS and arrange DDS characteristic and guidelines from various evaluations. By highlighting the DDS capabilities, we explore and evaluate the potential and constraints of DDS and ROS2. Yuya Maruyama, Shinpei Kato, Takuya Azumi |
EMSOFT | 2 |
| 2016 | Precise and efficient model-based vehicle tracking method using Rao-Blackwellized and scaling series particle filtersabstractA precise and efficient tracking technique is essential for an intelligent vehicle to interact with the surrounding vehicles on the road. In this paper, we propose a model-based vehicle tracking method using the Rao-Blackwellized particle filter (RBPF) and scaling series particle filter (SSPF). Mengwen He, Eijiro Takeuchi, Yoshiki Ninomiya, Shinpei Kato |
IROS | 4 |
| 2016 | Robust virtual scan for obstacle Detection in urban environmentsabstractObstacle detection is an essential technique for intelligent vehicles. Environmental sensing especially plays a vital role to achieve accurate obstacle detection. Unlike classical 2D scan, emerging 3D Light Detection and Ranging (LiDAR) sensors can scan dense point cloud at one time, which represents detailed information of urban environments. The downside of obstacle detection using 3D LiDAR, on the other hand, is its computational cost posed by a large amount of 3D data. The virtual scan (VScan), first introduced by Petrovskaya et al. [1] for efficient vehicle detection and tracking, is a 2D compression of 3D point cloud to represent free space, obstacles and unknown areas. To overcome the computational problem of obstacle detection using 3D LiDAR, therefore, VScan is suitable. In addition, it can bridge across new-born 3D LiDAR sensors and many matured applications based on 2D scan, including occupancy grid map, SLAM, planning, detection, and tracking, due to its 2D representation of 3D point cloud. Mengwen He, Eijiro Takeuchi, Yoshiki Ninomiya, Shinpei Kato |
Intelligent Vehicles Symposium | 4 |
| 2016 | GPUvm: GPU Virtualization at the HypervisorabstractGraphic processing units (GPUs) provide a massively-parallel computational power and encourage the use of general-purpose computing on GPUs (GPGPU). The distinguished design ofdiscrete GPUshelps them to provide the high throughput, scalability, and energy efficiency needed for GPGPU applications. Despite the previous study on GPU virtualization, the tradeoffs between the virtualization approaches remain unclear, because of a lack of designs for or quantitative evaluations of the hypervisor-level virtualization for discrete GPUs. Shedding light on these tradeoffs and the technical requirements for the hypervisor-level virtualization would facilitate the development of an appropriate GPU virtualization solution.$\sf{GPUvm}$, which is an open architecture for hypervisor-level GPU virtualization with a particular emphasis on using the Xen hypervisor, is presented in this paper.$\sf{GPUvm}$offers three virtualization modes: the full-, naive para-, and high-performance para-virtualization.$\sf{GPUvm}$exposes low- and high-level interfaces such as memory-mapped I/O and DRM APIs to the guest virtual machines (VMs). Our experiments using a relevant commodity GPU showed that$\sf{GPUvm}$incurs different overheads as the level of the exposed interfaces is changed. The results also showed that a coarse-grained fairness on the GPU among multiple VMs can be achieved using GPU scheduling. Yusuke Suzuki, Shinpei Kato, Kenji Kono |
IEEE Trans. Computers | 2 |
| 2016 | GPUrpc: Exploring Transparent Access to Remote GPUs
Yuki Iida, Yusuke Fujii, Takuya Azumi, Nobuhiko Nishio, Shinpei Kato |
ACM Trans. Embed. Comput. Syst. | 5 |
| 2016 | Accelerated Deformable Part Models on GPUsabstractObject detection is a fundamental challenge facing intelligent applications. Image processing is a promising approach to this end, but its computational cost is often a significant problem. This paper presents schemes for accelerating the deformable part models (DPM) on graphics processing units (GPUs). DPM is a well-known algorithm for image-based object detection, and it achieves high detection rates at the expense of computational cost. GPUs are massively parallel compute devices designed to accelerate data-parallel compute-intensive workload. According to an analysis of execution times, approximately 98 percent of DPM code exhibits loop processing, which means that DPM could be highly parallelized by GPUs. In this paper, we implement DPM on the GPU by exploiting multiple parallelization schemes. Results of an experimental evaluation of this GPU-accelerated DPM implementation demonstrate that the best scheme of GPU implementations using an NVIDIA GPU achieves a speed up of 8.6x over a naive CPU-based implementation. Manato Hirabayashi, Shinpei Kato, Masato Edahiro, Kazuya Takeda, Seiichi Mita |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2014 | A server-based approach for overrun management in multi-core real-time systemsabstractThis paper presents a server-based framework for task overrun management in multi-core real-time systems. Unlike most existing scheduling methods which usually assume a single upper bound of the Worst-Case Execution Time (WCET) for each task, our approach targets scenarios with task overruns. The main idea of our framework is to employ Synchronized Deferrable Servers (SDS) to deal with globally scheduled task overruns, while a partitioned scheduling approach is applied on regular task executions. Moreover, we provide a deterministic Worst-Case Response Time (WCRT) analysis focusing on hard timing constraints, along with a probabilistic analysis of Deadline Miss Ratio (DMR) for soft real-time applications. In the evaluation phase, we have implemented two types of experiments evaluating different timing constraints. Meng Liu 0001, Moris Behnam, Shinpei Kato, Thomas Nolte |
ETFA | 3 |
| 2014 | Power and Performance Characterization and Modeling of GPU-Accelerated SystemsabstractGraphics processing units (GPUs) provide an order-of-magnitude improvement on peak performance and performance-per-watt as compared to traditional multicore CPUs. However, GPU-accelerated systems currently lack a generalized method of power and performance prediction, which prevents system designers from an ultimate goal of dynamic power and performance optimization. This is due to the fact that their power and performance characteristics are not well captured across architectures, and as a result, existing power and performance modeling approaches are only available for a limited range of particular GPUs. In this paper, we present power and performance characterization and modeling of GPU-accelerated systems across multiple generations of architectures. Characterization and modeling both play a vital role in optimization and prediction of GPU-accelerated systems. We quantify the impact of voltage and frequency scaling on each architecture with a particularly intriguing result that a cutting-edge Kepler-based GPU achieves energy saving of 75% by lowering GPU clocks in the best scenario, while Fermi- and Tesla-based GPUs achieve no greater than 40% and 13%, respectively. Considering these characteristics, we provide statistical power and performance modeling of GPU-accelerated systems simplified enough to be applicable for multiple generations of architectures. One of our findings is that even simplified statistical models are able to predict power and performance of cutting-edge GPUs within errors of 20% to 30% for any set of voltage and frequency pair. Yuki Abe 0001, Hiroshi Sasaki 0001, Shinpei Kato, Koji Inoue, Masato Edahiro, Martin Peres |
IPDPS | 3 |
| 2014 | An adaptive server-based scheduling framework with capacity reclaiming and borrowingabstractIn this paper, we present a new reservation based scheduling framework for soft real-time systems using EDF algorithm (called CARB-EDF). This framework has the features of Capacity Adaptation, Reclaiming and Borrowing. This framework can simplify the initial configuration of the system, where the system designer does not need to provide any estimations of task execution times. We also present a Chebyshev's inequality based predictor to estimate task execution times. A number of simulation-based experiments have been implemented. According to the results compared with some related works, our scheduling framework can provide a better performance with acceptable extra scheduling overhead. Meng Liu 0001, Moris Behnam, Shinpei Kato, Thomas Nolte |
RTCSA | 3 |
| 2014 | GDM: device memory management for gpgpu computingabstractGPGPUs are evolving from dedicated accelerators towards mainstream commodity computing resources. During the transition, the lack of system management of device memory space on GPGPUs has become a major hurdle. In existing GPGPU systems, device memory space is still managed explicitly by individual applications, which not only increases the burden of programmers but can also cause application crashes, hangs, or low performance. Kaibo Wang, Xiaoning Ding, Rubao Lee, Shinpei Kato, Xiaodong Zhang 0001 |
SIGMETRICS | 4 |
| 2014 | GPUvm: Why Not Virtualizing GPUs at the Hypervisor?
Yusuke Suzuki, Shinpei Kato, Kenji Kono |
USENIX ATC | 2 |
| 2013 | Data Transfer Matters for GPU ComputingabstractGraphics processing units (GPUs) embrace many-core compute devices where massively parallel compute threads are offloaded from CPUs. This heterogeneous nature of GPU computing raises non-trivial data transfer problems especially against latency-critical real-time systems. However even the basic characteristics of data transfers associated with GPU computing are not well studied in the literature. In this paper, we investigate and characterize currently-achievable data transfer methods of cutting-edge GPU technology. We implement these methods using open-source software to compare their performance and latency for real-world systems. Our experimental results show that the hardware-assisted direct memory access (DMA) and the I/O read-and-write access methods are usually the most effective, while on-chip micro controllers inside the GPU are useful in terms of reducing the data transfer latency for concurrent multiple data streams. We also disclose that CPU priorities can protect the performance of GPU data transfers. Yusuke Fujii, Takuya Azumi, Nobuhiko Nishio, Shinpei Kato, Masato Edahiro |
ICPADS | 4 |
| 2013 | Multiprocessor Real-Time Scheduling with a Few Migrating TasksabstractWe present HIME, a new EDF-based semi-partitioned scheduling algorithm which allows at most one migrating task per processor. In a system with m processors, this arrangement limits the migrating tasks to at most m/2 and the number of migrations per job to at most m-1. HIME has a utilisation bound of at least 74.9%, and can be configured to achieve 75%, the theoretical limit for semi-partitioned schemes with at most m/2 migrating tasks. Experiments show that the average system utilisation achieved by HIME is about 95%. J. Augusto Santos Junior, George Lima 0001, Konstantinos Bletsas 0001, Shinpei Kato |
RTSS | 4 |
| 2012 | DS-Bench Toolset: Tools for dependability benchmarking with simulation and assuranceabstractToday's information systems have become large and complex because they must interact with each other via networks. This makes testing and assuring the dependability of systems much more difficult than ever before. DS-Bench Toolset has been developed to address this issue, and it includes D-Case Editor, DS-Bench, and D-Cloud. D-Case Editor is an assurance case editor. It makes a tool chain with DS-Bench and D-Cloud, and exploits the test results as evidences of the dependability of the system. DS-Bench manages dependability benchmarking tools and anomaly loads according to benchmarking scenarios. D-Cloud is a test environment for performing rapid system tests controlled by DS-Bench. It combines both a cluster of real machines for performance-accurate benchmarks and a cloud computing environment as a group of virtual machines for exhaustive function testing with a fault-injection facility. DS-Bench Toolset enables us to test systems satisfactorily and to explain the dependability of the systems to the stakeholders. Hajime Fujita 0002, Yutaka Matsuno, Toshihiro Hanawa, Mitsuhisa Sato, Shinpei Kato, Yutaka Ishikawa |
DSN | 5 |
| 2012 | QBox: guaranteeing I/O performance on black box storage systemsabstractMany storage systems are shared by multiple clients with different types of workloads and performance targets. To achieve performance targets without over-provisioning, a system must provide isolation between clients. Throughput-based reservations are challenging due to the mix of workloads and the stateful nature of disk drives, leading to low reservable throughput, while existing utilization-based solutions require specialized I/O scheduling for each device in the storage system. Dimitrios Skourtis, Shinpei Kato, Scott A. Brandt |
HPDC | 2 |
| 2012 | ExSched: An External CPU Scheduler Framework for Real-Time SystemsabstractScheduling theory and algorithms have been well studied in the real-time systems literature. Many useful approaches and solutions have appeared in different problem domains. While their theoretical effectiveness has been extensively discussed, the community is now facing implementation challenges that show the impact of the algorithms in practice. In this paper, we propose a scheduler framework, called ExSched, which enables different schedulers to be developed for different operating system (OS) platforms without any modifications to the OS itself, using a unified interface. The framework will easily keep up with changes in the kernel since it is only dependent on a few kernel primitives. The usefulness of this framework is that scheduling policies can be implemented as external plug-ins. They can simply use the ExSched interface instead of platform-dependent functions, since platform details are abstracted by ExSched. The advantage for industry is that they would more easily keep up with new kernel versions since ExSched does not require patches. The advantage for academia is that we could focus on the development of schedulers instead of tedious and time-consuming installations of patched kernels. Our prototype implementation of ExSched supports Linux and Vx Works and it comes with example schedulers which include hierarchical and multi-core schedulers in addition to traditional fixed-priority scheduling (FPS) and earliest deadline first (EDF) algorithms. Mikael Asberg, Thomas Nolte, Shinpei Kato, Ragunathan Rajkumar |
RTCSA | 3 |
| 2012 | Supporting Low-Latency CPS Using GPUs and Direct I/O SchemesabstractGraphics processing units (GPUs) are increasingly being used for general purpose parallel computing. They provide significant performance gains over multi-core CPU systems, and are an easily accessible alternative to supercomputers. The architecture of general purpose GPU systems(GPGPU), however, poses challenges in efficiently transferring data among the host and device(s). Although commodity many core devices such as NVIDIA GPUs provide more than one way to move data around, it is unclear which method is most effective given a particular application. This presents difficulty in supporting latency-sensitive cyber-physical systems (CPS). In this work we present a new approach to data transfer in a heterogeneous computing system that allows direct communication between GPUs and other I/O devices. In addition to adding this functionality our system also improves communication between the GPU and host. We analyze the current vendor provided data communication mechanisms and identify which methods work best for particular tasks with respect to throughput, and total time to completion. Our method allows a new class of real-time cyber-physical applications to be implemented on a GPGPU system. The results of the experiments presented here show that GPU tasks can be completed in 34 percent less time than current methods. Furthermore, effective data throughput is at least as good as the current best performers. This work is part of concurrent development of Gdev, an open-source project to provide Linux operating system support of many-core device resource management. Jason Aumiller, Scott A. Brandt, Shinpei Kato, Nikolaus Rath |
RTCSA | 3 |
| 2012 | Gdev: First-Class GPU Resource Management in the Operating System
Shinpei Kato, Michael McThrow, Carlos Maltzahn, Scott A. Brandt |
USENIX ATC | 1 |
| 2012 | FPSL, FPCL and FPZL schedulability analysis
Robert I. Davis 0001, Shinpei Kato |
Real Time Syst. | 2 |
| 2011 | Towards real-time scheduling of virtual machines without kernel modificationsabstractVirtualization is a well used technique in the area of internet server systems for managing several (legacy) applications on a single physical machine. These applications do not have strict time deadlines, which also reflects how these applications are scheduled. Using virtualization in an embedded real-time systems context is of course attractive, since we want to pack as much software as possible on a, as small as possible, hardware platform. The problem is that this kind of software does not easily cope well together, in the aspect of time related properties. Hence, we need a new mechanism, i.e., a scheduler, that can satisfy the timing requirements of each application. However, scheduler implementations typically require modifications to middleware or kernel and this is not acceptable in the area embedded systems, due to stability and reliability reasons. Hence, in this paper, we propose a framework for scheduling (soft real-time) applications residing in separate operating systems (virtual machines) using hierarchical fixed-priority preemptive scheduling, without the requirement of kernel modifications.1 Mikael Asberg, Nils Forsberg, Thomas Nolte, Shinpei Kato |
ETFA | 4 |
| 2011 | Resource Sharing in GPU-Accelerated Windowing SystemsabstractRecent windowing systems allow graphics applications to directly access the graphics processing unit (GPU) for fast rendering. However, application tasks that render frames on the GPU contend heavily with the windowing server that also accesses the GPU to blit the rendered frames to the screen. This resource-sharing nature of direct rendering introduces core challenges of priority inversion and temporal isolation in multi-tasking environments. In this paper, we identify and address resource-sharing problems raised in GPU-accelerated windowing systems. Specifically, we propose two protocols that enable application tasks to efficiently share the GPU resource in the X Window System. The Priority Inheritance with X server (PIX) protocol eliminates priority inversion caused in accessing the GPU, and the Reserve Inheritance with X server (RIX) protocol addresses the same problem for resource-reservation systems. Our design and implementation of these protocols highlight the fact that neither the X server nor user applications need modifications to use our solutions. Our evaluation demonstrates that multiple GPU-accelerated graphics applications running concurrently in the X Window System can be correctly prioritized and isolated by the PIX and the RIX protocols. Shinpei Kato, Karthik Lakshmanan, Yutaka Ishikawa, Ragunathan Rajkumar |
IEEE Real-Time and Embedded Technology and Applications Symposium | 1 |
| 2011 | A Loadable Task Execution Recorder for Hierarchical Scheduling in LinuxabstractThis paper presents a Hierarchical Scheduling Framework (HSF) recorder for Linux-based operating systems. The HSF recorder is a loadable kernel module that is capable of recording tasks and servers without requiring any kernel modifications. Hence, it complies with the reliability and stability requirements in the area of embedded systems where proven versions of Linux are preferred. The recorder is built upon the loadable real-time scheduler framework RESCH (Real-time Scheduler). We evaluate our recorder by comparing the overhead of this solution against another (patched) recorder. Also, the tracing accuracy of the HSF recorder is tested by running a media-processing task together with periodic real-time Linux tasks in combination with servers. The tests are recorded with the HSF recorder, and the Ftrace recorder, in order to show the correctness of the experiments and the HSF recorder itself. Mikael Asberg, Thomas Nolte, Shinpei Kato |
RTCSA (1) | 3 |
| 2011 | RGEM: A Responsive GPGPU Execution Model for Runtime EnginesabstractGeneral-purpose computing on graphics processing units, also known as GPGPU, is a burgeoning technique to enhance the computation of parallel programs. Applying this technique to real-time applications, however, requires additional support for timeliness of execution. In particular, the non-preemptive nature of GPGPU, associated with copying data to/from the device memory and launching code onto the device, needs to be managed in a timely manner. In this paper, we present a responsive GPGPU execution model (RGEM), which is a user-space runtime solution to protect the response times of high-priority GPGPU tasks from competing workload. RGEM splits a memory-copy transaction into multiple chunks so that preemption points appear at chunk boundaries. It also ensures that only the highest-priority GPGPU task launches code onto the device at any given time, to avoid performance interference caused by concurrent launches. A prototype implementation of an RGEM-based CUDA runtime engine is provided to evaluate the real-world impact of RGEM. Our experiments demonstrate that the response times of high-priority GPGPU tasks can be protected under RGEM, whereas their response times increase in an unbounded fashion without RGEM support, as the data sizes of competing workload increase. Shinpei Kato, Karthik Lakshmanan, Mihir Kelkar, Yutaka Ishikawa, Ragunathan Rajkumar |
RTSS | 1 |
| 2011 | TimeGraph: GPU Scheduling for Real-Time Multi-Tasking Environments
Shinpei Kato, Karthik Lakshmanan, Ragunathan Rajkumar, Yutaka Ishikawa |
USENIX ATC | 1 |
| 2011 | Global EDF-based scheduling with laxity-driven priority promotion
Shinpei Kato, Nobuyuki Yamasaki |
J. Syst. Archit. | 1 |
| 2011 | CPU scheduling and memory management for interactive real-time applications
Shinpei Kato, Yutaka Ishikawa, Ragunathan Rajkumar |
Real Time Syst. | 1 |
| 2010 | AIRS: Supporting Interactive Real-Time Applications on Multicore PlatformsabstractModern real-time systems increasingly operate with multiple interactive applications. While these systems often require reliable quality of service (QoS) for the applications, even under heavy workloads, many existing CPU schedulers are not very capable of satisfying such requirements. In this paper, we design and implement an Advanced Interactive and Real-time Scheduler, called AIRS. AIRS is aimed at supporting systems that run multiple interactive real-time applications, particularly on multicore platforms. It provides a new CPU reservation mechanism to enhance the QoS of the overall system. The reservation algorithm is based on the prior Constant Bandwidth Server (CBS) algorithm, but is more flexible and efficient, when multiple applications reserve CPU bandwidth. It also provides a new multicore scheduler to improve the absolute CPU bandwidth available for the applications to perform well. The scheduling algorithm is subject to the prior Earliest Deadline First with Window-constraint Migration (EDF-WM) algorithm, but is extended to work with the new CPU reservation mechanism. Experimental evaluation shows that AIRS delivers higher quality to simultaneous playback of multiple movies than the existing real-time scheduler. It also demonstrates that AIRS offers hard timing guarantees for randomly-generated task sets with heavy workloads. Shinpei Kato, Ragunathan Rajkumar, Yutaka Ishikawa |
ECRTS | 1 |
| 2010 | Towards hierarchical scheduling in Linux/multi-core platformabstractThis paper proposes the implementation of 4 different scheduling strategies for combining multi-core scheduling with hierarchical scheduling. Three of the scheduling schemes are analyzable with state-of-the-art schedulability analysis theory, available in the real-time systems community. Our idea is to implement these hierarchical multi-core scheduling strategies in a Linux based operating system, without modifying the kernel, and evaluate them. As of now, we have developed/implemented a prototype two-level hierarchical scheduling framework (HSF) in Linux (uni-core), which supports fixed priority preemptive scheduling (FPPS) of periodic servers at the top level, and FPPS of periodic tasks at the second level. The HSF is based on the REal-time SCHeduler (RESCH) framework. Mikael Asberg, Thomas Nolte, Shinpei Kato |
ETFA | 3 |
| 2010 | Scheduling Parallel Real-Time Tasks on Multi-core ProcessorsabstractMassively multi-core processors are rapidly gaining market share with major chip vendors offering an ever increasing number of cores per processor. From a programming perspective, the sequential programming model does not scale very well for such multi-core systems. Parallel programming models such as OpenMP present promising solutions for more effectively using multiple processor cores. In this paper, we study the problem of scheduling periodic real-time tasks on multiprocessors under the fork join structure used in OpenMP. We illustrate the theoretical best-case and worst-case periodic fork-join task sets from a processor utilization perspective. Based on our observations of these task sets, we provide a partitioned preemptive fixed-priority scheduling algorithm for periodic fork-join tasks. The proposed multiprocessor scheduling algorithm is shown to have a resource augmentation bound of 3.42, which implies that any task set that is feasible on m unit speed processors can be scheduled by the proposed algorithm on m processors that are 3:42 times faster. Karthik Lakshmanan, Shinpei Kato, Ragunathan Rajkumar |
RTSS | 2 |
| 2009 | Semi-partitioned Scheduling of Sporadic Task Systems on MultiprocessorsabstractThis paper presents a new algorithm for scheduling of sporadic task systems with arbitrary deadlines on identical multiprocessor platforms. The algorithm is based on the concept of semi-partitioned scheduling, in which most tasks are fixed to specific processors, while a few tasks migrate across processors. Particularly, we design the algorithm so that tasks are qualified to migrate only if a task set cannot be partitioned any more, and such migratory tasks migrate from one processor to another processor only once in each period. The scheduling policy is then subject to earliest deadline first. simulation results show that the algorithm delivers competitive scheduling performance to the state-of-the-art, with a smaller number of context switches. Shinpei Kato, Nobuyuki Yamasaki, Yutaka Ishikawa |
ECRTS | 1 |
| 2009 | Execution Time Monitoring in LinuxabstractThis paper presents an implementation of an Execution Time Monitor (ETM) which can be applied in a resource management framework, such as the one proposed in the Open Media Platform (OMP) [4]. OMP is a European project which aims at creating an open, flexible and resource efficient software architecture for mobile devices such as cell phones and handsets. One of its goals is to open up the possibility for software portability and fast integration of applications, in order to decrease development costs. The task of the ETM is to measure task execution time and provide this information to the scheduler which then can schedule tasks in a more efficient and dynamic way. This implementation is our first step towards a full resource management framework that later will include a hierarchical scheduler, for soft real-time systems. Mikael Asberg, Thomas Nolte, Clara Otero Pérez, Shinpei Kato |
ETFA | 4 |
| 2009 | Extended RT-Component Framework for RT-MiddlewareabstractModular component-based robot systems require not only an infrastructure for component management, but also scalability as well as real-time properties. Robot technology (RT)-middleware is a software platform for such component-based robot systems. Each component in the RT-Middleware, so-called "RT-component'' supporting particular robot functions, is based on common object request broker architecture (CORBA). Unfortunately, the RT-Middleware lacks the mechanism for real-time control. In this paper, we extend the framework of the RT-Components to take care of timing constraints. We first enable tasks to have different periods within each RT-Component. We then modify the packet format of the General Inter-ORB Protocol (GIOP) to transfer the information of timing constraints over RT-Components. The performance evaluation on ART-Linux shows that the extended RT-Component framework improves the schedulability of distributed real-time tasks, without causing critical overheads in unmarshaling the modified GIOP packets. Hiroyuki Chishiro, Yuji Fujita, Akira Takeda, Yuta Kojima, Kenji Funaoka, Shinpei Kato, Nobuyuki Yamasaki |
ISORC | 6 |
| 2009 | Semi-partitioned Fixed-Priority Scheduling on MultiprocessorsabstractThis paper presents a new algorithm for fixed-priority scheduling of sporadic task systems on multiprocessors.The algorithm is categorized to such a scheduling class that qualifies a few tasks to migrate across processors, while most tasks are fixed to particular processors. We design the algorithm so that a task is qualified to migrate, only if it cannot be assigned to any individual processors, in such a way that it is never returned to the same processor within the same period, once it is migrated from one processor to another processor. The scheduling policy is then conformed to deadline monotonic. According to the simulation results, the new algorithm significantly outperforms the traditional fixed-priority algorithms in terms of schedulability. Shinpei Kato, Nobuyuki Yamasaki |
IEEE Real-Time and Embedded Technology and Applications Symposium | 1 |
| 2009 | Periodic and Aperiodic Communication Techniques for Responsive LinkabstractResponsive Link, an ISO/IEC communication standard, provides many functional capabilities for distributed realtime systems. This paper is focused on periodic and aperiodic communication techniques for Responsive Link. In periodic communication, the priority is assigned to each packet so that the network utilization is improved. A schedulability test for connection establishments is also derived to ensure timing guarantees. In aperiodic communication, meanwhile, the bandwidth is reserved to improve response time as much as possible without periodic timing violations. The effectiveness of the presented techniques is demonstrated through a series of simulations. Shinpei Kato, Yuji Fujita, Nobuyuki Yamasaki |
RTCSA | 1 |
| 2009 | Gang EDF Scheduling of Parallel Task SystemsabstractThe preemptive real-time scheduling of sporadic parallel task systems is studied. We present an algorithm, called gang EDF, which applies the earliest deadline first (EDF) policy to the traditional gang scheduling scheme. We also provide schedulability analysis of gang EDF. Specifically, the total amount of interference that is necessary to cause a deadline miss is first identified. The contribution of each task to the interference is then bounded. Finally, verifying that the total amount of contribution does not exceed the necessary interference for every task, the schedulability test is derived. Although the techniques proposed herein are based on the prior results for the sequential task model, we introduce new ideas for the parallel task model. Shinpei Kato, Yutaka Ishikawa |
RTSS | 1 |
| 2008 | Work-Conserving Optimal Real-Time Scheduling on MultiprocessorsabstractExtended T-N plane abstraction (E-TNPA) proposed in this paper realizes work-conserving and efficient optimal real-time scheduling on multiprocessors relative to the original T-N plane abstraction (TNPA). Additionally a scheduling algorithm named NVNLF (no virtual nodal laxity first) is presented for E-TNPA. E-TNPA and NVNLF relax the restrictions of TNPA and the traditional algorithm LNREF, respectively. Arbitrary tasks can be preferentially executed by both tie-breaking rules and time apportionment policies in accordance with various system requirements with several restrictions. Simulation results show that E-TNPA significantly reduces the number of task preemptions as compared to TNPA. Kenji Funaoka, Shinpei Kato, Nobuyuki Yamasaki |
ECRTS | 2 |
| 2008 | Portioned EDF-based scheduling on multiprocessorsabstractThis paper presents an EDF-based algorithm, called Earliest Deadline Deferrable Portion (EDDP), for efficient scheduling of recurrent real-time tasks on multiprocessor systems. The design of EDDP is based on the portioned scheduling technique which classifies each task into a fixed task or a migratable task. A fixed task is scheduled on the dedicated processor without migrations. A migratable task is meanwhile permitted to migrate between the particular two processors. In order to curb the cost of task migrations, EDDP makes at most M -- 1 migratable tasks on M processors. The scheduling analysis derives the condition for a given task set to be schedulable. It is also proven that no tasks ever miss deadlines, if the system utilization does not exceed 65%. Beyond the theoretical analysis, the effectiveness of EDDP is evaluated through simulation studies. Simulation results show that EDDP achieves high system utilization with a small number of preemptions, compared with the traditional EDF-based algorithms. Shinpei Kato, Nobuyuki Yamasaki |
EMSOFT | 1 |
| 2008 | Scheduling Aperiodic Tasks Using Total Bandwidth Server on MultiprocessorsabstractThis paper presents real-time scheduling techniques for reducing the response time of aperiodic tasks scheduled with real-time periodic tasks on multiprocessor systems. Two problems are addressed in this paper: (i) the scheduling of aperiodic tasks that can be dispatched to any processors when they arrive, and (ii) the scheduling of aperiodic tasks that must be executed on particular processors on which they arrive. In order to improve the responsiveness to both types of aperiodic tasks, efficient dispatching and migration algorithms are designed based on the Earliest Deadline First (EDF) algorithm and the Total Bandwidth Server (TBS) algorithm. The effectiveness of the designed algorithms is evaluated through simulation studies. Shinpei Kato, Nobuyuki Yamasaki |
EUC (1) | 1 |
| 2008 | Portioned static-priority scheduling on multiprocessorsabstractThis paper proposes an efficient real-time scheduling algorithm for multiprocessor platforms. The algorithm is a derivative of the rate monotonic (RM) algorithm, with its basis on the portioned scheduling technique. The theoretical design of the algorithm is well implementable for practical use. The schedulability of the algorithm is also analyzed to guarantee the worst-case performance. The simulation results show that the algorithm achieves higher system utilizations, in which all tasks meet deadlines, with a small number of preemptions compared to traditional algorithms. Shinpei Kato, Nobuyuki Yamasaki |
IPDPS | 1 |
| 2008 | Energy-Efficient Optimal Real-Time Scheduling on MultiprocessorsabstractOptimal real-time scheduling is effective to not only schedulability improvement but also energy efficiency for real-time systems. In this paper, we propose real-time static voltage and frequency scaling (RT-SVFS) techniques based on an optimal real-time scheduling algorithm for multiprocessors. The techniques are theoretically optimal when the voltage and frequency can be controlled both uniformly and independently among processors. Simulation results show that the independent RT-SVFS technique closely approaches the lower bound on energy consumption if the voltage and frequency can be controlled minutely. Kenji Funaoka, Shinpei Kato, Nobuyuki Yamasaki |
ISORC | 2 |
| 2008 | New Abstraction for Optimal Real-Time Scheduling on MultiprocessorsabstractT-R plane abstraction (TRPA) proposed in this paper is an abstraction technique of real-time scheduling on multiprocessors. This paper presents that NNLF (no nodal laxity first) based on TRPA is work-conserving and optimally solves the problem of scheduling periodic tasks on a multiprocessor system. TRPA can accommodate to dynamic environments due to its dynamic time reservation, while T-N plane abstraction (TNPA) and extended TNPA (E-TNPA) reserve processor time statically at every task release. Kenji Funaoka, Shinpei Kato, Nobuyuki Yamasaki |
RTCSA | 2 |
| 2008 | Global EDF-Based Scheduling with Efficient Priority PromotionabstractThis paper presents an algorithm, called Earliest Deadline Critical Laxity (EDCL), for the efficient scheduling of sporadic real-time tasks on multiprocessors systems. EDCL is a derivative of the Earliest Deadline Zero Laxity (EDZL) algorithm in that the priority of a job reaching certain laxity is imperiously promoted to the top, but it differs in that the occurrence of priority promotion is confined to at the release time or the completion time of a job. This modification enables EDCL to bound the number of scheduler invocations and to relax the implementation complexity of scheduler, while the schedulability is still competitive with EDZL. The schedulability test of EDCL is designed through theoretical analysis. In addition, an error in the traditional schedulability test of EDZL is corrected. Simulation studies demonstrate the effectiveness of EDCL in terms of guaranteed schedulability and exhaustive schedulability by comparing with traditional efficient scheduling algorithms. Shinpei Kato, Nobuyuki Yamasaki |
RTCSA | 1 |
| 2007 | Real-Time Scheduling with Task Splitting on MultiprocessorsabstractThis paper presents a real-time scheduling algorithm with high schedulability and few preemptions for multiprocessor systems. The algorithm is based on an unorthodox method called portioned scheduling that assigns each task to a particular processor like partitioned scheduling but can split a task into two processors if there is not enough capacity remaining on a processor. We describe an algorithm for assigning tasks to processors as well as an algorithm for scheduling the assigned tasks on per-processor. The schedulability analysis provides a formula to calculate the upper bound of the schedulable per-processor utilization for the algorithm. We then prove that the least upper bound of the whole system utilization is 50%. In addition, we propose heuristic procedures to improve schedulability. The simulation results show that the algorithm can often successfully schedule a task set with system utilization much higher than 50%, though the least upper bound is 50%. We also show that the algorithm achieves higher schedulability with fewer preemptions compared to the existing algorithms. Shinpei Kato, Nobuyuki Yamasaki |
RTCSA | 1 |
| 2006 | Extended U-Link Scheduling to Increase the Execution Efficiency for SMT Real-Time SystemsabstractThis paper extends U-Link scheduling to increase the average execution efficiency of the system. We first define the execution efficiency. Then we propose a new algorithm that establishes the co-scheduled sets where the execution efficiency can be increased. Also we present the static estimation of the execution time and provide the schedulability analysis for the extended U-Link scheduling. In the experiments, we evaluate the advancement of the extended U-Link scheduling from the viewpoint of the execution efficiency and real-time processing Shinpei Kato, Nobuyuki Yamasaki |
RTCSA | 1 |
| 2005 | U-Link Scheduling: Bounding Execution Time of Real-Time Tasks with Multi-Case Execution Time on SMT ProcessorsabstractThe goal of this paper is to achieve hard real-time processing with admitting as many tasks as possible on simultaneous multithreaded (SMT) processors. For this goal we propose U-link scheduling scheme that determines the co-scheduled set that is the fixed combinations of co-scheduled tasks to bound the task execution time. Also we present practical algorithms, RR-DUP for building co-scheduled sets and UL-EDF for task scheduling. The performance evaluation shows that UL-EDF with RR-DUP outperforms the conventional scheduling algorithms, EDF-FF and EDF-US, in the point of execution time stability, task rejection ratio and deadline miss ratio. Shinpei Kato, Hidenori Kobayashi, Nobuyuki Yamasaki |
RTCSA | 1 |