VLDB 2026 Research / reviewers in the wild / expert
Takuya Azumi
dblp:20/621
· DBLP profile ↗
72ranked-venue papers
4as first author
38since 2021 · last 2026
0000-0003-0767-4086ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 18 · 1 first-author · 7 since 2021Artificial intelligence and machine learning · 12 · 1 first-author · 8 since 2021Human-computer interaction and ubiquitous computing · 8 · 6 since 2021Software engineering, systems software and programming languages · 7 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 5 since 2021Security and privacy · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SC-MII: Infrastructure LiDAR-based 3D Object Detection on Edge Devices for Split Computing with Multiple Intermediate Outputs Integrationabstract3D object detection using LiDAR-based point cloud data and deep neural networks is essential in autonomous driving technology. However, deploying state-of-the-art models on edge devices present challenges due to high computational demands and energy consumption. Additionally, single LiDAR setups suffer from blind spots. This paper proposes SC-MII, multiple infrastructure LiDAR-based 3D object detection on edge devices for Split Computing with Multiple Intermediate outputs Integration. In SC-MII, edge devices process local point clouds through the initial DNN layers and send intermediate outputs to an edge server. The server integrates these features and completes inference, reducing both latency and device load while improving privacy. Experimental results on a real-world dataset show a 2.19× speed-up and a 71.6% reduction in edge device processing time, with at most a 1.09% drop in accuracy. Taisuke Noguchi, Takayuki Nishio, Takuya Azumi |
CCNC | 3 |
| 2026 | Probabilistic Schedulability Analysis for Mixed-Criticality DAG Tasks on MultiprocessorsabstractMixed-criticality DAG task systems on multiprocessors require schedulability analysis to guarantee that safety-critical tasks meet their deadlines. Conventional approaches rely on worst-case execution times (WCETs), which account for extremely rare pathological scenarios and consequently lead to significant over-provisioning of processor cores. This paper proposes a probabilistic schedulability analysis that exploits the statistical rarity of multiple vertices within a DAG exceeding their expected execution budgets. By grouping vertices within each DAG into small clusters and bounding the probability that more than one vertex in a cluster overruns, the method assigns each cluster a tighter execution budget than the sum of individual WCETs, thereby reducing the number of cores required. The clustering configuration is optimized via simulated annealing to minimize total core usage while maintaining a designer-specified bound on the system-level probability of deadline misses. Experiments on synthetic task sets demonstrate that the proposed method reduces the required number of cores by up to 36% compared to a deterministic baseline, with larger gains for DAGs exhibiting higher internal parallelism. Hiroto Takahashi, Atsushi Yano, Takuya Azumi |
ECRTS | 3 |
| 2026 | ipc_shared_ptr: A Publish/Subscribe-Aware Smart Pointer for Cross-Process Object Lifetime Management
Takahiro Ishikawa-Aso, Atsushi Yano, Koichi Imai, Takuya Azumi, Shinpei Kato |
ISORC | 4 |
| 2025 | D-AWSIM: Distributed Autonomous Driving Simulator for Dynamic Map Generation FrameworkabstractAutonomous driving systems have achieved significant advances, and full autonomy within defined operational design domains near practical deployment. Expanding these domains requires addressing safety assurance under diverse conditions. Information sharing through vehicle-to-vehicle and vehicle-to-infrastructure communication, enabled by a Dynamic Map platform built from vehicle and roadside sensor data, offers a promising solution. Real-world experiments with numerous infrastructure sensors incur high costs and regulatory challenges. Conventional single-host simulators lack the capacity for large-scale urban traffic scenarios. This paper proposes DAWSIM, a distributed simulator that partitions its workload across multiple machines to support the simulation of extensive sensor deployment and dense traffic environments. A Dynamic Map generation framework on D-AWSIM enables researchers to explore information-sharing strategies without relying on physical testbeds. The evaluation shows that D-AWSIM increases throughput for vehicle count and LiDAR sensor processing substantially compared to a single-machine setup. Integration with Autoware demonstrates applicability for autonomous driving research. Shunsuke Ito, Chaoran Zhao, Ryo Okamura, Takuya Azumi |
DSD | 4 |
| 2025 | Partitioned Scheduling for DAG Tasks Considering Probabilistic Execution TimeabstractAutonomous driving systems, critical for safety, require real-time guarantees and can be modeled as DAGs. Their acceleration features, such as caches and pipelining, often result in execution times below the worst-case. Thus, a probabilistic approach ensuring constraint satisfaction within a probability threshold is more suitable than worst-case guarantees for these systems. This paper considers probabilistic guarantees for DAG tasks by utilizing the results of probabilistic guarantees for single processors, which have been relatively more advanced than those for multi-core processors. This paper proposes a task set partitioning method that guarantees schedulability under the partitioned scheduling. The evaluation on randomly generated DAG task sets demonstrates that the proposed method schedules more task sets with a smaller mean analysis time compared to existing probabilistic schedulability analysis for DAGs. The evaluation also compares four bin-packing heuristics, revealing Item-Centric Worst-Fit-Decreasing schedules the most task sets. Fuma Omori, Atsushi Yano, Takuya Azumi |
HPCC | 3 |
| 2025 | TECS/Rust-OE: Optimizing Exclusive Control in Rust-Based Component Systems for Embedded DevicesabstractThe diversification of functionalities and the development of the IoT are making embedded systems larger and more complex in structure. Ensuring system reliability, especially in terms of security, necessitates selecting an appropriate programming language. As part of existing research, TECS/Rust has been proposed as a framework that combines Rust and component-based development (CBD) to enable scalable system design and enhanced reliability. This framework represents system structures using static mutable variables, but excessive exclusive controls applied to ensure thread safety have led to performance degradation. This paper proposes TECS/Rust-OE, a memory-safe CBD framework that optimizes exclusive control by utilizing call flows to address these limitations. The proposed Rust code leverages real-time OS exclusive control mechanisms, optimizing performance without compromising reusability. Rust code is automatically generated based on component descriptions. Evaluations demonstrate reduced overhead due to optimized exclusion control and high reusability of the generated code. Nao Yoshimura, Hiroshi Oyama 0002, Takuya Azumi |
ISORC | 3 |
| 2025 | CART: Combined AUTOSAR AP and ROS 2 Tracing Framework
Ryudai Iwakami, Hiroyuki Hanyu, Tasuku Ishigooka, Takuya Azumi |
RSP | 4 |
| 2025 | Work in Progress: Middleware-Transparent Callback Enforcement in Commoditized Component-Oriented Real-Time SystemsabstractReal-time scheduling in commoditized componentoriented real-time systems, such as ROS 2 systems on Linux, has been studied under nested scheduling: OS thread scheduling and middleware layer scheduling (e.g., ROS 2 Executor). However, by establishing a persistent one-to-one correspondence between callbacks and OS threads, we can ignore the middleware layer and directly apply OS scheduling parameters (e.g., scheduling policy, priority, and affinity) to individual callbacks. We propose a middleware model that enables this idea and implements CallbackIsolatedExecutor as a novel ROS 2 Executor. We demonstrate that the costs (user-kernel switches, context switches, and memory usage) of CallbackIsolatedExecutor remain lower than those of the MultiThreadedExecutor, regardless of the number of callbacks. Additionally, the cost of CallbackIsolatedExecutor relative to SingleThreadedExecutor stays within a fixed ratio (1.4x for inter-process and 5x for intra-process communication). Future ROS 2 real-time scheduling research can avoid nested scheduling, ignoring the existence of the middleware layer. Takahiro Ishikawa-Aso, Atsushi Yano, Takuya Azumi, Shinpei Kato |
RTAS | 3 |
| 2025 | Autoware Toolbox 2: Model-Based Benchmark Suite of ROS 2 Nodes for Autonomous Driving SystemabstractThe rapid development of autonomous driving technology calls for efficient and reliable methods for performance evaluation and optimization. This paper proposes Autoware Toolbox 2, which is a benchmark suite for the autonomous driving system, Autoware, using the Model-Based Development (MBD) approach. By modeling and simulating ROS 2 nodes in MATLAB/Simulink, the benchmark suite enables systematic evaluation of Autoware. Compared to traditional code-based benchmarks, the MBD-based approach improves modularity, reusability, and visualization, streamlining the testing process and reducing development time. Experimental results demonstrate that the proposed benchmark suite is effective in evaluating Autoware, and validate the utility of the benchmark suite and its potential to enhance both development efficiency and functional performance of Autoware. Hiroshi Fujimoto, Takuya Azumi |
SoMeT | 3 |
| 2025 | Exploring Power Usage and Inference Speed Evaluation of Low-power Clustered Many-core PlatformsabstractHigh-performance platforms capable of running deep neural network (DNN)-based applications are necessary for embedded systems such as autonomous-driving systems. These systems must be compact and power-efficient, rather than relying on rich-computational power platforms such as graphics processing units (GPUs). Additionally, platforms capable of executing multiple applications in parallel in such large-scale systems are required. Clustered many-core platforms, such as the Kalray massively parallel processor array (MPPA) 3-80 Coolidge, have been designed to meet these requirements. Coolidge employs more computing cores compared with single-core or multi-core processors, allowing for reduced clock frequencies per core. Consequently, Coolidge is a high-performance platform with low-power usage. Additionally, in Coolidge, cores are grouped into clusters, each of which can independently run different applications. This enables a single Coolidge platform to support multiple applications simultaneously. In the realm of cyber-physical systems (CPS), which bridge the physical and digital domains, these platforms become crucial. CPS relies on real-time embedded systems, like those in autonomous vehicles, which necessitate low-power, high-performance platforms that can perform complex computations like DNNs for object detection. In our evaluation, we examine DNN inference speed and explore the performance of the Coolidge platform. The DNN task for the evaluation employs object detection, which is commonly used in autonomous-driving systems. The acceptable speed threshold for real-time applications in object detection is 30 frames per second. Our study reveals that the “you only look once” v5 model can exceed performance benchmarks with only one cluster in Coolidge for int8 data type and two clusters for the higher precision fp16 data type. Moreover, when two clusters are utilized for DNN tasks, the remaining clusters are available for other non-DNN applications, underlining the platform’s flexibility and resource efficiency. Masahiro Hasumi, Takuma Yabe, Takuya Azumi |
ACM Trans. Cyber Phys. Syst. | 3 |
| 2024 | AUTOSAR AP and ROS 2 Collaboration FrameworkabstractThe field of autonomous vehicle research is advancing rapidly, necessitating platforms that meet real-time performance, safety, and security requirements for practical deployment. AUTOSAR Adaptive Platform (AUTOSAR AP) is widely adopted in development to meet these criteria; however, licensing constraints and tool implementation challenges limit its use in research. Conversely, Robot Operating System 2 (ROS 2) is predominantly used in research within the autonomous driving domain, leading to a disparity between research and development platforms that hinders swift commercialization. This paper proposes a collaboration framework that enables AUTOSAR AP and ROS 2 to communicate with each other using a Data Distribution Service for Real-Time Systems (DDS). In contrast, AUTOSAR AP uses Scalable service-Oriented Middleware over IP (SOME/IP) for communication. The proposed framework bridges these protocol differences, ensuring seamless interaction between the two platforms. We validate the functionality and performance of our bridge converter through empirical analysis, demonstrating its efficiency in conversion time and ease of integration with ROS 2 tools. Furthermore, the availability of the proposed collaboration framework is improved by generating a configuration file automatically for the proposed bridge converter. Ryudai Iwakami, Hiroyuki Hanyu, Tasuku Ishigooka, Takuya Azumi |
DSD | 5 |
| 2024 | Deadline Miss Early Detection Method for DAG Tasks Considering Variable Execution Time
Hayate Toba, Takuya Azumi |
ECRTS | 2 |
| 2024 | Model-Based Development for Autonomous Driving Software Considering ParallelizationabstractIn recent years, autonomous vehicles have attracted attention as one of the solutions to various social problems. However, autonomous driving software requires real-time performance as it considers a variety of functions and complex environments. Therefore, this paper proposes a parallelization method for autonomous driving software using the Model-Based Development (MBD) process. The proposed method extends the existing Model-Based Parallelizer (MBP) method to facilitate the implementation of complex processing. As a result, execution time was reduced. The evaluation results demonstrate that the proposed method is suitable for the development of autonomous driving software, particularly in achieving real-time performance. Kenshin Obi, Takumi Onozawa, Hiroshi Fujimoto, Takuya Azumi |
ETFA | 4 |
| 2024 | Parallelized Code Generation from Simulink Models for Event-driven and Timer-driven ROS 2 NodesabstractIn recent years, the complexity and scale of embedded systems, especially in the rapidly developing field of autonomous driving systems, have increased significantly. This has led to the adoption of software and hardware approaches such as Robot Operating System (ROS) 2 and multi-core processors. Traditional manual program parallelization faces challenges, including maintaining data integrity and avoiding concurrency issues such as deadlocks. While model-based development (MBD) automates this process, it encounters difficulties with the integration of modern frameworks such as ROS 2 in multi-input scenarios. This paper proposes an MBD framework to overcome these issues, categorizing ROS 2-compatible Simulink models into event-driven and timer-driven types for targeted parallelization. As a result, it extends the conventional parallelization by MBD and supports parallelized code generation for ROS 2-based models with multiple inputs. The evaluation results show that after applying parallelization with the proposed framework, all patterns show a reduction in execution time, confirming the effectiveness of parallelization. Kenshin Obi, Ryo Yoshinaka, Hiroshi Fujimoto, Takuya Azumi |
SEAA | 4 |
| 2024 | ROS-lite2: Autonomous-driving Software Platform for Clustered Many-core ProcessorabstractIn intelligent robotics, which spans from assistive devices to automation, the development of autonomous systems merges computational and physical capabilities, such as in autonomous wheelchairs. Many-core processor, essential for real-time operations, pose challenges to software adaptability. This paper introduces ROS-lite2, which is a software platform for autonomous vehicles, utilizing a ROS 2 framework based on many-core processor to boost flexibility and simplify deployment. Our method facilitates complex function integration and intuitive operation with less hardware knowledge. Experiments with an autonomous wheelchair demonstrate the platform’s effectiveness in improving autonomy and reducing development effort, advancing robotic assistance. Yuta Tajima, Shuhei Tsunoda, Takuya Azumi |
IROS | 3 |
| 2024 | TECS/Rust: Memory-safe Component Framework for Embedded SystemsabstractAs embedded systems grow in complexity and scale due to increased functional diversity, component-based development (CBD) emerges as a solution to streamline their architecture and enhance functionality reuse. CBD typically utilizes the C programming language for its direct hardware access and low-level operations, despite its susceptibility to memory-related issues. To address these concerns, this paper proposes TOPPERS Embedded Component Systems/Rust (TECS/Rust), a Rust-based framework specifically designed for TECS, which is a component framework for embedded systems. It leverages Rust’s compile-time memory-safe features, such as lifetime and borrowing, to mitigate memory vulnerabilities common with C. The proposed framework not only ensures memory safety but also maintains the flexibility of CBD, automates Rust code generation for CBD components, and supports efficient integration with real-time operating systems. An evaluation of the amount of generated code indicates that the code generated by this paper framework accounts for a large percentage of the actual code. Compared to code developed without the proposed framework, the difference in execution time is minimal, indicating that the overhead introduced by the proposed framework is negligible. Nao Yoshimura, Hiroshi Oyama 0002, Takuya Azumi |
ISORC | 3 |
| 2024 | Point Cloud Automatic Annotation Framework for Autonomous DrivingabstractIn autonomous driving systems, infrastructure LiDAR technology provides advanced point cloud information of the road, allowing for preemptive analysis, which increases decision-making time. 3D object detection affords autonomous vehicles the ability to recognize and understand surrounding environmental objects accurately. To further investigate the optimal deployment locations and impacts of infrastructure LiDAR in autonomous driving systems, we have developed an automated annotation framework integrated into an autonomous driving simulator. This framework enables the automated labeling of point cloud data and the rapid construction of datasets, significantly reducing the time required for users to create such datasets. Additionally, we enhanced the usability of the autonomous driving simulator, allowing for real-time adjustments of LiDAR settings during operation, and the generation of vehicle NPCs in accordance with the OpenSCENARIO 2.0 standard. Finally, utilizing this automatic annotation framework, we conducted an evaluation of the impact of various types of LiDAR (dense point clouds and sparse point clouds) and their quantities on the accuracy of 3D object detection models. The experimental evaluation shows that the number of points in infrastructure point clouds and the detection range have a significant impact on 3D detection models. Upon replacing VLP-16 with MID70, the performance of various models improved significantly, with a maximum increase of 50% mAP. Chaoran Zhao, Takuya Azumi |
IV | 3 |
| 2024 | Preliminary Approach to Parallelizing Autonomous Driving Applications Using High-Performance Many-core ProcessorabstractAdvancements in autonomous driving technologies leverage real-time computing and embedded systems to enable vehicles to make quick decisions based on dynamic road conditions. These increasingly complex systems face rising computational demands. This study introduces a preliminary approach to parallelizing autonomous driving applications using a high-performance many-core processor. By distributing tasks across multiple cores, the approach enhances concurrent execution, reduces conflicts, and minimizes resource contention, improving the efficiency and performance of autonomous driving systems. Xuankeng He, Takuya Azumi |
RTCSA | 2 |
| 2024 | Work-in-Progress: Multi-Deadline DAG Scheduling Model for Autonomous Driving SystemsabstractAutoware is an autonomous driving system implemented on Robot Operation System (ROS) 2, where an end-to-end timing guarantee is crucial to ensure safety. However, existing ROS 2 cause-effect chain models for analyzing end-toend latency struggle to accurately represent the complexities of Autoware, particularly regarding sync callbacks, queue consumption patterns, and feedback loops. To address these problems, we propose a new scheduling model that decomposes the end-to-end timing constraints of Autoware into local relative deadlines for each sub-DAG. This multi-deadline DAG scheduling model avoids the need for complex analysis of data flows through queues and loops, while ensuring that all callbacks receive data within correct intervals. Furthermore, we extend the Global Earliest Deadline First (GEDF) algorithm for the proposed model and evaluate its effectiveness using a synthetic workload derived from Autoware. Atsushi Yano, Takuya Azumi |
RTSS | 2 |
| 2023 | TILDE: Topic-Tracking Infrastructure for Dynamic Message Latency and Deadline Evaluator for ROS 2 ApplicationabstractAutonomous-driving systems require real-time processing for safety. Detecting the deadline miss (disabling to execute the process until the designated time) of autonomous-driving software is essential for executing real-time processing. We proposed the application of a dynamic message tracking system (Topic-tracking Infrastructure for Dynamic Message Latency and Deadline Evaluator) to software based on Robot Operating System (ROS) 2, e.g., Autoware (autonomous-driving software). TILDE transmits MessageTrackingTag with ROS 2 messages to trace and dynamically specify the data flow of the messages. In particular, MessageTrackingTag contains input and output times at each node with the data flow. The information of the MessageTrackingTag is acquired by a deadline detector to identify the time from the input to output (hereafter latency) of the data flow. The latency is compared to the deadline to detect the occurrence of a deadline miss. In principle, the deadline detector is a package of TILDE for notifying deadline misses. In addition, this paper conducted two evaluations of TILDE. First, we measured the overhead caused by embedding TILDE on ROS 2 nodes through the transmission of a String-type message from one node to another. Second, we measured the deadline miss detection rate using CARET (a latency analysis tool for autonomous-driving systems) to validate the accuracy of the deadline detector in determining deadline misses. Xuankeng He, Hiromi Sato, Yoshikazu Okumura, Takuya Azumi |
DS-RT | 4 |
| 2023 | DAG Scheduling for Clustered Many-Core Processor Considering Execution-Time-Reduction EffectivenessabstractIn recent years, there has been a surge in research on autonomous systems such as unmanned aerial vehicles and autonomous-driving systems. These systems require high computing power with low power consumption and strict real-time requirements. Correspondingly, clustered many-core processors are attracting attention as a solution to these requirements. In this study, we propose a core allocation method for DAG scheduling using clustered many-core processors. Specifically, we propose two methods for selecting nodes to be allocated based on the execution-time-reduction effectiveness, assuming that each node in the DAG can perform intra-node processing in parallel. In addition, we compared the response times obtained from scheduling using each of the two methods to determine which method is superior. Yutaro Nozaki, Takuya Azumi |
DS-RT | 2 |
| 2023 | Estimation of Deadline Miss Rate for DAG Mixed Timer-Driven and Event-Driven NodesabstractAutonomous-driving systems are becoming increasingly complex, making it more difficult to predict timing. Autonomous-driving applications, such as localization and path planning, can be represented by a Directed Acyclic Graph (DAG) consisting of timer-driven and event-driven nodes. Various estimation methods have been proposed for determining the response time of a DAG. However, most existing studies rely on worst-case execution time estimation, leading to pessimistic results. Additionally, there is a lack of research on DAGs that incorporate a combination of timer-driven and event-driven nodes. To address this issue, a proposed method estimates the response time and deadline miss rates considering the variation in execution time for DAGs with mixed node types. By considering execution time variation, the proposed method allows for multiple values of response time and deadline miss rate based on probability. When a sufficient number of cores are available, the proposed method accurately estimates the deadline miss rate. Daichi Yamazaki, Takuya Azumi |
DS-RT | 2 |
| 2023 | HRMP3+TECS: Component Framework for Multiprocessor Real-time Operating System with Memory ProtectionabstractThe scale and demand for protection functionalities in embedded systems continue to grow as Internet of things technology develops. Simultaneously, multiprocessor real-time operating systems (RTOSs) with memory protection functionalities are broadening in use. On the other hand, the large development effort and poor reusability of multiprocessor RTOS-based development remain hindrances to their use. To solve this problem, this paper proposes a component framework for multiprocessor RTOSs that includes memory protection. In the proposed framework, OS functionalities are reframed as components. The allocation of objects to processors and protection settings, which are supported by the OS, can then be configured based on the component description. Plugins are implemented here to generate files for object generation from component descriptions of OS functionalities. In addition, access to the protection domain can be defined based on the component description. Finally, test programs are used in a performance evaluation. The proposed framework enables extensions to the component-based OS while maintaining most of the functionality and performance of the target OS. Yoshitada Takaso, Hiroshi Oyama 0002, Hiroaki Takada, Takuya Azumi |
ISORC | 4 |
| 2023 | Automated Testing Framework for Embedded Component SystemsabstractEmbedded systems in equipment have recently become larger and more complex. To solve these problems and improve development efficiency, a Component-Based Development (CBD) method could be used to divide the system into components. However, CBD presents challenges in system testing, such as an increase in the number of test objects and occurrence of failures in team development. To address these issues, a development method called Continuous Integration (CI) is sometimes used, which automatically performs building, and testing. This paper proposes a CI framework for automated testing that can be used for embedded components. In addition to automated testing, the proposed framework can perform line coverage measurement and display, Boundary Value Testing, and Equivalence Partitioning Testing. Furthermore, the evaluation of the proposed framework yielded the following contributions: Automated Workflow, Automated Code Generation, and Line Coverage Measurement Hinata Tomimori, Hiroshi Oyama 0002, Takuya Azumi |
ISORC | 3 |
| 2023 | RD-Gen: Random DAG Generator Considering Multi-rate Applications for Reproducible Scheduling EvaluationabstractReal-time systems have various requirements such as the deadline and resource constraints. In addition, real-time systems are becoming larger and more complex, and studies on performance analysis and efficient scheduling algorithms are becoming increasingly important. Directed acyclic graph (DAG) models, which can express task dependencies and parallelism, are used for such studies. Random DAG sets are used to demonstrate the effectiveness and objectivity of methods proposed for real-time systems. However, there is no random DAG generation tool available that can generate a DAG set that considers the latest multi-rate applications. Therefore, researchers need to generate random DAG sets on their own, leading to additional effort and reduced reliability and reproducibility. To solve this problem, we propose a random DAG generator considering multi-rate applications for reproducible scheduling evaluation (RD-Gen). RD-Gen also enables batch generation of random DAG sets with different parameters. Case studies are used to demonstrate that RD-Gen can manage various problem settings and DAG study requirements. Atsushi Yano, Takuya Azumi |
ISORC | 2 |
| 2023 | Work-in-Progress: Federated and Bundled-Based DAG SchedulingabstractIn the rapidly evolving landscape of real-time systems, particularly autonomous-driving technologies, directed acyclic graphs (DAGs) have emerged as an essential model for delineating task dependencies and facilitating parallel processing. This growing complexity and computational demand are making it increasingly challenging to efficiently schedule tasks on multi-core processors. Among various DAG scheduling algorithms, federated scheduling stands out for its efficacy in managing tasks with implicit deadlines. However, it comes at the cost of potentially over-consuming core resources and diminishing the overall schedulability of task set. To address these challenges, this paper proposes a hybrid scheduling method tailored for sporadic DAG tasks, which integrates principles from bundled scheduling to optimize core utilization and enhance system schedulability. Tomoya Kobayashi, Takuya Azumi |
RTSS | 2 |
| 2023 | Model-based Development for ROS 2-based Autonomous-driving SoftwareabstractAutonomous vehicles have attracted increasing research attention as a solution to various contemporary social problems. Current code-based development methods require time to understand the processing of open-source, complex autonomous-driving systems, and parallelization is performed manually when verifying the operation of such systems. Thus, a method for parallelizing autonomous-driving software using the Model-based Development (MBD) process was proposed. The parallelized code for autonomous-driving software can be generated semi-automatically. The proposed method realizes the generation of code that assumes operation on real-world hardware and reduces execution time. Takumi Onozawa, Hiroshi Fujimoto, Takuya Azumi |
TrustCom | 3 |
| 2023 | Performance Evaluation Framework for Arbitrary Nodes of Autonomous-driving SystemsabstractAutonomous-driving software is increasingly based on Robot Operating System (ROS) 2, middleware for developing robotic applications. This paper proposes a performance evaluation framework for arbitrary nodes of Autonomous-driving software. The proposed framework empowers ROS 2 application developers to evaluate particular nodes requiring evaluation efficiently. The proposed framework can activate each node by considering the logic when the application is run as a whole. In addition, input data specific to each node are required for node evaluation. Because each node exchanges various types of data, generating input data to evaluate a particular node can be difficult. Thus, the proposed framework generates input data suitable for node evaluation. The proposed framework also provides a graphical user interface (GUI). A developer can easily evaluate arbitrary nodes by manipulating the proposed framework on the GUI. Experimental results show that the proposed framework successfully activates the evaluation target node and evaluates arbitrary nodes of autonomous-driving software. In addition, each functionality of the proposed framework is demonstrated to be working properly. The results of user functionality testing of the proposed framework demonstrate its usefulness. Yuta Tajima, Tatsuya Miki, Takuya Azumi |
TrustCom | 3 |
| 2022 | DAG Scheduling Considering Parallel Execution for High-Load Processing on Clustered Many-core ProcessorsabstractIn recent years, high computational power has been required for computer platforms to support complex systems such as self-driving systems. Clustered many-core processors and directed acyclic graphs (DAGs), which can represent dependencies and parallelism of task processing, have attracted much attention as solutions to this problem. Previous studies on scheduling DAGs on multi-core processors have attempted to reduce the makespan (i.e., time it takes for a task to complete) by increasing the number of processes that can be executed in parallel. However, in self-driving systems, such as those utilizing clustered many-core processors, it is impossible to sufficiently increase the utilization of processor cores due to high-load processing. In this paper, a scheduling method is proposed to improve the utilization of processor cores by parallel executing high-load processes in parallel across multiple cores. The proposed method can reduce the makespan of DAGs performing high-load processing on clustered many-core processors. Ryo Okamura, Takuya Azumi |
DS-RT | 2 |
| 2022 | CARET: Chain-Aware ROS 2 Evaluation ToolabstractThis paper presents a tool to evaluate the latency of Robot Operating System (ROS) 2 applications called Chain-Aware ROS 2 Evaluation Tool (CARET). ROS 2 is designed to enhance the modularity of real-time robotic applications, including self-driving software such as Autoware. To analyze the performance of ROS 2 applications, CARET supports measurement functionalities for the callback latency, node latency, communication time between nodes, and end-to-end latency. To calculate each latency, CARET provides tracepoints, an architecture file, information on tracepoint connections, and a message tracking functionality. Furthermore, CARET can visualize different types of latency, such as bottlenecks and lost message, to analyze ROS 2 applications. The experimental results demonstrate that CARET can successfully measure the end-to-end latency of Autoware.Universe, which is ROS 2-based self-driving software. Takahisa Kuboichi, Atsushi Hasegawa, Keita Miura, Kenji Funaoka, Shinpei Kato, Takuya Azumi |
EUC | 7 |
| 2022 | Component Framework for Multiprocessor Real-Time Operating SystemsabstractMultiprocessor real-time operating systems (RTOSs) are in high demand to deal with the large complexity of embedded systems. Developers can reduce power consumption, cope with increased size, and improve processing performance using multiprocessor RTOSs. However, multiprocessor RTOSs for embedded systems have two major problems. The handling of kernel objects, functionalities of Multiprocessor RTOSs, makes system reusability difficult. The same goes for the low reusability of interprocessor communication in multiprocessor systems. To solve these problems, this paper proposes a component framework for multiprocessor RTOSs. Kernel objects were componentized based on TOPPERS embedded component systems, a component-based development framework. The proposed framework supports object allocation on multiprocessors, and implements automatic code generation plugins for system reusability and flexible interprocessor communication methods depending on the system. The reusability of the system and inter-processor communication will be verified to be easy while maintaining real-time performance and resource constraints. Yoshitada Takaso, Hiroshi Oyama 0002, Takuya Azumi |
EUC | 3 |
| 2022 | Communication Overhead Schema Independent of Libraries for Software/Hardware InterfaceabstractEvery year, embedded systems, such as self-driving systems, are becoming larger and more complex which requires high computing power yet low power consumption. To meet these requirements, there is an increase in the usage of many-core processors. However, thorough understanding of the concept of many-core processors is needed for these to be adopted. Therefore, Software-Hardware Interface for Multi-Many-Core (SHIM), a hardware abstraction description, was developed to reduce the burden on developers. However, SHIM cannot fully demonstrate communication overhead as it describes overhead in instruction units. For this reason, the communication overhead cannot be used to estimate execution time. With this, we propose a new schema that can describe the communication overhead in API units. In other words, the communication API overhead can be described without relying on communication libraries. To improve the usability of the existing schema, which is SHIM, we incorporated the proposed schema into SHIM. We set up what is required of the proposed schema, considered use cases, and actually created instance diagrams to evaluate the proposed schema. We also compared the proposed schema with SHIM and determined what needs to be modified when incorporating the schema. As a result, the proposed schema can express communication overhead that varies with the combination of cores and message size, while significantly reducing the amount of description. Yutaro Kobayashi, Hiroshi Fujimoto, Takuya Azumi |
SoMeT | 3 |
| 2022 | Self-Driving Software Benchmark for Model-Based DevelopmentabstractThis study describes the MATLAB/Simulink benchmark of an open-source self-driving system based on Robot Operating System 2 (ROS 2). In recent years, self-driving systems have been the subject of research and development worldwide. Due to the lack of open source models for self-driving systems, model-based development, a common approach to the development of in-vehicle systems, is not yet fully used in the development of self-driving systems. The provided MATLAB/Simulink benchmarks support the design of ROS 2-based self-driving systems using MATLAB/Simulink. Improvements to the benchmark’s design issues are discussed. Furthermore, Simulink’s profiling function makes it easier to make redesign decisions. The model can run using only sensor data, and the runtime evaluation revealed that the benchmark models could reduce the runtime, although the number of cores used was different. Takumi Onozawa, Takuya Azumi |
SoMeT | 2 |
| 2022 | Autoware_Perf: A tracing and performance analysis framework for ROS 2 applications
Atsushi Hasegawa, Takuya Azumi |
J. Syst. Archit. | 3 |
| 2021 | Federated Scheduling in Clustered Many-core ProcessorsabstractHigh-performance embedded systems, such as self-driving systems require platforms that reduce power consumption and perform high-performance processing. As satisfying both requirements, multi-tmany-core processors are attracting attention. This paper focuses on a clustered many-core processors represented by Kalray MPPA. Clustered many-core processors regard multiple cores as a single cluster, with each cluster having private memory. A bus within each cluster performs communication. A network-on-chip (NoC) is used between clusters. These two communication speeds are different. However, it is inappropriate to always assume that the worst calculation is the slow NoC communication speed. This paper discusses a method to improve worst-case communication time estimation when applying federated scheduling on clustered many-core processors with routes at a different speed. Furthermore, this paper investigates an approach to assign dedicated cores in multiple clusters to reduce the number of required cores. Ryotaro Koike, Takuya Azumi |
DS-RT | 2 |
| 2021 | Contention-Free Scheduling Algorithm Using LET Paradigm for Clustered Many-core ProcessorabstractSelf-driving systems require multi-/many-core platforms with high computing power and low power consumption. However, for hard real-time applications, multiple demands on shared resources can impede real-time performance. Therefore, making the timing of memory access deterministic is important. The logical execution time (LET) paradigm has gained attention as a means to achieve this purpose. However, this approach lacks scalability owing to the overhead caused because the LET paradigm is set longer than the actual execution time of the task. This paper proposes a theoretical scheduling method for a model applying the LET paradigm to the directed acyclic graph (DAG) nodes for a multi-/many-core platform. The proposed method considers communication timing and generates a schedule that does not cause communication contentions. In addition, the proposed method performs a parallel calculation of tasks to deal with the overhead caused by adopting the LET paradigm. Atsushi Yano, Shingo Igarashi, Takuya Azumi |
DS-RT | 3 |
| 2021 | Dynamically Interchangeable Framework for Component Behavior of Embedded Component SystemsabstractThe development of information technology and the improved performance of embedded devices have led to embedded software of increasing scale and complexity. Component-Based Development (CBD), which divides software into components and subsystems, is effective for efficiently developing large-scale embedded software. However, components created in CBD are not interchangeable dynamically while the application is running, and components written in C must be compiled to change and run. To address this issue, this paper proposes the component framework using mruby to change the component behavior dynamically. The proposed framework works on TOPPERS Embedded Component System written in C. Experimental results show that the time it takes to change the behavior of the component is shorter using the proposed framework than using a component written in C. The proposed framework can realize CBD more efficiently. Ryota Shimomura, Hiroshi Oyama 0002, Takuya Azumi |
EUC | 3 |
| 2021 | Work-in-Progress: Reinforcement Learning-Based DAG Scheduling Algorithm in Clustered Many-Core PlatformabstractEmbedded systems have become extensive, complex, and automated; thus, increasingly, computing platforms for such systems are being transformed into multi-/many-core platforms. Typically, self-driving systems, involve various applications that run simultaneously, and such systems require low power consumption and large-scale computation. A many-core processor with instructions, multiple data architecture can satisfy these requirements. Shortening the time required to execute all tasks (i.e., makespan) is an important objective in task scheduling for parallel real-time systems, such as self-driving system. Machine learning algorithms have been introduced to solve this kind of problem. This paper proposes a reinforcement learning-based scheduling algorithm for parallel real-time systems represented by a directed acyclic graph (DAG), and Kalray MPPA3-80 is used as a target many-core processor. Atsushi Yano, Takuya Azumi |
RTSS | 2 |
| 2020 | Heuristic Contention-Free Scheduling Algorithm for Multi-core Processor using LET ModelabstractEmbedded systems, e.g., self-driving systems and advanced driver-assistance systems (ADAS), require computing platforms with high computing power and low power consumption. Multi-/many-core platforms satisfy these requirements effectively. However, for hard real-time applications, multiple demands on shared resources can impede real-time performance, and memory is one resource that can impair the desired performance significantly. Therefore, it is important that memory access timing be deterministic to facilitate predictability. To realize this, the Logical Execution Time (LET) paradigm is currently attracting attention. This paper proposes a theoretical scheduling method for a model applying the LET paradigm to directed acyclic graph (DAG) nodes for a multi-/many-core platform. The proposed method considers communication timing between nodes and generates a schedule that does not cause communication contentions. In addition, the proposed method attempts to distribute tasks and reduce LET intervals to address increased execution times due to the implementation of the LET paradigm. In the evaluation, we observed that the proposed method improved the schedule length by up to 40%. Shingo Igarashi, Tasuku Ishigooka, Tatsuya Horiguchi, Ryotaro Koike, Takuya Azumi |
DS-RT | 5 |
| 2020 | Model-Based Development Considering Self-Driving Systems for Many-Core ProcessorsabstractEmbedded systems, such as self-driving systems, consist of multiple applications interacting in a complex way. Many-core processors can execute high-load arithmetic processing for self-driving systems with low power consumption. Applications must be parallelized to achieve high-speed processing with many-core processors; however, manual parallelization is difficult. Model-based development makes it possible to automate the parallelization of one application (model) for many-core processors. However, a system composed of multiple models, such as self-driving systems, cannot be parallelized for many-core processors. In this paper, we propose a model-based parallelization method compatible with the Robot Operating System for parallelizing a system composed of multiple models. Experimental evaluation revealed that the code generated by the proposed method has the same performance as those manually written by the code. In addition, we propose a data parallelization method to support a model that inputs very large data, such as a self-driving system. The evaluation demonstrates that the proposed method improves the data parallelism of the Simulink model. Ryo Yoshinaka, Takuya Azumi |
ETFA | 2 |
| 2020 | Estimation Method Considering OS Overheads for Embedded Many-Core PlatformabstractEmbedded systems such as automotive systems require high computational power and low power consumption. To meet these requirements, many-core processors have attracted attention. Compared with single-core processors, many-core processors can execute multiple processes in parallel, allowing for improved power consumption and performance. Due to their performance, many-core processors can be used in automotive systems which have strict real-time requirements. Therefore, it is important to obtain accurate hardware and software information. In this paper, we propose a method for estimating the execution time of an application supported by MATLAB/Simulink on a many-core platform. The proposed method considers operating system (OS) overheads with a real-time OS and Kalray MPPA-256 cluster structure which contains many-core processors. The effectiveness, and the experimental results demonstrate that the proposed method is more accurate than existing methods. Kentaro Honda, Hiroshi Fujimoto, Takuya Azumi |
EUC | 3 |
| 2020 | Converting Driving Scenario Framework for Testing Self-Driving SystemsabstractThis paper presents a converting driving scenarios to generate a format suitable for LGSVL simulator. Autonomous vehicles are developed worldwide. To reduce the test cost, virtual simulators are used in the automotive industry. However, the virtual simulators require simulation environments, such as roads, pedestrians, and vehicles. To set these parameters of many test cases is not efficient for developers. MATLAB/Simulink can create the scenarios graphically. Therefore, we proposed a framework to convert scenarios created with MATLAB/Simulink to a format suitable for LGSVL simulator which is an autonomous vehicle simulator. The converted scenarios can perform in LGSVL simulator like the scenarios defined with MATLAB/Simulink. Moreover, LGSVL simulator has cooperated with Autoware, which is an open-source self-driving system. This cooperation facilitates testing the self-driving systems. This framework can help developers create scenarios. Keita Miura, Takuya Azumi |
EUC | 2 |
| 2020 | Unit Testing Framework for Embedded Component SystemsabstractEffective construction of large, complex embedded systems involve the enhancement of potential reusability by dividing the software into subsystems and converting them into distinct parts (i.e., component-based development (CBD) for embedded systems). CBD can be used to improve development efficiency and reduce costs and is also applied to ensure software reusability: however, there are few approaches for testing CBD systems. General CBD systems do not provide methods for evaluating whether each component behaves as expected, making it necessary to manually connect individual components, which makes it difficult to test and fix bugs. To address these issues, this paper describes a unit testing framework for the embedded component systems (called TECSUnit), allowing the assessment of the behavior of each component. This framework increases the efficiency of testing component systems based on a design focusing on flexibility and efficiency. Shuichiro Morisaki, Seito Shirata, Hiroshi Oyama 0002, Takuya Azumi |
EUC | 4 |
| 2020 | ROS-lite: ROS Framework for NoC-Based Embedded Many-Core PlatformabstractThis paper proposes ROS-lite, a robot operating system (ROS) development framework for embedded many- core platforms based on network-on-chip (NoC) technology. Many-core platforms support the high processing capacity and low power consumption requirement of embedded systems. In this study, a self-driving software platform module is parallelized to run on many-core processors to demonstrate the practicality of embedded many-core platforms. The experimental results show that the proposed framework and the parallelized applications have met the deadline for low-speed self-driving systems. Takuya Azumi, Yuya Maruyama, Shinpei Kato |
IROS | 1 |
| 2020 | Mapping Method of MATLAB/Simulink Model for Embedded Many-Core PlatformabstractMulti-/many-core processors are being increasingly used to reduce power consumption and improve performance. In addition, the use of Model-Based Development for embedded systems has been increasing. Relative to these trends, Model-Based Parallelizer (MBP) has an essential role in parallelizing applications (i.e., Simulink blocks) at the model level. MBP maps Simulink blocks to cores using various types of information such as block characteristics, a C code, and the multi-/many-core hardware implementation. However, MBP does not consider many-core hardware with cluster structures. This paper proposes an algorithm that decides on core allocations by considering cluster structures. The proposed algorithm combines two other algorithms: one algorithm uses the core allocation of MBP and path analysis at the cluster-level and considers the influence of communication contention to decide on cluster allocations, and the other algorithm uses the results of MBP and remaps cluster allocations. The proposed algorithm produces better results than its component algorithms could separately. Evaluations demonstrate that the proposed algorithm obtained the better results than the existing method in terms of execution time on random and real models. Kentaro Honda, Sasuga Kojima, Hiroshi Fujimoto, Masato Edahiro, Takuya Azumi |
PDP | 5 |
| 2020 | Accurate Contention Estimate Scheduling Method Using Multiple Clusters of Many-core PlatformabstractEmbedded systems such as self-driving systems require a computing platform with high computing power and low power consumption. Multi-/many-core platforms satisfy exactly these requirements. However, for hard real-time applications, multiple demands on shared resources can hinder real-time performance. Memory is among the resources that can most dramatically impair the desired performance. Therefore, we addressed contentions induced by the shared memory. We improve the predictability of contentions by dividing tasks into the memory access phase and the execution phase using a Directed Acyclic Graph (DAG). Existing methods are able to make accurate contention estimations for one Compute Cluster (CC) of a Clustered many-core processor. Our method is able to do the same for multiple CCs, thereby doubling the scalability in consideration of contentions. Using an Integer Linear Programming (ILP) formulation, we produced a static, non-preemptive, partitioned, time-triggered schedule. We also conducted an experiment in order to minimize the makespan. The evaluation confirmed that our new method reduced the makespan by increasing the number of CCs. Shingo Igarashi, Yuto Kitagawa, Takuro Fukunaga, Takuya Azumi |
PDP | 4 |
| 2019 | Resource Manager for Scalable Performance in ROS Distributed EnvironmentsabstractThis paper presents a resource manager to achieve scalable performance in Robot Operating System (ROS) for distributed environments. In robotics, using ROS in distributed environments via multiple host machines is trending for large-scale data processing, for example, cloud/edge computing and the data communication of point clouds and images in dynamic map composition. However, ROS is unable to manage the resources (e.g., the CPUs, memory, and disks) on each host machine. Therefore, it is difficult to use distributed environmental resources efficiently and achieve scalable performance. This paper proposes a resource management mechanism for ROS distributed environments using a master-slave model to execute ROS processes efficiently and smoothly. We manage the resource usage of each host machine and construct a mechanism to adaptively distribute the load to be balanced. Evaluations show that scalable performance can be achieved in ROS distributed environments comprising ten host machines using a real application (SLAM: simultaneous localization and mapping) processing large-scale point cloud data. Daisuke Fukutomi, Takuya Azumi, Shinpei Kato, Nobuhiko Nishio |
DATE | 2 |
| 2019 | Multi-rate DAG Scheduling Considering Communication Contention for NoC-based Embedded Many-core ProcessorabstractComputing platforms for embedded systems are increasingly being transformed into multi/many-core platforms because embedded systems have become extensive, complex, and automated. In the case of an autonomous driving system, various applications are simultaneously running, and low power consumption and large-scale calculation are required. Many-core processors with a multiple instruction, multiple data (MIMD) architecture can meet these requirements. This paper proposes a scheduling algorithm for an automotive driving system expressed in a directed acyclic graph (DAG) and we use Kalray MPPA-256 as the target many-core processor. On the basis of the architecture of Kalray MPPA-256, task processing that requires large-scale calculation and intercore communication is performed while avoiding communication contention by using a proposed grouping computational resource. In addition, we propose a scheduling method for a multi-rate DAG which is a DAG with multiple periods. This method generates a DAG task in a hyperperiod and schedules the DAG with dependency on tasks that have been released closely. The formulas for prioritization and processor selection are proposed for various generated tasks in a hyperperiod. Evaluation results show that the proposed algorithm is superior to existing DAG scheduling algorithms with regard to schedulability and deadline miss ratio. Shingo Igarashi, Yuto Kitagawa, Tasuku Ishigooka, Tatsuya Horiguchi, Takuya Azumi |
DS-RT | 5 |
| 2019 | MATLAB/Simulink Benchmark Suite for ROS-based Self-driving Software PlatformabstractIn recent years, self-driving systems have been developed worldwide, and the technology has been making remarkable progress. One approach to the development of the autonomous vehicle is using ROS which is an open-source middleware framework used for developing robot applications. On the other hand, the popular approach in the automotive industry is using MATLAB/Simulink which is the software for modeling, simulating, and analyzing. MATLAB/Simulink has an interface connecting ROS and MATLAB/Simulink. However, it is not used much in the development of self-driving systems because there are not enough samples for the self-driving systems. Therefore, we provide a MATLAB/Simulink benchmark suite for a ROS-based self-driving system called Autoware. Autoware provides an abundant set of self-driving modules and enables to simulate and operate the autonomous vehicle. The provided benchmark is a set of MATLAB code and Simulink model samples. They assist to design the self-driving systems using MATLAB/Simulink. Shota Tokunaga, Keita Miura, Takuya Azumi |
ISORC | 3 |
| 2019 | Autoware Toolbox: MATLAB/Simulink Benchmark Suite for ROS-based Self-driving Software PlatformabstractThis paper describes a MATLAB/Simulink benchmark suite for an open-source self-driving system based on Robot Operating System (ROS). In recent years, self-driving systems have been developed worldwide, and the technology has been making remarkable progress. One approach to the development of the self-driving systems is the utilization of ROS which is an open-source middleware framework used for developing robot applications. On the other hand, the popular approach in the automotive industry is the utilization of MATLAB/Simulink which is the software for modeling, simulating, and analyzing. MATLAB/Simulink provides an interface between ROS and MATLAB/Simulink that enables to create functionalities of ROS-based robots in MATLAB/Simulink. However, it has not been fully utilized in the development of the self-driving systems yet because there are not enough open-source models for self-driving. Thus the co-development is difficult. Therefore, we provide a MATLAB/Simulink benchmark suite for a ROS-based self-driving system called Autoware. Autoware is popular open-source software that provides a complete set of self-driving modules. The provided benchmark contains MATLAB/Simulink models. They help to design the ROS-based self-driving systems using MATLAB/Simulink. Moreover, we investigated other benchmarks for self-driving and indicated that this is the first work aiming to the ROS-based self-driving systems with MATLAB/Simulink. Keita Miura, Shota Tokunaga, Noriyuki Ota, Yoshiharu Tange, Takuya Azumi |
RSP | 5 |
| 2019 | Work in Progress: Considering Heuristic Scheduling for NoC-Based Clustered Many-Core Processor Using LET ModelabstractEmbedded systems such as self-driving systems and advanced driver-assistance systems (ADAS) require a computing platform with high computing power and low power consumption. Many-core platforms satisfy exactly these requirements. However, for hard real-time applications, multiple demands on shared resources can hinder real-time performance. Memory is among the resources that can most dramatically impair desired performance. In particular, it is important that the timing for accessing the memory is deterministic. As a method for realizing this, the Logical Execution Time (LET) paradigm is currently attracting attention. Research on the LET model is mainly conducted for control systems such as engines. However, scalability is required for application to autonomous driving systems, and high-speed task scheduling and mapping to the core must be considered. This paper summarizes the scheduling issues of hard real-time applications on a many-core processor using the LET model and discusses possible approaches. Shingo Igarashi, Takuya Azumi |
RTSS | 2 |
| 2018 | DAG Scheduling Algorithm for a Cluster-Based Many-Core ArchitectureabstractThis paper proposes a directed acyclic graph (DAG) scheduling algorithm for cluster-based many-core architecture. Most of DAG scheduling methods that consider multiple processors and communication delays use a heuristic approach because it is difficult to shorten a schedule length (i.e.,makespan). Unfortunately, existing heuristic algorithms do not consider tasks that require a large number of computational resources. Such tasks typically need to be offloaded to a large-scale computational resource such as a many-core system or graphics processing unit (GPU). Therefore, we propose a DAG scheduling algorithm that uses a cluster-based many-core architecture to offload such computation; here, we use Kalray MPPA-256 as the many-core architecture. Our algorithm divides many-core computational resources for such tasks. Comparing our results with existing algorithms, we succeeded in shortening makespan with many types of DAGs while also considering Amdahl's law. We also compared the deadline miss ratio and execution time of our algorithm to existing algorithms, and show that the proposed algorithm is superior to the existing algorithms for any evaluation metrics. Yuto Kitagawa, Tasuku Ishigooka, Takuya Azumi |
EUC | 3 |
| 2018 | Runtime Component Information on Embedded Component SystemsabstractTo efficiently build large embedded software systems, dividing the embedded software into components and subsystems and enhance reusability by converting these into distinct parts are essential. Component-Based Development (CBD) is used to reduce costs and improve development efficiency and can be applied to reusable software development. However, CBD systems generally do not support obtaining runtime component information about interfaces and variables generated at runtime, which makes testing for, verifying, and fixing bugs difficult. To deal with this issue, we propose the component framework for obtaining information about component-based systems. This makes it possible to obtain static information on generated components, including interfaces, and runtime component information, including component states and component variables. This makes it an invaluable tool for discovering mistakes and bugs. In addition, the proposed framework was designed with flexibility and efficiency in mind. Seito Shirata, Hiroshi Oyama 0002, Takuya Azumi |
EUC | 3 |
| 2018 | GPUhd: Augmenting YARN with GPU Resource ManagementabstractThis paper presents GPUhd, a graphics processing unit (GPU) resource management approach that combines Hadoop and a GPU to obtain scale-out and scale-up functionality. There are several researches that combine Hadoop and GPU. However, there are no researches that can schedule tasks in consideration of GPU resource on Hadoop. Moreover, these researches cannot use multiple distributed frameworks. GPUhd extends the Yet Another Resource Negotiator (YARN) management mechanism and distributed processing frameworks for the coordinated use of GPU resources in Hadoop. We extend the YARN scheduling algorithm to consider GPU resources and incorporate a resources monitoring function. GPU resources can be managed on the basis of existing development methods because GPUhd simply handles GPU resources as host memory and CPU resources. In addition, GPUhd achieves high-speed processing, e.g., the computational time required to calculate 2048 x 2048 matrix multiplication is approximately 25 times less than that required when using only a CPU with Hadoop. GPUhd achieves high scalability and excellent response times in a heterogeneous distributed environment. Daisuke Fukutomi, Yuki Iida, Takuya Azumi, Shinpei Kato, Nobuhiko Nishio |
HPC Asia | 3 |
| 2018 | Real-Time ROS Extension on Transparent CPU/GPU Coordination MechanismabstractRobot Operating System (ROS) promotes fault isolation, faster development, modularity, and core reusability and is therefore widely studied and used as the de facto standard for autonomous driving systems. Graphics processing units (GPUs) also facilitate high-performance computing and are therefore used for autonomous driving. As the requirements for real-time processing increase, methods for satisfying real-time constraints for ROS and GPUs are being developed. Unfortunately, scheduling algorithms specifying ROS's transportation (publish/subscribe) model, which can have execution order restrictions, are not being investigated, leading to the introduction of waiting time and degrading the responsiveness of the entire system. Furthermore, GPU tasks on ROS are also affected by the ROS transportation model, because central processing unit (CPU) time is occupied when GPU functions are launched. This paper proposes a loadable kernel module framework, called real-time ROS extension on transparent CPU/GPU coordination mechanism (ROSCH-G), for scheduling ROS in a heterogeneous environment without modifying the OS kernel and device drivers and then evaluates it experimentally. ROSCH-G provides a scheduling algorithm that considers ROS's execution order restrictions and a CPU/GPU coordination mechanism. Experimental results demonstrate that the proposed algorithm reduces the deadline miss rate and, compared with previous studies, makes effective use of the benefits of parallel processing. In addition, the results for the coordination mechanism demonstrate that ROSCH-G can schedule multiple GPU applications successfully. Yuhei Suzuki, Takuya Azumi, Shinpei Kato, Nobuhiko Nishio |
ISORC | 2 |
| 2018 | ROSCH: Real-Time Scheduling Framework for ROSabstractThis paper presents a real-time scheduling framework for the robot operating system (ROS) called ROSCH. ROS is an open-source software platform and a meta-operating system designed for robots. ROS provides extensive development libraries and has been widely used for autonomous driving systems. However, ROS does not guarantee real-time performance; hence, a ROS-based autonomous driving car could cause a traffic accident. Therefore, ROSCH comprises three functionalities that do not exist in the ROS to guarantee real-time performance: (1) a synchronization system, (2) fixed-priority based directed acyclic graph (DAG) scheduling framework, and (3) fail-safe function. In particular, the synchronization system guarantees that the timestamp gaps between sensor measurements will be less than or equal to the calculated value. The fixed-priority based DAG scheduling framework guarantees that end-to-end latency is less than or equal to an estimated value. Operating both mechanisms simultaneously guarantees the final output topic frequency. The fail-safe functionality provides a danger avoidance action at a deadline miss. In addition, these functionalities are designed to be compatible with legacy software. No previous method that guarantees real-time operation provides a comparable range of features while maintaining the compatibility with the existing software. Experimental results showed that the time differences between sensor timestamps remained below a calculated value. Moreover, correctly scheduled tasks incur an end-to-end latency that is within an estimated value. In addition, the throughput of the end node is achieved. Finally, the fail-safe functionality overhead is 45 μsec on an average, which is 0.004% of the average execution time for one node. Yukihiro Saito, Futoshi Sato, Takuya Azumi, Shinpei Kato, Nobuhiko Nishio |
RTCSA | 3 |
| 2017 | Exploring Scalable Data Allocation and Parallel Computing on NoC-Based Embedded Many CoresabstractIn embedded systems, high processing requirements and low power consumption need heterogeneous computing platforms. Considering embedded requirements, applications need to be designed based on scalable data allocation and parallel computing with non-uniform memory access (NUMA) many cores. In this paper, we use one of the embedded commercial off-the-shelf (COTS) multi/many-core components, the Massively Parallel Processor Arrays (MPPA) 256 developed by Kalray, and conduct evaluations of data transfer and parallelization of a practical application. We investigate currently achievable data transfer latencies between distributed memories on network-on-chip (NoC), memory access characteristics, and parallelization potential with many cores. Subsequently, we run a practical application, the core of the autonomous driving system, on many-core processors and acceleration by parallelization indicates practicality of many cores. By highlighting many-core computing capabilities, we explore the scalable data allocation and parallel computing on NoC-based embedded many cores. Yuya Maruyama, Shinpei Kato, Takuya Azumi |
ICCD | 3 |
| 2017 | Anomaly Prediction Based on k-Means Clustering for Memory-Constrained Embedded DevicesabstractThis paper proposes an anomaly prediction method based on k-means clustering that assumes embedded devices with memory constraints to predict control system anomalies. With this method, by checking control system behavior, it is possible to predict anomalies. However, continuing clustering is difficult because data accumulate in memory similar to existing k-means clustering method, which is problematic for embedded devices with low memory capacity. Therefore, we also propose k-means clustering to continue clustering for infinite stream data. The proposed k-means clustering method is based on online k-means clustering of sequential processing. The proposed k-means clustering method only stores data required for anomaly prediction and releases other data from memory. Experimental results show that anomalies can be predicted by k-means clustering, and the proposed method can predict anomalies similar to standard k-means clustering while reducing memory consumption. Moreover, the proposed k-means clustering demonstrates better results of anomaly prediction than existing online k-means clustering. Yuto Kitagawa, Tasuku Ishigooka, Takuya Azumi |
ICMLA | 3 |
| 2017 | Demo Abstract: Co-simulation Framework for Autonomous Driving Systems with MATLAB/SimulinkabstractAutonomous driving vehicles are currently being developed in many countries. Autonomous driving systems are developed using the Robot Operating System (ROS), which is suitable for the development of robotics and used for various systems of autonomous driving vehicles. However, in the automotive industry, these systems have often been designed using MATLAB/Simulink, can simulate and evaluate models created for autonomous driving. These models cannot be used with the systems based on ROS. To use a model created using MATLAB/Simulink in ROS, it is necessary to rewrite the model for ROS and incorporate it into the autonomous driving system, reducing development efficiency. Therefore, we propose an integrated development framework that can simulate and operate an autonomous driving system based on ROS with MATLAB/Simulink. The proposed framework improves the development efficiency because the model created by MATLAB/Simulink can be used in the systems without a separate incorporation step. In this demonstration, we perform a co-simulation with the autonomous driving system using the proposed framework. Shota Tokunaga, Takuya Azumi |
RTAS | 2 |
| 2017 | Scheduling parallel and distributed processing for automotive data stream management system
Jaeyong Rho, Takuya Azumi, Mayo Nakagawa, Kenya Sato, Nobuhiko Nishio |
J. Parallel Distributed Comput. | 2 |
| 2017 | Real-Time GPU Resource Management with Loadable Kernel ModulesabstractGraphics processing unit (GPU) programming environments have matured for general-purpose computing on GPUs. Significant challenges for GPUs include system software support for bounded response times and guaranteed throughput. In recent years, GPU technologies have been applied to real-time systems by extending the operating system modules to support real-time GPU resource management. Unfortunately, such a system extension makes it difficult to maintain the system with version updates because the OS kernel and device drivers must be modified at the source-code level, thereby preventing continuous research and development of GPU technologies for real-time systems. A loadable kernel module (LKM) framework, called Linux Real-Time eXtention with GPUs (Linux-RTXG), for managing real-time GPU resources with Linux without modifying the OS kernel and device drivers is proposed and evaluated experimentally. Linux-RTXG provides mechanisms for interrupt interception and independent synchronization to achieve real-time scheduling and resource reservation capabilities for GPU applications on top of existing device drivers and runtime libraries. Experimental results demonstrate that the overhead incurred by introducing the proposed Linux-RTXG is comparable to that of introducing existing kernel-dependent approaches. In addition, the results demonstrate that multiple GPU applications can be scheduled successfully by Linux-RTXG to meet their priority and quality-of-service requirements in real time. Yuhei Suzuki, Yusuke Fujii, Takuya Azumi, Nobuhiko Nishio, Shinpei Kato |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2016 | Exploring the performance of ROS2abstractMiddleware for robotics development must meet demanding requirements in real-time distributed embedded systems. The Robot Operating System (ROS), open-source middleware, has been widely used for robotics applications. However, the ROS is not suitable for real-time embedded systems because it does not satisfy real-time requirements and only runs on a few OSs. To address this problem, ROS1 will undergo a significant upgrade to ROS2 by utilizing the Data Distribution Service (DDS). DDS is suitable for real-time distributed embedded systems due to its various transport configurations (e.g., deadline and fault-tolerance) and scalability. ROS2 must convert data for DDS and abstract DDS from its users; however, this incurs additional overhead, which is examined in this study. Transport latencies between ROS2 nodes vary depending on the use cases, data size, configurations, and DDS vendors. We conduct proof of concept for DDS approach to ROS and arrange DDS characteristic and guidelines from various evaluations. By highlighting the DDS capabilities, we explore and evaluate the potential and constraints of DDS and ROS2. Yuya Maruyama, Shinpei Kato, Takuya Azumi |
EMSOFT | 3 |
| 2016 | Extended mapping algorithm based on modularity from synchronous block diagrams to AUTOSAR runnablesabstractModel-based development (MBD) has become important in the automobile domain. Automobile control systems consist of various software applications, and with MATLAB/Simulink, developers can design such applications using synchronous reactive models represented by synchronous block diagrams (SBD). The automotive open system architecture (AUTOSAR), a global development partnership formed to create open and standardized software architecture for automotive electronic control units (ECU), can provide highly reusable middleware. In this case, developers must map blocks of the SBD to AUTOSAR runnables, i.e., ECU processing units, and then assign the runnables to the ECUs. Most sample models are single-rate models. However, multi-rate control models will become essential due to the increasing complexity and scale of such automotive systems. This paper proposes top-down mapping algorithms from multi-rate control SBDs to runnables in consideration of schedulability, modularity, and code size. Note that proposed algorithms do not consider reusability. Evaluation results demonstrate that algorithms provide runnable sets with superior modularity than an existing algorithm. Shunsuke Hori, Takuya Azumi |
ETFA | 2 |
| 2016 | RTM-TECS: Collaboration Framework for Robot Technology Middleware and Embedded Component SystemabstractRobot technologies, such as robot technology middleware (RTM) that is a component-oriented platform, are popular. However, RTM does not ensure stable real-time processing in common object request broker architecture. In this paper, a collaboration framework of RTM and TOPPERS embedded component system (TECS) is proposed to address this problem. TECS, a system that satisfies real-time processing requirements, is employed to enhance real-time processing in the proposed framework. To implement the collaboration of RTM and TECS, we have adopted remote procedure call and one-way communication. In addition, extending a generator enables the generation of robot technology components from TECS components. We have evaluated the processor cycle counts of the proposed framework in comparison with those of a conventional method. In addition, we evaluated the execution time of serial communication and a motor application using the proposed framework. The evaluation results show that the proposed framework is functionally employed in a hard real-time system. Furthermore, we evaluated the amount of code generated by the proposed framework. The evaluation results reveal that the code generated by the proposed framework is reusable and can enhance productivity. Ryo Hasegawa, Naofumi Yawata, Noriaki Ando, Nobuhiko Nishio, Takuya Azumi |
ISORC | 5 |
| 2016 | GPUrpc: Exploring Transparent Access to Remote GPUs
Yuki Iida, Yusuke Fujii, Takuya Azumi, Nobuhiko Nishio, Shinpei Kato |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2015 | mruby on TECS: Component-Based Framework for Running Script ProgramabstractScripting languages are attractive for embedded system due to their high productivity. However, it is difficult to use scripting languages in a practical application because their performance and libraries for managing embedded devices are immature compared to the C programming language. This paper proposes a framework for effectively running an mruby script program on embedded systems based on the TOPPERS Embedded Component System (TECS). TECS generates glue code for invocation from mruby programs to legacy code in C language. It also supports configuration of mruby. Experimental results demonstrate the effectiveness of the proposed framework. Takuya Azumi, Yuki Nagahara, Hiroshi Oyama 0002, Nobuhiko Nishio |
ISORC | 1 |
| 2013 | VISA synthesis: Variation-aware Instruction Set Architecture synthesisabstractWe present VISA: a novel Variation-aware Instruction Set Architecture synthesis approach that makes effective use of process variation from both software and hardware points of view. To achieve an efficient speedup, VISA selects custom instructions based on statistical static timing analysis (SSTA) for aggressive clocking. Furthermore, with minimum performance overhead, VISA dynamically detects and corrects timing faults resulting from aggressive clocking of the underlying processor. This hybrid software/hardware approach generates significant speedup without degrading the yield. Our experimental results on commonly used ISA synthesis benchmarks demonstrate that VISA achieves significant performance improvement compared with a traditional deterministic worst case-based approach (up to 78.0%) and an existing SSTA-based approach (up to 49.4%). Yuko Hara-Azumi, Takuya Azumi, Nikil Dutt |
ASP-DAC | 2 |
| 2013 | Data Transfer Matters for GPU ComputingabstractGraphics processing units (GPUs) embrace many-core compute devices where massively parallel compute threads are offloaded from CPUs. This heterogeneous nature of GPU computing raises non-trivial data transfer problems especially against latency-critical real-time systems. However even the basic characteristics of data transfers associated with GPU computing are not well studied in the literature. In this paper, we investigate and characterize currently-achievable data transfer methods of cutting-edge GPU technology. We implement these methods using open-source software to compare their performance and latency for real-world systems. Our experimental results show that the hardware-assisted direct memory access (DMA) and the I/O read-and-write access methods are usually the most effective, while on-chip micro controllers inside the GPU are useful in terms of reducing the data transfer latency for concurrent multiple data streams. We also disclose that CPU priorities can protect the performance of GPU data transfers. Yusuke Fujii, Takuya Azumi, Nobuhiko Nishio, Shinpei Kato, Masato Edahiro |
ICPADS | 2 |
| 2013 | HR-TECS: Component technology for embedded systems with memory protectionabstractA software partitioning has been used to develop safety-critical systems in recent years. In addition, software component technologies supporting a software partitioning have been developed. This paper describes the new component technology for embedded software that requires memory protection, which is one of the important features for the partitioning. HR-TECS is a new component technology based on the real-time operating system supporting the static memory layout. Developers can easily allocate components to partitions in order to protect memory areas. In addition, HR-TECS supports inter-partition communications so that developers can implement components without consideration for inter-partition communications. The results of evaluation demonstrate the effectiveness of HR-TECS. Takuya Ishikawa, Takuya Azumi, Hiroshi Oyama 0002, Hiroaki Takada |
ISORC | 2 |
| 2012 | Enhancement of Real-Time Processing by Cooperation of RTM and TECSabstractRecently, RTM (Robot Technology Middleware) is attracting attention as a component oriented platform for robot development. However, RTM is unable to ensure real-time processing in CORBA because CORBA manages packets in a FIFO manager. In this paper, we propose a communication method from RTM to TECS in an effort to enhance realtime processing. TECS is a component system for embedded systems and suitable for real-time systems. We remove a part of real-time processing in RTM. Moreover, TECS is added to enhance real-time processing because TECS support a real-time processing requirements. In addition, it is possible to generate the components to communicate from RTM to TECS by using a plug-in. In the evaluation, cycle counts between RTM and TECS are compared with those between RTM and RTM. Furthermore, the amount of codes which are generated code with the purpose method and written code by developers are compared. Naofumi Yawata, Takuya Azumi, Nobuhiko Nishio |
RTCSA | 2 |
| 2010 | Wheeled Inverted Pendulum with Embedded Component System: A Case StudyabstractSoftware component techniques have been widely used for enhancement and the cost reduction of software development. We herein introduce a component system with a real-time operating system (RTOS). A case study of a two-wheeled inverted pendulum balancing robot with the component system is presented. The component system can deal with RTOS resources, such as tasks and semaphores, as components. Moreover, a trace functionality which is a new functionality to confirm the state of components or calling components without modification of C source code is introduced. Takuya Azumi, Hiroaki Takada, Takayuki Ukai, Hiroshi Oyama 0002 |
ISORC | 1 |
| 2007 | A New Specification of Software Components for Embedded SystemsabstractIn the last decade, the size and complexity of the software in embedded systems have increased. The present study attempts to decrease the complexity and difficulty of software development in embedded systems. We herein introduce a new component system that is suitable for embedded systems. It is possible to estimate the memory consumption of an entire application since the proposed system adopts a static configuration. In addition, this system takes into account to be used in several domains of embedded systems because several particle sizes of component are supported. Moreover, the concept of the component for a distributed application is presented Takuya Azumi, Masanari Yamamoto, Yasuo Kominami, Nobuhisa Takagi, Hiroshi Oyama 0002, Hiroaki Takada |
ISORC | 1 |