EDBT 2026 Demo / reviewers in the wild / expert
Hiroyuki Tomiyama
dblp:68/1314
· DBLP profile ↗
45ranked-venue papers
4as first author
8since 2021 · last 2025
0000-0003-1655-7877ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 24 · 4 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 3 since 2021Software engineering, systems software and programming languages · 6 · 1 first-authorHuman-computer interaction and ubiquitous computing · 3 · 3 since 2021Computer networks · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Security and privacy · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
Embedded and real-time systems · 99% Processor architecture and microarchitecture · 1% | |
| Theoretical computer science
1 paper |
Algorithms and data structures · 100% | |
| Software engineering, system software, and programming languages
1 paper |
Compilers and program optimization · 100% |
Topics — the 3 heaviest of 5, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Algorithms and data structures
dynamic programming |
0.1 | 1 | 2019 | Work-in-Progress: Routing of Delivery Drones with Load-Dependent Flight Speed · RTSS 2019 |
Compilers and program optimization › compiler construction
compiler generation |
0.0 | 1 | 1997 | Memory-CPU Size Optimization for Embedded System Designs · DAC 1997 |
Embedded and real-time systems
embedded system design |
0.0 | 1 | 1997 | Memory-CPU Size Optimization for Embedded System Designs · DAC 1997 |
Methods — techniques the papers use, named apart from their topics
dynamic programming · 0.8system cost minimization · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Poster: A Comparative Analysis of Machine Learning Models for SAT Runtime Prediction
Tomohisa Kawakami, Tomoyasu Shimada, Xiangbo Kong, Hiroyuki Tomiyama, Shigeru Yamashita |
RTCSA | 4 |
| 2025 | YOLO-EMR: Efficient Multi-Scale and Rotated Object Detection in UAV Aerial ImageryabstractWith the rapid development of UAVs, object detection in aerial imagery has become an important techniques. However, deploying real-time detection models on UAV platforms remains highly challenging due to limited computational costs, as well as the multi-scale objects and dense distribution caused by high-altitude imaging. Numerous studies have contributed valuable insights to address these challenges, yet opportunities remain for improving the balance between model efficiency and detection accuracy. Moreover, current research mainly focuses on small object detection, without fully considering the detection requirements for multi-scale and rotated objects in high-altitude imagery. To address these issues, this paper proposes a lightweight object detection model specifically designed for multi-scale and rotated object detection in high-altitude imagery. Furthermore, to better evaluate the model’s performance in UAV-based object detection in high-altitude imagery, this work uses VisDrone 2019 dataset to assess the model’s real-world performance. As a results, compared to existing object detection approaches, the proposed model achieves a strong balance between detection accuracy, model efficiency, and inference speed, reducing parameters by approximately 31% and improving inference speed by 25%, while maintaining the same detection accuracy as the best baseline methods. Haimin Yan, Xiangbo Kong, Tomoyasu Shimada, Hiroyuki Tomiyama |
SMC | 5 |
| 2024 | YOLO-UTS: Lightweight YOLOv5 for UAV Traffic Monitoring and SurveillanceabstractIn the application of drone-based target detection, both the lightweighting of the model and its accuracy are crucial. Therefore, this paper aims to achieve model lightweighting and accuracy improvement through structural improvements to the YOLO model. To meet this objective, this work optimizes the structure of the existing model to address the challenge of operational limitations in small drones, which have restricted computational capacities. By integrating lightweight modules as the backbone of the model, it enhances the feasibility of deployment on devices with limited processing power. Furthermore, the introduction of a self-attention mechanism, which is placed in the Neck of the model improves the ability to prioritize critical regions within the image. This enhancement is crucial for accurately detecting overlapping or blurred objects encountered by the moving drone. Additionally, the modification of the Intersection over Union (IoU) metric, which now considers the aspect ratio and shape of the bounding boxes further refines the target detection capabilities, ensuring more precise and reliable object localization. Since drones do not remain at a constant height or position, the angles and distances of the vehicles which they capture are bound to change. This adjustment, which modifies how the IoU quantifies overlap, allows the IoU to more accurately quantify the degree of overlap of vehicle detection bounding boxes at different angles and distances, thereby providing more reliable target detection performance. Experimental results indicate that compared to existing YOLO models, our method achieves 30% reduction in model size and 1% improvement in accuracy. Haimin Yan, Xiangbo Kong, Hiroyuki Tomiyama |
SMC | 4 |
| 2024 | YOLO-ELD: Efficient and Lightweight Detection for UAV Aerial ImageryabstractObject detection in UAV imagery has become a hot topic in recent years. However, deploying real-time detection models on UAV platforms is highly challenging due to limited computational power and memory. Moreover, the large size of UAV-captured images, the small size of objects, and their dense distribution all impact detection efficiency. Many researchers have made a series of improvements to address these issues, but they have not maintained a good balance among model size, inference speed, and accuracy. To address above difficulties, this paper proposes an efficient and lightweight model that maintains moderate detection accuracy while achieving model lightweighting and reduced inference time. Concretely, considering the higher resolution of drone-captured images, we have designed a backbone with more lightweight downsampling modules, enhancing deployment efficiency on devices with limited resources. Additionally, this work incorporates a self-attention mechanism in the feature extraction component of the model, which significantly improves the ability to process critical areas in the image, crucial for detecting small-scale and densely distributed targets. Moreover, this work designs an IoU tailored for drone aerial images, which calculates losses by focusing on the shape and scale of bounding boxes, thereby enhancing the accuracy of bounding box regression. Additionally, it uses a ratio of scale factors to control the generation of auxiliary bounding boxes, which aids in loss calculation and accelerates convergence. Experimental results on the VisDrone 2019 dataset show that compared to existing detection methods used for drones, our model is more lightweight and efficient, while also achieving medium accuracy. Also, compared to our baseline method, YOLO-ELD reduces the number of model parameters by about 40%, increases the inference speed of the model by 10%, and also improves the precision of model by 3%. Haimin Yan, Xiangbo Kong, Hiroyuki Tomiyama |
SMC | 4 |
| 2024 | Table Tennis Stroke Classification from Game Videos Using 3D Human KeypointsabstractIn this paper, we propose a classification method for table tennis strokes using player’s 3D joint coordinates. In existing studies on stroke classification, classification is performed based on videos taken by a camera installed on a table tennis table or data obtained by attaching an inertial sensor to a player. However, in these existing methods, sensors and cameras interfere with the game, and it is difficult to adapt to the game videos. Therefore, in this paper, we classify strokes from videos that can be taken during games under more practical conditions. For the classification, we use the player’s 3D joint coordinates as input to classify the game video. In our method, we use deep learning to learn two kinds of information, the recorded video and the player’s joint coordinates obtained by 3D pose estimation and perform classification. As a result of the experiment, in the classification of the rally video dataset, the accuracy is improved by 4.2~15.5% in the validation data, which is the stroke of the learned player, and by 1.8~15.0% in the test data, which is the stroke of the unlearned player, compared with the existing method which input the videos and 2D joint coordinates. Yuta Fujihara, Xiangbo Kong, Ami Tanaka, Hiroki Nishikawa, Hiroyuki Tomiyama |
VCIP | 5 |
| 2024 | Dynamic Point-Pixel Feature Alignment for Multimodal 3-D Object DetectionabstractDetection of small or distant objects is a major challenge in 3-D object detection in autonomous driving either through RGB images or LiDAR point clouds. Despite the growing popularity of sensor fusion in this task, existing fusion methods have not adequately taken into account the challenges associated with 3-D small object detection, such as semantic misalignment of small objects, caused by occlusion and calibration errors. To address this issue, we propose dynamic point-pixel feature alignment network (DPPFA-Net) for multimodal 3-D small object detection by introducing memory-based point-pixel fusion (MPPF) modules, deformable point-pixel fusion (DPPF) modules, and semantic alignment evaluator (SAE) modules. More concretely, the proposed MPPF module automatically performs intramodal and cross-modal feature interactions. The intramodal interaction reduces sensitivity to noise points, while the explicit cross-modal feature interaction based on the memory bank facilitates easier network learning and enables a more comprehensive and discriminative feature representation. The DPPF module establishes interactions exclusively with key position pixels based on a sampling strategy. This design not only guarantees a low-computational complexity but also enables adaptive fusion functionality, especially beneficial for high-resolution images. The SAE module guarantees semantic alignment of the fused features, thereby enhancing the robustness and reliability of the fusion process. Furthermore, we construct a simulated multimodal noise data set, which enables quantitative analysis of the robustness of multimodal methods under varying degrees of multimodal noise. Extensive experiments on the KITTI benchmark and challenging multimodal noisy cases show that DPPFA-Net achieves a new state-of-the-art, highlighting its effectiveness in detecting small objects. Our proposed method is compared to the first place on the KITTI leaderboard and achieves better performance by 2.07%, 6.52%, 7.18%, and 6.22% of the average precision on the varying degrees of multimodal noise cases. Xiangbo Kong, Hiroki Nishikawa, Qiuyou Lian, Hiroyuki Tomiyama |
IEEE Internet Things J. | 5 |
| 2023 | Message from the Chairs: RTCSA 2023abstractIt is our pleasure to welcome you to the 29th IEEE International Conference on Embedded and Real-Time Computing Systems and Applications (RTCSA 2023), held in Niigata, Japan. This year, we are honored that RTCSA is sponsored by IEEE, IEEE Computer Society, and IEEE Technical Committee on Real-Time Systems (TCRTS). The objective of the conference is to bring together researchers and developers from academia and industry for advancing the technology of embedded and real-time systems and their emerging applications, including the Internet of Things (IoT) and Cyber-Physical Systems (CPS). Hiroyuki Tomiyama, Nan Guan, Sebastian Steinhorst |
RTCSA | 3 |
| 2021 | A Comprehensive Analysis of Low-Impact Computations in Deep Learning WorkloadsabstractDeep Neural Networks (DNNs) have achieved great successes in various machine learning tasks involving a wide range of domains. Though there are multiple hardware platforms available, such as GPUs, CPUs, FPGAs, and etc, CPUs are still preferred choices for machine learning applications, especially in low-power and resource-constrained computation environments such as embedded systems. However, the power and performance efficiency become critical issues in such computation environments when applying DNN techniques. An attractive optimization to DNNs is to remove redundant computations to enhance the execution efficiency. To this end, this paper conducts extensive experiments and analyses on popular state-of-the-art deep learning models. The experimental results include the numbers of instructions, branches, branch prediction misses, cache misses, and etc, during the execution of the models. Besides, we also investigate the performance and sparsity of each layer in the models. Based on the analysis results, this paper also proposes an instruction-level optimization, which achieves the performance improvement ranging from 10.26% to 28.0% for certain convolution layers. Zhichen Wang, Xuebin Yue, Wenwen Wang 0001, Hiroyuki Tomiyama, Lin Meng 0001 |
ACM Great Lakes Symposium on VLSI | 5 |
| 2020 | Frame Detection and Text Line Segmentation for Early Japanese Books Understanding
Bing Lyu, Hiroyuki Tomiyama, Lin Meng 0001 |
ICPRAM | 2 |
| 2020 | Scheduling of moldable fork-join tasks with inter- and intra-task communicationsabstractThis paper proposes scheduling techniques for moldable fork-join tasks on multicore architecture. The proposed techniques decide the number of cores and execution start time for each task during scheduling and mapping, with taking into account inter- and intra-task communications. The proposed techniques based on integer programming formulation aim at minimization of the overall schedule length. Experimental results are compared with the state-of-the-art techniques. Hiroki Nishikawa, Kana Shimada, Ittetsu Taniguchi, Hiroyuki Tomiyama |
SCOPES | 4 |
| 2019 | Work-in-Progress: Routing of Delivery Drones with Load-Dependent Flight SpeedabstractDrones draws increasing attention as vehicles for home delivery services. Delivery time is one of the most critical concerns of both customers and delivery service providers. The delivery time depends not only on the flight distance but also on the flight speed, and the flight speed depends on the payload. This paper studies a routing problem for delivery drones considering load-dependent flight speed. This paper formally defines Flight Speed-aware Vehicle Routing Problem (FSVRP) and proposes a dynamic programming algorithm to efficiently solve the problem. Experiments show the effectiveness of the proposed algorithm in terms of quality of results and algorithm runtime. Yusuke Funabashi, Ittetsu Taniguchi, Hiroyuki Tomiyama |
RTSS | 3 |
| 2019 | QoE-Constrained Concurrent Request Optimization Through Collaboration of Edge ServersabstractCloud computing, which is claimed to provide plentiful storage, computational, and other resources, has become a promising platform to support resource-intensive applications. Due to the wide adoption of smart things to support domain applications and considering the delay-sensitivity of certain requests and limited network capacity compared with huge data packets to be transmitted, the quality of experience (QoE) may be hard to be satisfied when requests are solely supported by cloud computing. In this setting, edge computing has become an infrastructure to facilitate request satisfaction at the network edge. This article proposes a mechanism to optimize the collaboration of heterogeneous edge servers with certain QoE constraints. Specifically, concurrent requests, which are usually represented in terms of SQL queries, are rewritten as atomic queries, and these atomic queries are optimally assigned to edge servers through adopting an algorithm inspired by the minimum spanning tree, where QoE factors, including the delay, size of data packets, and number of operators, are considered. Evaluation results indicate that the proposed mechanism can effectively improve the QoE of requests compared with the state-of-the-art's mechanisms. Yaqiang Zhang, Lin Meng 0001, Xiao Xue 0001, Zhangbing Zhou, Hiroyuki Tomiyama |
IEEE Internet Things J. | 5 |
| 2018 | Communication-aware scheduling of data-parallel tasks: work-in-progress
Kana Shimada, Ittetsu Taniguchi, Hiroyuki Tomiyama |
CASES | 3 |
| 2018 | Synthesis of Full Hardware Implementation of RTOS-Based SystemsabstractThis paper presents a method of automatically synthesizing a hardware design from a set of source codes for a real-time system utilizing an RTOS. It generates a full hardware implementation where all the tasks and handlers in the system as well as all the necessary services provided by the RTOS kernel are implemented as hardware. Every task and handler is synthesized into an independent hardware module so that it may run in parallel with the other tasks/handlers as soon as it is ready. This leads to task switching with extremely low overhead and reduced computation time both by parallel and hardware execution. Moreover, this eliminates the necessity of the task queue management; task scheduling is realized by a relatively simple manager hardware which instructs each task/handler to run or stall based on the values of its status variables. Since most of the API calls from tasks/handlers are reduced to reads/writes of these status variables, they can be expanded inline into the tasks/handlers source codes which are compiled into hardware designs by a high-level synthesizer. We have implemented a prototype synthesis system which assume the use of the TOPPERS/ASP3 real-time kernel. A hardware implementation synthesized from a sample 1.c code, bundled in the TOPPERS/ASP3 release, took 23 cycles for waking up a waiting task and only 1 cycle for activating an interrupt handler. Yuuki Oosako, Nagisa Ishiura, Hiroyuki Tomiyama, Hiroyuki Kanbara |
RSP | 3 |
| 2018 | Scheduling of Malleable Tasks Based on Constraint ProgrammingabstractThis paper proposes a scheduling method for malleable tasks based on constraint programming (CP). For a given task-graph, the proposed method decides the execution order of tasks and the number of cores to execute each task simultaneously in such a way that the overall schedule length is minimized. Experimental results show that our CP-based scheduling method could find better schedules than the state-of-the-art method which is based on integer linear programming. Hiroki Nishikawa, Kana Shimada, Ittetsu Taniguchi, Hiroyuki Tomiyama |
TENCON | 4 |
| 2017 | Binary synthesis implementing external interrupt handler as independent moduleabstractThis article presents a method of synthesizing hardware from a given executable binary code with an external interrupt handler, where the normal flow and the interrupt handling are executed by separate hardware modules. Our previous method synthesized the whole program into a single hardware module, in which register save/restore imposed limitations on the timing to start interrupt handling and also impaired efficiency of the synthesized hardware. By executing the two tasks on separate modules, register save/restore can be eliminated, which allows interrupt handler to start at arbitrary timing and reduces the response time and cost of the hardware. By allowing two processes to run in parallel, total execution time is also reduced. An experiment with a simple program has shown that the execution cycles and the delay were reduced by about 80% and 20%, respectively, as compared with MIPS CPU. A motor controller driven by periodical interrupts from a timer has been successfully synthesized from C and assembly programs, which runs more than 20 times faster than the MIPS CPU. Naoya Ito, Yuuki Oosako, Nagisa Ishiura, Hiroyuki Kanbara, Hiroyuki Tomiyama |
RSP | 5 |
| 2016 | Energy-aware task migration for multiprocessor real-time systems
Yutaka Matsubara, Hiroyuki Tomiyama, Hiroaki Takada |
Future Gener. Comput. Syst. | 3 |
| 2015 | Profiling-driven multi-cycling in FPGA high-level synthesis
Stefan Hadjis, Andrew Canis, Ryoya Sobue, Yuko Hara-Azumi, Hiroyuki Tomiyama, Jason Helge Anderson |
DATE | 5 |
| 2014 | Cache Simulation for Instruction Set Simulator QEMUabstractIn embedded system design, there is an increasing demand for modeling techniques that can provide both accurate measurements of delay and fast simulation speed. Modeling latency effects of a cache can greatly increase accuracy of the simulation and assist developers to optimize their software. Current solutions have not succeeded in balancing three important factors: speed, accuracy and usability. In this research, we created a cache simulation module inside a well-known instruction set simulator QEMU. Our implementation can simulate various cases of cache configuration and obtain every memory access. In full system simulation, speed is kept at around 73 MIPS on a personal host computer which is close to native execution of ARM Cortex-M3(125 MIPS at 100 MHz). Compared to the widely used cache simulation tool, Valgrind, our simulator is three time faster. Tran Van Dung, Ittetsu Taniguchi, Hiroyuki Tomiyama |
DASC | 3 |
| 2014 | Fast Design-Space Exploration Method for SW/HW Codesign on FPGAs
Yuki Ando, Seiya Shibata, Shinya Honda, Hiroyuki Tomiyama, Hiroaki Takada |
FCCM | 4 |
| 2013 | SMYLE OpenCL: A programming framework for embedded many-core SoCsabstractEmbedded SoC architecture has shifted from single-core to multi/many-core paradigm because of better power/performance efficiency. In order to exploit the potential power/performance efficiency of the many-core architecture, a parallel computing framework is necessary. OpenCL is one of the most popular parallel computing frameworks in the field of general-purpose computing on GPUs and multicore servers. However, the existing OpenCL implementations are not suitable to embedded real-time systems because of the large runtime overhead. In this paper, we describe a lightweight OpenCL framework for embedded multi/many-core SoCs. Our OpenCL framework minimizes the runtime overhead by statically creating threads and mapping them onto cores. Preliminary experiments on an FPGA prototype board with a five-core architecture shows a significant reduction in runtime overhead compared with an existing OpenCL framework. Hiroyuki Tomiyama, Takuji Hieda, Naoki Nishiyama, Noriko Etani, Ittetsu Taniguchi |
ASP-DAC | 1 |
| 2012 | Clock-constrained simultaneous allocation and binding for multiplexer optimization in high-level synthesisabstractThis paper proposes a novel simultaneous allocation and binding method in high-level synthesis, which minimizes the circuit area including multiplexers (MUXs) under a clock constraint. Most existing works on binding minimize MUXs under given allocation by minimizing the number of interconnections, but do not care where the MUXs would be inserted in a circuit. As a result, they cannot guarantee the required clock frequency and often violate the clock constraint. On the contrary, our work globally optimizes binding and allocation for FUs and registers while meeting the clock constraint by considering where MUXs would be inserted. Our work is formulated as an ILP problem. Also, an effective ILP-based heuristic for non-small designs is presented. Experimental results demonstrate that our work satisfies the clock constraint with the minimum circuit area. Yuko Hara-Azumi, Hiroyuki Tomiyama |
ASP-DAC | 2 |
| 2011 | An Energy Aware Design Space Exploration for VLIW AGU Model with Fine Grained Power GatingabstractReducing energy consumption is a crucial for the embedded system design, and especially the leakage energy reduction is now big problem for the low power design. In order to reduce the leakage energy at standby time, power gating scheme is well known as a promising technique to realize partial power shutdown. However, the power gating usually causes penalties for shutdown and wakeup time, and this brings tradeoff between leakage energy reduction and latency penalty. This paper proposes energy aware design space exploration for power gated VLIW AGU model with fine grained power management. Contribution of this paper is an energy aware design space exploration with fast scheduling exploration for power gated VLIW AGU model. Experimental results show that proposed method can realize low power scheduling considering fine grained power management, and proposed architecture exploration method enables optimal design space exploration in practical time. Ittetsu Taniguchi, Mitsuya Uchida, Hiroyuki Tomiyama, Masahiro Fukui, Praveen Raghavan, Francky Catthoor |
DSD | 3 |
| 2011 | An integrated optimization framework for reducing the energy consumption of embedded real-time applications
Hideki Takase, Lovic Gauthier, Hirotaka Kawashima, Noritoshi Atsumi, Tomohiro Tatematsu, Yoshitake Kobayashi, Shunitsu Kohara, Takenori Koshiro, Tohru Ishihara, Hiroyuki Tomiyama, Hiroaki Takada |
ISLPED | 11 |
| 2010 | Minimizing inter-task interferences in scratch-pad memory usage for reducing the energy consumption of multi-task systemsabstractThis paper presents a new technique for reducing the energy consumption of a multi-task system by sharing its scratchpad memory (SPM) space among the tasks. With this technique, tasks can interfere by using common areas of the SPM. However, this requires to update these areas during context switches, which involves considerable overheads. Hence, an integer linear programming formulation is used at compile time for finding the best assignment of memory objects to the SPM and their respective locations inside it. Experiments show that the technique achieves up to 85% energy reduction with 8Kb of SPM and surpasses other sharing approaches. Lovic Gauthier, Tohru Ishihara, Hideki Takase, Hiroyuki Tomiyama, Hiroaki Takada |
CASES | 4 |
| 2010 | Partitioning and allocation of scratch-pad memory for priority-based preemptive multi-task systemsabstractScratch-pad memory has been employed as a partial or entire replacement for cache memory due to its better energy efficiency. In this paper, we propose scratch-pad memory management techniques for priority-based preemptive multi-task systems. Our techniques are applicable to a real-time environment. The three methods which we propose, i.e., spatial, temporal, and hybrid methods, bring about effective usage of the scratch-pad memory space, and achieve energy reduction in the instruction memory subsystems. We formulate each method as an integer programming problem that simultaneously determines (1) partitioning of scratch-pad memory space for the tasks, and (2) allocation of program code to scratch-pad memory space for each task. It is remarkable that periods and priorities of tasks are considered in the formulas. Additionally, we implement an RTOS-hardware cooperative support mechanism for a runtime code allocation to the scratch-pad memory space. We have made the experiments with the fully functional real-time operating system. The experimental results with four task sets have demonstrated the effectiveness of our techniques. Up to 73% energy reduction compared to a standard method was achieved. Hideki Takase, Hiroyuki Tomiyama, Hiroaki Takada |
DATE | 2 |
| 2010 | A Novel Mechanism for Effective Hardware Task Preemption in Dynamically Reconfigurable SystemsabstractExtending the idea of preemptive multitasking to DPRS (Dynamic Partial Reconfiguration Systems) has far-reaching implications as many mechanisms supporting the concept, such as context saving and restoring, have to be built practically from scratch. This paper addresses previously neglected issues, related to design of effective preemption mechanisms for Flip-Flop-based and RAM-based hardware tasks. Furthermore, a very efficient and complete solution to hardware task preemption for Virtex4-based DPRS is presented featuring in bitstream manipulation tool intended for PC and embedded system infrastructure with a DMA-based, instruction-driven reconfiguration/readback controller. Taking advantage of the developed lightweight bus, enhancing management of reconfigurable hardware modules, controller takes care of all essential hardware aspects related to context-switching thereby reducing CPU utilization to necessary minimum. Krzysztof Jozwik, Hiroyuki Tomiyama, Shinya Honda, Hiroaki Takada |
FPL | 2 |
| 2010 | Automatic communication synthesis with hardware sharing for design space explorationabstractIn this paper, we present a hardware sharing method for design space exploration of multiprocessor embedded systems. In our prior work, we had developed a system-level design tool which automatically synthesizes communications among the processes. In this work, we have extended our tool so that the tool can automatically synthesize communications which realize sharing of hardware among different processes. With the tool, designers only need to change the mapping information for hardware sharing. Designers therefore can easily explore wider design space with hardware sharing. A case study shows the effectiveness of our hardware sharing method. Yuki Ando, Seiya Shibata, Shinya Honda, Hiroyuki Tomiyama, Hiroaki Takada |
ISCAS | 4 |
| 2009 | Analyzing and optimizing energy efficiency of algorithms on DVS systems a first step towards algorithmic energy minimizationabstractThe energy efficiency at the algorithmic level on DVS systems and its analysis and optimization methods are presented. Given a problem the most energy efficient algorithm is not uniquely determined but dependend on multiple factors, including intratask dynamic voltage scaling (IntraDVS) policies, the size of intermediate data structure, and the size of inputs. We show that at the algorithmic level principles behind energy optimization and performance optimization are not identical. We propose a metric for evaluating optimal energy efficiency of static voltage scaling (SVS) and a few new effective IntraDVS policies employing data flow information. Experimental results on sorting algorithms show the existence of several tradeoffs in terms of energy consumption. Transforming algorithms by employing problem specific knowledge and data flow information successfully improves their energy efficiency. Tetsuo Yokoyama, Hiroyuki Tomiyama, Hiroaki Takada |
ASP-DAC | 3 |
| 2009 | Practical Energy-Aware Scheduling for Real-Time Multiprocessor SystemsabstractEnergy-aware real-time multiprocessor scheduling has been studied extensively so far. However, some of the constraints associated with the practical DVS applications have been ignored for simplicity. These constraints include discrete speed, idle power, inefficient speed, and application-specific power characteristics etc. This work targets energy-aware scheduling of periodic real-time tasks on the DVS-equipped multiprocessor systems with practical constraints. An adaptive minimal bound first-fit (AMBFF) algorithm with consideration of these realistic constraints is proposed for both dynamic-priority and fixed-priority multiprocessor scheduling. Simulation results on three commercial processor models show that our algorithm can save significantly more energy than existing algorithms. Tetsuo Yokoyama, Hiroyuki Tomiyama, Hiroaki Takada |
RTCSA | 3 |
| 2008 | A Generalized Framework for System-Wide Energy Savings in Hard Real-Time Embedded SystemsabstractA generalized dynamic energy performance scaling (DEPS) framework is proposed for exploring application-specific energy-saving potential in hard real-time embedded systems. This software-centric framework focuses on system-wide energy reduction and takes advantage of possible power control mechanisms to trade off performance for energy savings. Three existing technologies, i.e., dynamic hardware resource configuration (DHRC), dynamic voltage frequency scaling (DVFS), and dynamic power management (DPM) have been employed in this framework to achieve the maximal energy savings. Static and dynamic schemes of DEPS are proposed to deal with stable or variable workload in the embedded systems. Through a case study, its effectiveness has been validated. Hiroyuki Tomiyama, Hiroaki Takada, Tohru Ishihara |
EUC (1) | 2 |
| 2008 | CHStone: A benchmark program suite for practical C-based high-level synthesisabstractIn general, standard benchmark suites are critically important for researchers to quantitatively evaluate their new ideas and algorithms. This paper presents CHStone, a suite of benchmark programs for C-based high-level synthesis. CHStone consists of a dozen of large, easy-to-use programs written in C, which are selected from various application domains. This paper also presents synthesis results which will be served as a baseline for researchers to compare their new techniques with. In addition, we present a case study on function-level transformation using a program in the CHStone suite. Yuko Hara-Azumi, Hiroyuki Tomiyama, Shinya Honda, Hiroaki Takada, Katsuya Ishii |
ISCAS | 2 |
| 2007 | RTOS and Codesign Toolkit for Multiprocessor Systems-on-ChipabstractMultiprocessor designs have become popular in embedded domains for achieving the power and performance requirements. In this paper, we present principles and techniques for design and implementation of RTOS for embedded multiprocessor systems. We also present a system-level design toolkit for rapid design and evaluation of embedded multiprocessor systems. Shinya Honda, Hiroyuki Tomiyama, Hiroaki Takada |
ASP-DAC | 2 |
| 2007 | A Software Framework for Energy and Performance Tradeoff in Fixed-Priority Hard Real-Time Embedded Systems
Hiroyuki Tomiyama, Hiroaki Takada |
EUC | 2 |
| 2007 | Complexity-constrainted partitioning of sequential programs for efficient behavioral synthesisabstractThis paper proposes a behavioral level partitioning method for efficient behavioral synthesis from a large sequential program consisting of a set of functions. Our method optimally determines functions to be inlined into the main module and ones to be synthesized into sub modules in such a way that the overall datapath is minimized while the complexity of individual modules is lower than a certain level. The partitioning problem is formulated as an integer programming problem. Experimental results show the effectiveness of the proposed method. Yuko Hara-Azumi, Hiroyuki Tomiyama, Shinya Honda, Hiroaki Takada, Katsuya Ishii |
ACM Great Lakes Symposium on VLSI | 2 |
| 2007 | Scheduling Algorithms for I/O Blockings with a Multi-frame Task ModelabstractA task that suspends itself to wait for an I/O completion or to wait for an event from another node in distributed environments is called an I/O blocking task. In conventional hard real-time scheduling theories, there exist several approaches to schedule such I/O blocking tasks within the conventional framework of rate monotonic analysis (RMA). However, most of them are pessimistic. In this paper, we propose effective algorithms that can schedule a task set which includes I/O blocking tasks under dynamic priority assignment. We present a new critical instant theorem for multi-frame task set under dynamic priority assignment. The schedulability is analyzed under the new critical instant theorem. For the schedulability analysis , this paper presents saturation summation which is used to calculate maximum interference function (MIF). With the saturation summation, the schedulability of a task set including I/O blocking tasks can be analyzed more accurately. We propose an algorithm which is based on a frame laxity monotonic scheduling (FLMS). Genetic algorithm is also applied. From our experiments, we can conclude that the FLMS can significantly reduce the time of the calculation time, and GA can improve task schedulability ratio than the FLMS. Shan Ding, Hiroyuki Tomiyama, Hiroaki Takada |
RTCSA | 2 |
| 2006 | Function Call Optimization in Behavioral SynthesisabstractBehavioral synthesis, which automatically synthesizes an RTL circuit from a sequential program, is one of promising technologies to improve the design productivity. However, behavioral synthesis has not become popular yet in industry since the quality of generated circuits is not satisfactory, especially in the synthesis from the large programs with a number of functions. This paper proposes a method to optimize function calls in behavioral synthesis. We formulate the optimization problem using integer linear programming. Our experimental results show that our method reduces the circuit area by 44.6%, compared with a traditional method Yuko Hara-Azumi, Hiroyuki Tomiyama, Shinya Honda, Hiroaki Takada |
DSD | 2 |
| 2005 | A GA-based scheduling method for FlexRay systemsabstractAn advanced communication system, the FlexRay system, has been developed for future automotive applications. It consists of time-triggered clusters, such as drive-by-wire in cars, in order to meet different requirements and constraints between various sensors, processors, and actuators. In this paper, an approach to static scheduling for FlexRay systems is proposed. Our experimental results show that the proposed scheduling method significantly reduces up to 36.3% of the network traffic compared with a past approach. Shan Ding, Naohiko Murakami, Hiroyuki Tomiyama, Hiroaki Takada |
EMSOFT | 3 |
| 2005 | An Efficient Search Algorithm of Worst-Case Cache Flush TimingsabstractIn recent years, the use of cache memory has been desired in hard real-time systems in order to reduce the memory access time. To enable it, accurate analysis of the worst-case execution time considering cache flushes is necessary since the cache may be flushed by preempting tasks in a multitask environment. This paper proposes a method to find the worst-case timing of cache flushes and demonstrates its effectiveness. Hiroshi Miyamoto, Shinichi Iiyama, Hiroyuki Tomiyama, Hiroaki Takada, Hiroshi Nakashima |
RTCSA | 3 |
| 2002 | Automatic Verification of In-Order Execution In Microprocessors with Fragmented Pipelines and Multicycle Functional UnitsabstractAs embedded systems continue to face increasingly higher performance requirements, deeply pipelined processor architectures are being employed to meet desired system performance. System architects critically need modeling techniques that allow exploration, evaluation, customization and validation of different processor pipeline configurations, tuned for a specific application domain. We propose a novel finite state machine (FSM) based modeling of pipelined processors and define a set of properties that can be used to verify the correctness of in-order execution in the presence of fragmented pipelines and multicycle functional units. Our approach leverages the system architect's knowledge about the behavior of the pipelined processor through architecture description language (ADL) constructs, and thus allows a powerful top-down approach to pipeline verification. We applied this methodology to the DLX processor to demonstrate the usefulness of our approach. Prabhat Mishra 0001, Nikil Dutt, Alexandru Nicolau, Hiroyuki Tomiyama |
DATE | 4 |
| 2001 | New directions in compiler technology for embedded systems (embedded tutorial)abstractTraditionally, compiler technology has focused on the generation of code with the goal of improving performance for a variety of applications running on general-purpose processor architectures. In the embedded system space, compiler technology is faced with many new challenges, including: code generation for specialized architectural features, requireing a highly flexible degree of retargetability; memory-aware code generation that exploits the timing and structure of the embedded system's memory organization; optimizing software to meet both real-time and performance constraints; energy- and power-aware software generation, both from the context of energy minimization, as well as power modulation; code size minimization for memory-constrained embedded systems; coarse-grain transformations for tightly-coupled, memory-constrained multi-processor architectures; and interaction with the operating system for active management of embedded system resources. This paper discusses new directions for compiler technology, surveys some of the current research efforts and illustrates proposed solutions to selected issues. Nikil Dutt, Alexandru Nicolau, Hiroyuki Tomiyama, Ashok Halambi |
ASP-DAC | 3 |
| 1998 | Module Selection Using Manufacturing InformationabstractSince manufacturing processes inherently fluctuate, LSI chips which are produced from the same design have different propagation delays. However, the difference in delays caused by the process fluctuation has rarely been considered in most high-level synthesis systems which were developed before. This paper presents a new approach to module selection in high-level synthesis, which exploits difference in functional unit delays. First, a module library model which assumes the probabilistic nature of functional unit delays is presented. Then, we propose a module selection problem and an algorithm which minimizes the cost per faultless chip. Experimental results demonstrate that the proposed algorithm finds the optimal module selection which would not have been explored without manufacturing information. Hiroyuki Tomiyama, Hiroto Yasuura |
ASP-DAC | 1 |
| 1998 | Instruction Scheduling for Power Reduction in Processor-Based System DesignabstractThis paper proposes an instruction scheduling technique to reduce power consumed for off-chip driving. The technique minimizes the switching activity of a data bus between an on-chip cache and a main memory when instruction cache misses occur. The scheduling problem is formulated and a scheduling algorithm is also presented. Experimental results demonstrate the effectiveness and the efficiency of the proposed algorithm. Hiroyuki Tomiyama, Tohru Ishihara, Akihiko Inoue, Hiroto Yasuura |
DATE | 1 |
| 1997 | Memory-CPU Size Optimization for Embedded System DesignsabstractEntire systems embedded in a chip and consistingof a processor, memory, and system-specific peripheral hardwareare now commonly contained in commodity electronicdevices. Cost minimization of these systems is of paramounteconomic importance to manufactures of these devices. Byemploying a variable configuration processor in conjunctionwith a multi-precision compiler generator there are situationsin which considerable system cost reduction can be obtainedby synthesizing a CPU that is narrower than the largest variablein the application program. Barry Shackleford, Mitsuhiro Yasuda, Etsuko Okushi, Hisao Koizumi, Hiroyuki Tomiyama, Hiroto Yasuura |
DAC | 5 |
| 1997 | Code placement techniques for cache miss rate reductionabstractIn the design of embedded systems with cache memories, it is important to minimize the cache miss rates to reduce power consumption of the systems as well as improve the performance. In this article, we propose two code placement methods ( a simplified method and a refined one) to reduce miss rates of instruction caches. We first define a simplified code placement problem without an attempt to minimize the code size. The problem is formulated as an integer linear programming (ILP) problem, by which an optimal placement can be found. Experimental results show that the simplified method reduces cache misses by an average of 30% (max. 77%). However, the code size obtained by the simplified method tends to be large, which inevitably leads to a larger memory size. In order to overcome this limitation, we further propose a refined code placement method in which the code size provided by the system designers must be satisfied. The effectiveness of the refined method is also demonstrated. Hiroyuki Tomiyama, Hiroto Yasuura |
ACM Trans. Design Autom. Electr. Syst. | 1 |