EDBT 2026 Demo / reviewers in the wild / expert
Shimpei Sato
dblp:78/7652
· DBLP profile ↗
20ranked-venue papers
4as first author
4since 2021 · last 2023
0000-0003-0292-1391ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 19 · 4 first-author · 4 since 2021Artificial intelligence and machine learning · 4 · 3 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Whole-Body Torque Control Without Joint Position Control Using Vibration-Suppressed Friction Compensation for Bipedal Locomotion of Gear-Driven Torque Sensorless HumanoidabstractHumanoids operate in repeated contact and non-contact with their environment and so the motion of humanoids such as walking on uneven terrain or in a narrow space requires the accurate force and position control. Joint torque control systems are suitable for position and force control, but are prone to friction and other modeling errors. To solve this problem, methods have been proposed to realize torque control in combination with joint position control systems or by improving joint structures such as sensors and actuators, but these methods have problems such as response delay and increased weight and volume. Thus, it is difficult to achieve motion of life-sized humanoids by whole-body torque control. In this paper, we solve challenges not with one specific layer, but rather with multiple layers that complement each other. We propose a hierarchical whole-body torque control method using four layers: friction compensation based on a vibration-suppressed model, whole-body resolved acceleration control using priority, center-of-gravity acceleration control based on foot-guided control, and landing position time modification based on capture point. We verify through walking experiments that the proposed methods can control the life-sized humanoid robot driven by high-reduction ratio joints by whole-body torque control without a torque sensor or joint position control, and that it enables the robot to move and even transport an object on outdoor uneven terrain. Takuma Hiraoka, Shimpei Sato, Naoki Hiraoka, Annan Tang, Kunio Kojima, Kei Okada, Masayuki Inaba, Koji Kawasaki |
IROS | 2 |
| 2023 | Humanoid Walking System with CNN-Based Uneven Terrain Recognition and Landing Control with Swing-Leg Velocity ConstraintsabstractIn order for a humanoid robot to traverse uneven terrain without falling over, the robot must control its landing position appropriately. To determine the landing position, there are two difficulties in terrain recognition and leg motion control. In terrain recognition, it is difficult to recognize and avoid terrain such as steps and obstacles that cannot be landed on in real-time. In leg motion control, it is necessary to land at appropriate positions and times to control the CoG trajectory while limiting the velocity of the swing-leg to suppress the landing impact. For solving these problems, we propose a recognition and walking control system on uneven terrain. In terrain recognition, we improved the recognition accuracy while satisfying real-time performance by using a CNN that learns the relationship between the foot and the geometric information of the surrounding terrain. In the leg motion control, landing impact was reduced by modifying the landing position under not only (1) terrain constraint and (2) robot stability constraint, but also (3) leg velocity constraint. We verified the effectiveness of the proposed system through uneven terrain walking and push recovery experiments using the actual robot. Shimpei Sato, Kunio Kojima, Naoki Hiraoka, Kei Okada, Masayuki Inaba |
IROS | 1 |
| 2022 | Robust Humanoid Walking System Considering Recognized Terrain and Robots' BalanceabstractWhen robots walk on uneven terrain, trajectory planning should take into account both the whole-body dy-namics and the ground geometry simultaneously. In uneven terrain environments, there are only a limited number of places where the robot is able to make stable contact with the ground without its feet wobbling or slipping because of the intricate round geometry. In such environments, the optional landing position and time to maintain the robot's balance and stable foot contact are not obvious and computationally expensive. In this study, we propose a robust walking system that integrates environment recognition using steppable regions and walking control for a humanoid robot to walk on uneven terrain. In this paper, a steppable region is defined as a two-dimensional convex hull that represents a region where a robot is capable of landing. We propose a method to compute the steppable region quickly by 2.SD projection of the environment points and spatial filtering. In this system, the walking controller integrates the steppable region with the Capture Region to modify the landing position from a two-dimensional geometric calculation. In addition, to cope with the environment recognition error, we have introduced a trajectory generation that allows the feet to penetrate the ground and hybrid control of position and torque. We verified the effectiveness of the proposed system through experiments in which a life-size humanoid robot walked on uneven terrain and recovered when pushed. Shimpei Sato, Yuta Kojio, Youhei Kakiuchi, Kunio Kojima, Kei Okada, Masayuki Inaba |
IROS | 1 |
| 2021 | Drop Prevention Control for Humanoid Robots Carrying Stacked BoxesabstractWe developed a method to enable a humanoid robot to carry stacked boxes. In order to transport objects efficiently, it is necessary to carry multiple objects at the same time, but in previous studies, humanoid robots have only been able to carry a single object. When a humanoid robot carries stacked boxes, the robot drops boxes when the positional relationship between un-grasped boxes changes. The causes for dropping the boxes can be divided into sudden changes attributed to robot making turns or losing balance, and the accumulation of small changes that occur because of the impact of landing while walking. We propose a method that prevents sudden changes in the stacked boxes by smoothing the hand trajectory and modifying the misalignment by tilting or shaking the entire stack. We verify the effectiveness of proposed method for enabling a humanoid robot to carry stacked boxes through experiments using a simulator and an actual robot. Shimpei Sato, Yuta Kojio, Kunio Kojima, Fumihito Sugai, Youhei Kakiuchi, Kei Okada, Masayuki Inaba |
IROS | 1 |
| 2020 | Tiny On-Chip Memory Realization of Weight Sparseness Split-CNNs on Low-end FPGAsabstractConsidering implementation in a low-end FPGA with more restrictions on on-chip memory resources and external memory bandwidth, existing methods are Limited by communication with external memory. The on-chip memory in a low-end FPGA is small, hence cannot store an entire feature map. It is imperative to rely on external memory for buffering. Although an on-chip memory in an FPGA is fast, implementing sparse CNN in low-end FPGAs is hindered by their limited memory size. To address the limitations of external memory bandwidth and its size, we employ a split-CNN [1] that splits an input image into small spatial patches and tests each patch using a CNN model as shown in Fig. 1. Since the feature-map to be processed at one time is reduced by the splitting, the amount of memory required for buffering is reduced. Akira Jinguji, Shimpei Sato, Hiroki Nakahara |
FCCM | 2 |
| 2019 | An FPGA-based Fine Tuning Accelerator for a Sparse CNNabstractFine-tuning learns abundant feature expression for a wide range of natural images by using a pre-trained CNN model. It can be applied to a wide range of the neural network (NN)based computer vision problems. This paper proposes an FPGA-based fine-tuning accelerator for a sparse convolutional neural network (CNN). The proposed architecture consists of sparse convolutional units and pooling units with distributed stacks those are suitable for a sparse CNN. Additionally, this paper presents a fine-tuning scheme, which loads a pre-trained sparse CNN to reduce the memory size for the training step. Thus, our fine-tuning scheme stores all parameters on BRAMs and Ultra RAMs in the case of the Xilinx Virtex UltraScale+ FPGA to accelerate the training computation and reduce power consumption by eliminating energy-costly DRAM accesses. We implemented on a Xilinx Virtex UltraScale+ VC1525 acceleration development kit. Experimental results show that the proposed sparse finetuning accelerator on the FPGA can achieve four times faster, 2.9 times lower power consumption, and 11.6 times better performance per power, compared to the existing NVidia GTX1080Ti GPU. Hiroki Nakahara, Akira Jinguji, Masayuki Shimoda, Shimpei Sato |
FPGA | 4 |
| 2019 | FPGA-Based Training Accelerator Utilizing Sparseness of Convolutional Neural NetworkabstractTraining of convolutional neural networks (CNNs) is almost exclusively performed on large clusters of GPUs. However, it consumes vast amounts of power. Thus, high-speed training systems superior in low-power consumption are desired. This paper proposes an FPGA-based training accelerator utilizing a sparseness of a CNN, which consists of universal convolutional units and pooling units with distributed stacks. The proposed universal convolution architecture supports various convolution operations, such as the point-wise, depth-wise, large kernel and atrous convolutions used in the modern CNN. Additionally, we utilize a fine-tuning scheme, which loads a pre-trained dense CNN to reduce the memory size for the training process, while it considers important connectivity to preserve recognition accuracy. Our training scheme reduces 85% parameters to accelerate the training computation and reduce on-chip size. Thus, it eliminates energy-consuming DRAM accesses. We implemented the proposed training accelerator on a Xilinx Virtex UltraScale+ VC1525 acceleration development board. Experimental results show that the proposed sparse CNN training accelerator on the FPGA can achieve four times faster, 2.9 times lower power consumption, and 11.6 times better performance per power, compared to the existing NVIDIA RTX2080Ti GPU. Hiroki Nakahara, Youki Sada, Masayuki Shimoda, Kouki Sayama, Akira Jinguji, Shimpei Sato |
FPL | 6 |
| 2019 | A Low Area Overhead Design for High-Performance General-Synchronous Circuits with Speculative ExecutionabstractIn order to obtain higher performance of digital circuits in advanced technology nodes, the effects of delay variability on performance should be reduced as much as possible with less overheads. Clock scheduling and speculative execution have potential to ease the influence of the delay difference among signal paths and the delay variability in each signal path, respectively. In this paper, we propose a high-performance digital circuit design method with speculative executions with less overhead by utilizing clock scheduling with delay insertions effectively. The necessity of speculations that cause overheads is effectively reduced by clock scheduling with delay insertion. Experiments show that a generated circuit achieves 26% performance improvement with 1.3% area overhead compared to a circuit without clock scheduling and without speculative execution. Shimpei Sato, Eijiro Sassa, Yuta Ukon, Atsushi Takahashi 0001 |
ISCAS | 1 |
| 2018 | A Lightweight YOLOv2: A Binarized CNN with A Parallel Support Vector Regression for an FPGAabstractA frame object detection problem consists of two problems: one is a regression problem to spatially separated bounding boxes, the second is the associated classification of the objects within realtime frame rate. It is widely used in the embedded systems, such as robotics, autonomous driving, security, and drones - all of which require high-performance and low-power consumption. This paper implements the YOLO (You only look once) object detector on an FPGA, which is faster and has a higher accuracy. It is based on the convolutional deep neural network (CNN), and it is a dominant part both the performance and the area. However, the object detector based on the CNN consists of a bounding box prediction (regression) and a class estimation (classification). Thus, the conventional all binarized CNN fails to recognize in most cases. In the paper, we propose a lightweight YOLOv2, which consists of the binarized CNN for a feature extraction and the parallel support vector regression (SVR) for both a classification and a localization. To our knowledge, this is the first time binarized CNN»s have been successfully used in object detection. We implement a pipelined based architecture for the lightweight YOLOv2 on the Xilinx Inc. zcu102 board, which has the Xilinx Inc. Zynq Ultrascale+ MPSoC. The implemented object detector archived 40.81 frames per second (FPS). Compared with the ARM Cortex-A57, it was 177.4 times faster, it dissipated 1.1 times more power, and its performance per power efficiency was 158.9 times better. Also, compared with the nVidia Pascall embedded GPU, it was 27.5 times faster, it dissipated 1.5 times lower power, and its performance per power efficiency was 42.9 times better. Thus, our method is suitable for the frame object detector for an embedded vision system. Hiroki Nakahara, Haruyoshi Yonekawa, Tomoya Fujii, Shimpei Sato |
FPGA | 4 |
| 2018 | A Demonstration of FPGA-Based You Only Look Once Version2 (YOLOv2)abstractWe implement the YOLO (You only look once) object detector on an FPGA, which is faster and has higher accuracy. It is based on the convolutional deep neural network (CNN), and it is a dominant part of both the performance and the area. It is widely used in the embedded systems, such as robotics, autonomous driving, security, and drones, all of which require high-performance and low-power consumption. A frame object detection problem consists of two problems: one is a regression problem to spatially separated bounding boxes, the second is the associated classification of the objects within realtime frame rate. We used the binary (1 bit) precision CNN for feature extraction and the half-precision (16 bit) precision CNN for both classification and localization. We implement a pipelined based architecture for the mixed-precision YOLOv2 on the Xilinx Inc. zcu102 board, which has the Xilinx Inc. Zynq Ultrascale+ MPSoC. The implemented object detector archived 35.71 frames per second (FPS), which is faster than the standard video speed (29.9 FPS). Compared with a CPU and a GPU, an FPGA based accelerator was superior in power performance efficiency. Our method is suitable for the frame object detector for an embedded vision system. Hiroki Nakahara, Masayuki Shimoda, Shimpei Sato |
FPL | 3 |
| 2018 | Demonstration of Object Detection for Event-Driven Cameras on FPGAs and GPUsabstractWe demonstrate an object detection system using a sliding window method for an event-driven camera[3] which outputs a subtracted frame (usually a binary value) when changes are detected in captured images. Fig. 1 shows the overall architecture of the object detector using a sliding window. First, our system extracts the object region candidates by applying a sliding window to the picture output by an event-driven camera. Since the event-driven camera outputs are binary, the sliding window determines whether an object exists or not based on the proportion of white pixels inside the bounding box as shown in Fig. 2. It reduces the number of proposed regions from 1,485 to average 10. The extracted pictures are resized to 40 × 40 size and applied to the ABCNN (All Binarized Convolutional Neural Network[7]), which is a BCNN (Binarized Convolutional Neural Network[4]) in which the first convolutional layer is done in binary form as well. All binarization is found to improve both the area and computation time, while the classification accuracy decreases. If the ABCNN infers the detected object as human, then it draws a box on the proposed region. Since more than one bounding box is commonly drawn against a single detected object, non-maximum suppression is used to reduce the extras to a single box. The proposed system is implemented on an FPGA and a GPU to show the comparison of them. Masayuki Shimoda, Shimpei Sato, Hiroki Nakahara |
FPL | 2 |
| 2018 | An FPGA Realization of OpenPose Based on a Sparse Weight Convolutional Neural NetworkabstractThe OpenPose is a kind of a deep learning based pose estimator which achieved a top accuracy for multiple person pose estimations. Even if using the OpenPose, it is necessary to used high-performance GPU since it requires massive parameters access with high-bandwidth off-chip GDDR5 memories and a higher operation clock frequency. Thus, the power consumption becomes a critical issue to realization. Also, its computation time is slower than the current video standard frame speed (29.97 FPS). In the paper, we introduce a sparse weight CNN to reduce the amount of memory size for weights, which is Then, we offer the indirect memory access architecture to realize the sparse CNN convolutional operation efficiently. Also, to increase throughput further, we applied the six stages of pipeline architecture with a pipeline buffer memory realization. Our implementation satisfied the timing constraint for real-time applications. Since our architecture computed an image with 42.6 msec, the number of frames per second (FPS) was 23.43. We measured the total board power consumption: It was 55 Watt. Thus, the performance per power efficiency was 0.444 (FPS/W). Compared with the NVidia Titan X Pascal architecture GPU, it was 3.49 times faster, it dissipated 3.54 times lower power, and its performance per power efficiency was 13.05 times better. As far as we know, this work is the first FPGA implementation of the OpenPose. Akira Jinguji, Tomoya Fujii, Shimpei Sato, Hiroki Nakahara |
FPT | 3 |
| 2018 | A Tri-State Weight Convolutional Neural Network for an FPGA: Applied to YOLOv2 Object DetectorabstractA frame object detection, such as the YOLO (You only look once), is used in embedded vision systems, such as a robot, an automobile, a security camera, and a drone. However, it requires highly performance-per-power detection by an inexpensive device. In the paper, we propose a tri-state weight CNN, which is a generalization of a low-precision and sparse (pruning) for CNN weight. In the former part, we set a weight {-1,0,+1} as a ternary CNN, while in the latter part, we set a {-w,0,+w} as a sparse weight CNN. The proposed tri-state CNN is a kind of a mixed-precision one, which is suitable for an object detector consisting of a bounding box prediction (regression) and a class estimation (classification). We apply an indirect memory access architecture to skip zero part and propose the weight parallel 2D convolutional circuit. It can efficiently be applied to the AlexNet based CNN, which has different size kernels. We design the AlexNet based YOLOv2 to reduce the number of layers toward low-latency computation. In the experiment, the proposed tri-state scheme CNN reduces the memory size for weight by 92%. We implement the proposed tri-state weight YOLOv2 on the AvNet Inc. UltraZed-EG starter kit, which has the Xilinx Inc. Zynq Ultrascale+ MPSoC ZU3EG. It archived 61.70 frames per second (FPS), which exceeds the standard video frame rate (29.97 FPS). Compared with the ARM Cortex-A57, it was 268.2 times faster, and its performance per power efficiency was 313.51 times better. Also, compared with the NVidia Pascal embedded GPU, it was 4.0 times faster, and its power performance efficiency was 11.35 times better. Hiroki Nakahara, Masayuki Shimoda, Shimpei Sato |
FPT | 3 |
| 2017 | A fully connected layer elimination for a binarizec convolutional neural network on an FPGAabstractA pre-trained convolutional deep neural network (CNN) is widely used for embedded systems, which requires highly power-and-area efficiency. In that case, the CPU is too slow, the embedded GPU dissipates much power, and the ASIC cannot keep up with the rapidly progress of the CNN variations. This paper uses a binarized CNN which treats only binary 2-values for the inputs and the weights. Since the multiplier is replaced into an XNOR circuit, we can realize a high-performance MAC circuit by using many XNOR circuits. In the paper, we eliminate internal FC layers excluding the last one, then, insert a binarized average pooling layer, which can be realized by a majority circuit for binarized (1/0) values. In that case, since the weight memory is replaced into the 1's counter, we can realize a compact and faster CNN than the conventional ones. We implemented the VGG-11 benchmark CNN for the CIFAR10 image classification task on the Xilinx Inc. Zedboard. Compared with the conventional binarized implementations on an FPGA, the classification accuracy was almost the same, the performance per power efficiency is 5.1 better, as for the performance per area efficiency, it is 8.0 times better, and as for the performance per memory, it is 8.2 times better. Hiroki Nakahara, Tomoya Fujii, Shimpei Sato |
FPL | 3 |
| 2017 | An object detector based on multiscale sliding window search using a fully pipelined binarized CNN on an FPGAabstractAn object detection problem consists of two problems: one is classification of detected object category and the other is localization. Frame object detection is used in an embedded vision systems, such as a robot, an automobile, a security camera, and a drone. These applications require high-performance computation and low-power consumption by an inexpensive device. This paper proposes multiscale sliding window based object detector using a fully pipelined binarized deep convolutional neural network (BCNN) on an FPGA. It consists of a sliding window part, a fully pipelined BCNN classifier, and an ARM processing unit for detection. Duplicate detections were filtered by using a non-maximum suppression algorithm running on the ARM processor. We propose the fully pipelined layers for the BCNN and its architecture for FPGA realization. Since the proposed BCNN circuit uses on-chip memories on the FPGA, its throughput is higher than a GPU based one with practical recognition accuracy. We trained the VGG11 based BCNN using the KITTI vision benchmark for the car detection scenario. Then, we implemented the proposed object detector on the Xilinx Inc. Zynq UltraScale+ MPSoC zcu102 evaluation board. The GPU based object detectors were too slow for the realtime application requirement (HD frame rate), with the exception of YOLOv2. As compared with the GPU implementation of YOLOv2, the proposed FPGA detector had higher recognition accuracy and lower power consumption. Compared with the YOLOv2, the proposed FPGA one is higher with respect to recognition accuracy, and its power consumption is lower than the GPU based YOLOv2. Thus, the FPGA based object detector suitable for the embedded realtime applications. Hiroki Nakahara, Haruyoshi Yonekawa, Shimpei Sato |
FPT | 3 |
| 2017 | All binarized convolutional neural network and its implementation on an FPGAabstractA pre-trained convolutional neural network (CNN) is a feed-forward computation perspective, which is widely used in the embedded systems requiring highly power-and-area efficiency. This paper realizes a binarized CNN which treats only binarized values (+1/-1) for the weights and the activation value. In this case, the multiplier is replaced by an XNOR circuit instead of a dedicated DSP block. Binarization for both weights and activation are more suitable for hardware implementation. However, the first convolutional layer still calculates in integer precision, since the input value is 8 bit RGB pixel, and not binarized. In this paper, we decompose the input value into maps of which each pixel is in 1-bit precision. The proposed method enables a binarized CNN to use bitwise operation in all layers, and shares a binarized convolutional circuit among all convolutional layers. We call this all binarized CNN. We compared our proposal with conventional ones. Since all binarized CNN do not require a dedicated DSP block, our proposal is smaller and 1.2 times faster than the typical CNNs, and almost maintains baseline classification accuracy. In addition, pipelined all binarized CNN achieved 1840 FPS, consumed 0.3 watts and its accuracy was 82.8%. Masayuki Shimoda, Shimpei Sato, Hiroki Nakahara |
FPT | 2 |
| 2017 | Fast and Cycle-Accurate Emulation of Large-Scale Networks-on-Chip Using a Single FPGAabstractModeling and simulation/emulation play a major role in research and development of novel Networks-on-Chip (NoCs). However, conventional software simulators are so slow that studying NoCs for emerging many-core systems with hundreds to thousands of cores is challenging. State-of-the-art FPGA-based NoC emulators have shown great potential in speeding up the NoC simulation, but they cannot emulate large-scale NoCs due to the FPGA capacity constraints. Moreover, emulating large-scale NoCs under synthetic workloads on FPGAs typically requires a large amount of memory and thus involves the use of off-chip memory, which makes the overall design much more complicated and may substantially degrade the emulation speed. This article presents methods for fast and cycle-accurate emulation of NoCs with up to thousands of nodes using a single FPGA. We first describe how to emulate a NoC under a synthetic workload using only FPGA on-chip memory (BRAMs). We next present a novel use of time-division multiplexing where BRAMs are effectively used for emulating a network using a small number of nodes, thereby overcoming the FPGA capacity constraints. We propose methods for emulating both direct and indirect networks, focusing on the commonly used meshes and fat-trees (k-aryn-trees). This is different from prior work that considers only direct networks. Using the proposed methods, we build a NoC emulator, called FNoC, and demonstrate the emulation of some mesh-based and fat-tree-based NoCs with canonical router architectures. Our evaluation results show that (1) the size of the largest NoC that can be emulated depends on only the FPGA on-chip memory capacity; (2) a mesh-based NoC with 16,384 nodes (128×128 NoC) and a fat-tree-based NoC with 6,144 switch nodes and 4,096 terminal nodes (4-ary 6-tree NoC) can be emulated using a single Virtex-7 FPGA; and (3) when emulating these two NoCs, we achieve, respectively, 5,047× and 232× speedups over BookSim, one of the most widely used software-based NoC simulators, while maintaining the same level of accuracy. Thiem Van Chu, Shimpei Sato, Kenji Kise |
ACM Trans. Reconfigurable Technol. Syst. | 2 |
| 2015 | Enabling Fast and Accurate Emulation of Large-Scale Network on Chip Architectures on a Single FPGAabstractNetwork on Chip (NoC) has become the de facto on-chip communication architecture of many-core systems. This paper proposes an FPGA-based NoC emulator which can achieve an ultra-fast simulation speed. We improve the scalability of the NoC emulator without simplifying the emulated architectures or using off-chip resources. We introduce a novel method which enables to accurately emulate NoC designs under synthetic workloads without using a large amount of memory by decoupling the time of the emulated NoC and the time of the traffic generators. Additionally, we propose a method based on the time-division multiplexing technique to emulate the behavior of the entire network using several physical nodes while effectively using FPGA resources. We show that an implementation of the proposed NoC emulator on a Virtex-7 FPGA can achieve 2, 745x simulation speedup over Booksim, one of the most widely used software-based NoC simulator, while maintaining the simulation accuracy. Thiem Van Chu, Shimpei Sato, Kenji Kise |
FCCM | 2 |
| 2015 | Ultra-fast NoC emulation on a single FPGAabstractNetwork-on-Chip (NoC) has become the de facto on-chip communication architecture for many-core systems. This paper proposes novel methods for emulating large-scale NoC designs on a single FPGA. Since FPGAs offer a highly parallel platform, FPGA-based emulation can be much faster than the software-based approach. However, emulating NoC designs with up to thousands of nodes is a challenging task due to the FPGA capacity constraints. We first describe how to accurately model synthetic workloads on FPGA by separating the time of the emulated network and the times of the traffic generation units. We next present a novel use of time-multiplexing in emulating the entire network using several physical nodes. Finally, we show the basic steps to apply the proposed methods to emulate different NoC architectures. The proposed methods enable ultrafast emulations of large-scale NoC designs with up to thousands of nodes using only on-chip resources of a single FPGA. In particular, more than 5,000× simulation speedup over BookSim, a widely used software-based NoC simulator, is achieved. Thiem Van Chu, Shimpei Sato, Kenji Kise |
FPL | 2 |
| 2009 | A Study of an Infrastructure for Research and Development of Many-Core ProcessorsabstractMany-core processors which have thousands of cores on a chip will be realized. We developed an infrastructure which accelerates the research and development of such many-core processors. This paper describes three main elements provided by our infrastructure. The first element is the definition of simple many-core processor architecture called M-Core. The second is SimMc, a software simulator of M-Core. The third is the software library MClib which helps the development of application programs for M-Core. The simulation speed of SimMc and the parallelization efficiency of M-Core are evaluated using some benchmark programs. We show that our infrastructure accelerates the research and development of many-core processors. Koh Uehara, Shimpei Sato, Takefumi Miyoshi, Kenji Kise |
PDCAT | 2 |