Hiroaki Inoue

dblp:18/7023 · DBLP profile ↗
← Back
28ranked-venue papers
16as first author
3since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 18 · 11 first-authorSoftware engineering, systems software and programming languages · 6 · 3 first-author · 2 since 2021Computer networks · 2 · 2 first-authorSecurity and privacy · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2025 FragmentFool: Fragment-based Adversarial Perturbation for Graph Neural Network-based Vulnerability Detection
abstract
Software vulnerability detection has achieved promising performance using graph neural networks (GNNs) to capture structural information in source code graph representations. However, these methods are vulnerable to various attacks. Exploiting GNN vulnerabilities can significantly compromise the robustness and reliability of these detection systems. In this study, we investigate the susceptibility of the current GNN-based source code vulnerability detection methods to structural perturbations. We propose a novel approach to perturb GNN inputs through fragment or subtree manipulation of the original graph representations, such as abstract syntax trees, targeting structure features as a critical feature when using GNN for vulnerability detection. Our comprehensive experiments demonstrate that the proposed perturbation approach effectively attacks common GNN models, thereby reducing the overall performance.
Muhammad Fakhrur Rozi, Tao Ban, Seiichi Ozawa, Hiroaki Inoue, Takeshi Takahashi 0001, Sajjad Dadkhah
PST4
2022 Automated formal analysis of temporal properties of Ladder programs
Cláudio Belo Lourenço, Denis Cousineau 0002, Florian Faissole, Claude Marché, David Mentré, Hiroaki Inoue
Int. J. Softw. Tools Technol. Transf.6
2021 Automated Verification of Temporal Properties of Ladder Programs
Cláudio Belo Lourenço, Denis Cousineau 0002, Florian Faissole, Claude Marché, David Mentré, Hiroaki Inoue
FMICS6
2019 A type system for first-class layers with inheritance, subtyping, and swapping
Hiroaki Inoue, Atsushi Igarashi
Sci. Comput. Program.1
2018 ContextWorkflow: A Monadic DSL for Compensable and Interruptible Executions
abstract
Context-aware applications, whose behavior reactively depends on the time-varying status of the surrounding environment - such as network connection, battery level, and sensors - are getting more and more pervasive and important. The term "context-awareness" usually suggests prompt reactions to context changes: as the context change signals that the current execution cannot be continued, the application should immediately abort its execution, possibly does some clean-up tasks, and suspend until the context allows it to restart. Interruptions, or asynchronous exceptions, are useful to achieve context-awareness. It is, however, difficult to program with interruptions in a compositional way in most programming languages because their support is too primitive, relying on synchronous exception handling mechanism such as try-catch. We propose a new domain-specific language ContextWorkflow for interruptible programs as a solution to the problem. A basic unit of an interruptible program is a workflow, i.e., a sequence of atomic computations accompanied with compensation actions. The uniqueness of ContextWorkflow is that, during its execution, a workflow keeps watching the context between atomic actions and decides if the computation should be continued, aborted, or suspended. Our contribution of this paper is as follows; (1) the design of a workflow-like language with asynchronous interruption, checkpointing, sub-workflows and suspension; (2) a formal semantics of the core language; (3) a monadic interpreter corresponding to the semantics; and (4) its concrete implementation as an embedded domain-specific language in Scala.
Hiroaki Inoue, Tomoyuki Aotani, Atsushi Igarashi
ECOOP1
2018 Parallel Rate Distortion Optimized Quantization for 4K Real-time GPU-based HEVC Encoder
abstract
We proposed a highly parallel rate distortion optimized quantization (RDOQ) for 4K real-time GPU-based HEVC encoder. RDOQ optimizes a quantized value in the view of a tradeoff relationship between picture quality and compression efficiency. While it brings better compression efficiency for HEVC, it is difficult to process on GPU. This is because two parts which compose RDOQ; the cost calculation and the optimization, are sequential. The proposed method parallelizes both of parts and accelerates RDOQ on GPU. For the cost calculation, the proposed method uses the history data of previous frame. Furthermore, to parallelize the optimization part, the proposed method applies bi-directional parallel scan which can be processed on GPU. Experimental results show that the proposed method improved 26.43 % of BD-rate compared with the conventional GPU-based encoder without RDOQ which enables 4K/60FPS real-time encoding. Furthermore, the proposed method is 5x faster than x265 which is the most practical CPU-based encoder under similar conditions of BD-rate.
Hiroaki Igarashi, Fumiyo Takano, Takashi Takenaka, Hiroaki Inoue, Tatsuji Moriyoshi
VCIP4
2015 A Sound Type System for Layer Subtyping and Dynamically Activated First-Class Layers
Hiroaki Inoue, Atsushi Igarashi
APLAS1
2014 Low latency FPGA acceleration of market data feed arbitration
abstract
A critical source of information in automated trading is provided by market data feeds from financial exchanges. Two identical feeds, known as the A and B feeds, are used in reducing message loss. This paper presents a reconfigurable acceleration approach to A/B arbitration, operating at the network level, and supporting any messaging protocol. The key challenges are: providing efficient, low latency operations; supporting any market data protocol; and meeting the requirements of downstream applications. To facilitate a range of downstream applications, one windowing mode prioritising low latency, and three dynamically configurable windowing methods prioritising high reliability are provided. We implement a new low latency, high throughput architecture and compare the performance of the NASDAQ TotalView-ITCH, OPRA and ARCA market data feed protocols using a Xilinx Virtex-6 FPGA. The most resource intensive protocol, TotalView-ITCH, is also implemented in a Xilinx Virtex-5 FPGA within a network interface card. We offer latencies 10 times lower than an FPGA-based commercial design and 4.1 times lower than the hardware-accelerated IBM PowerEN processor, with throughputs more than double the required 10Gbps line rate.
Stewart Denholm, Hiroaki Inoue, Takashi Takenaka, Tobias Becker, Wayne Luk
ASAP2
2014 Mapping complex algorithm into FPGA with High Level Synthesis reconfigurable chips with High Level Synthesis compared with CPU, GPGPU
abstract
This presentation discusses on the comparison between "Reconfigurable Chip with High Level Synthesis" and "CPU, GPCPU with compiler such as CUDA" from the compiler perspective. Initially, we introduce several demands for acceleration with FPGA to achieve low latency calculation and control. As an application example, we show a High Frequency Trading. We accelerate it by FPGA NIC with C-based and SQL-based HLS, and show the necessity of high level language customizable reconfigurable chip. Then, we illustrate the difference of FPGA and processor (CPU, GPGPU) with the "FSM+Datapath" model and examine how the architecture difference affects delay and parallelism of operations. Next, we discuss parallelization of operations, threads with High Level Synthesis for FPGA and software compiler for processors. The main advantage of the former method is it is able to parallelize operations beyond control dependencies while the latter method has to obey control dependencies. Finally, some experimental results prove that "FPGA and HLS" generate better performance than a processor for control intensive algorithm.
Kazutoshi Wakabayashi, Takashi Takenaka, Hiroaki Inoue
ASP-DAC3
2014 Caching memcached at reconfigurable network interface
abstract
Memcached is a technology that improves response speed of web servers by caching data on DRAMs in distributed servers. In order to achieve higher performance, memcached has been evaluated on various platforms. Among them, FPGA seems to be the most efficient platform to run memcached, and several research groups are trying to achieve higher throughput with it. However, it is difficult to utilize a large amount of memory (several dozen gigabytes) with an FPGA. Some groups are trying to solve this problem by using an embedded CPU for memory allocation and another group is employing an SSD. Unlike other approaches that try to replace memcached itself on FPGAs, our approach augments the software memcached running on the host CPU by caching its data and some operations at the FPGA-equipped network interface card (NIC) mounted on the server. The locality of memcached data enables the FPGA NIC to have a fairly high hit rate with a smaller memory. We first explore the cache parameters by software simulations and estimate the effectiveness of our approach, and then prototype a system to prove its effectiveness. Through our evaluation with YCSB, a standard key-value store (KVS) benchmarking tool, we estimate that the latency improved by an order of magnitude over software memcached running on a high performance CPU.
Eric Shun Fukuda, Hiroaki Inoue, Takashi Takenaka, Dahoo Kim, Tsunaki Sadahisa, Tetsuya Asai, Masato Motomura
FPL2
2014 Achieving higher performance of memcached by caching at network interface
abstract
As the volume of data that web services handle is becoming larger, many web service providers are utilizing memcached, an in-memory key-value store to improve their web server's performance. While memcached usually runs on a server with a high performance processor, various hardware platforms has been evaluated for running memcached in order to achieve higher performance. Recently, several works that use FPGAs have successfully achieved higher performance than Xeon. These works, however, struggles to utilize a large memory with FPGAs. In this paper, we propose a system that enables us to overcome this problem and enhances memcached by caching a part of software memcached's commands and data to the network interface card equipped with an FPGA and a DRAM. Our evaluation showed that the NIC cache has less than 30% of hit rate for workload with Latest key selection distribution, and 30% to 60% for Zipf distribution workloads.
Eric Shun Fukuda, Hiroaki Inoue, Takashi Takenaka, Dahoo Kim, Tsunaki Sadahisa, Tetsuya Asai, Masato Motomura
FPT2
2013 Application-specific customisation of market data feed arbitration
abstract
Messages are transmitted from financial exchanges to update their members about changes in the market. As UDP packets are used for message transmission, members subscribe to two identical message feeds from the exchange to lower the risk of message loss or delay. As financial trades can be time sensitive, low latency arbitration between these market data feeds is of particular importance. Members must either provide generic arbitration for all of their financial applications, increasing latency, or arbitrate within each application which wastes resources and scales poorly. We present a reconfigurable accelerated approach for market feed arbitration operating at the network level. Multiple arbitrators can operate within a single FPGA to output customised feeds to downstream financial applications. Application-specific customisations are supported by each core, allowing different market feed messaging protocols, windowing operations and message buffering parameters. We model multiple-core arbitration and explore the scalability and performance improvements within and between cores. We demonstrate our design within a Xilinx Virtex-6 FPGA using the NASDAQ TotalView-ITCH 4.1 messaging standard. Our implementation operates at 16Gbps throughput, and with resource sharing, supports 12 independent cores, 33% more than simple core replication. A 56ns (7 clock cycles) windowing latency is achieved, 2.6 times lower than a hardware-accelerated CPU approach.
Stewart Denholm, Hiroaki Inoue, Takashi Takenaka, Wayne Luk
FPT2
2013 C-Based Complex Event Processing on Reconfigurable Hardware
abstract
This brief presents an efficient complex event-processing framework, designed to process a large number of sequential events on field-programmable gate arrays (FPGAs). Unlike conventional structured query language based approaches, our approach features logic automation constructed with a new C-based event language that supports regular expressions on the basis of C functions, so that a wide variety of event-processing applications can be efficiently mapped to FPGAs. Evaluations on an FPGA-based network interface card show that we can achieve 12.3 times better event-processing performance than does a CPU software in a financial trading application.
Hiroaki Inoue, Takashi Takenaka, Masato Motomura
IEEE Trans. Very Large Scale Integr. Syst.1
2012 Dynamic query switching for complex event processing on FPGAs
abstract
This paper presents an FPGA event processing system which can dynamically switch queries with the following features. (1) The switch can be performed without a need for a server shutdown. (2) The output stream includes results of the old query just until the new query starts producing results. (3) The old results and new results are not disordered in the output stream. Switching queries is essential to maintain reliability, availability, and serviceability and to increase business profit. We propose (i) a query management mechanism which makes both a new and old query run in parallel and (ii) a network which selects and send out proper result data from mixed result data. We have validated the design on a FPGA board with financial trading queries and shown the feasibility of the query switch with 32 queries running simultaneously, each of which is fed input events of 20Gbps throughput.
Masamichi Takagi, Takashi Takenaka, Hiroaki Inoue
FPL3
2012 A scalable complex event processing framework for combination of SQL-based continuous queries and C/C++ functions
abstract
SQL-based languages are widely used for Software-based Complex Event Processing (CEP) systems. This paper proposes a FPGA acceleration framework to compile a SQL-based event processing language, which is based on the ANSI standard proposal to support event processing, into a high-performance CEP engine on FPGAs. Besides the SQL's primitives such as partitioning, windowing, aggregation and pattern matching, the proposed framework allows C/C++ functions to implement complex algorithms required by real-world applications. The CEP engine also scales very efficiently processing multiple streams in parallel by making use of inexpensive block memories. Experimental results show that our proposed CEP engine, which is compiled for a financial analysis application calculating a popular trading benchmark to capture meaningful market trends, achieves 150M events/sec (20Gbps) processing performance for over 16,000 streams.
Takashi Takenaka, Masamichi Takagi, Hiroaki Inoue
FPL3
2011 20Gbps C-Based Complex Event Processing
abstract
This paper presents the world's fastest complex event processing system, designed to process a large number of events on FPGAs. Unlike conventional SQL-based approaches, our approach features logic automation constructed with a new C-based event language that supports regular expressions on the basis of C functions, so that a wide variety of event-processing applications can be efficiently mapped to FPGAs. Evaluations on an FPGA-based NIC show that we have achieved 20Gbps event processing performance in a financial trading application with only a small logic increase, 2.7%.
Hiroaki Inoue, Takashi Takenaka, Masato Motomura
FPL1
2011 Greening of Many-Core Processors in Network-Optimized Computing
abstract
This paper presents adaptive performance control for many-core processors in response to network traffic in order to reduce power dissipation. It presents two new techniques: smart core-number control, which averages out both performance and power dissipation by disabling inactive CPU cores, and smart core-performance control, which quickly adjusts both performance and power dissipation in each active CPU core by using a dynamic thermal controller. Evaluation results for a 1.8GHz 16-core/64-thread IBM PowerEN processor show that smart core-number control reduces power dissipation by 46% under low network traffic conditions and smart core-performance control reduces power dissipation by 21% even when large network traffic fluctuates by 80% for a short time.
Hiroaki Inoue, Kazuhisa Ishizaka, Junji Sakai
GLOBECOM1
2011 Test compression for dynamically reconfigurable processors
abstract
We present the world's first test compression technique that features automation of compression rules for test time reduction on dynamically reconfigurable processors. Evaluations on an actual 40-nm product show that our technique achieves a 2.7 times compression ratio for original configuration information (better than does GZIP), the peak decompression bandwidth of 1.6 GB/s, and 2.7 times shorter test times.
Hiroaki Inoue, Junya Yamada, Hideyuki Yoneda, Katsumi Togawa, Masato Motomura, Koichiro Furuta
ACM Trans. Reconfigurable Technol. Syst.1
2010 Test Compression for Dynamically Reconfigurable Processors
abstract
We present the world’s first test compression technique that features automation of compression rules for test time reduction on dynamically reconfigurable processors. Evaluations on an actual 40-nm product show that our technique achieves a 2.7 times compression ratio for original configuration information (better than does GZIP), the peak decompression bandwidth of 1.6 GB/s, and 2.7 times shorter test times.
Hiroaki Inoue, Junya Yamada, Hideyuki Yoneda, Katsumi Togawa, Koichiro Furuta
FPL1
2010 Onix: A Distributed Control Platform for Large-scale Production Networks
Teemu Koponen, Martín Casado, Natasha Gude, Jeremy Stribling, Leonid B. Poutievski, Rajiv Ramanathan, Yuichiro Iwata, Hiroaki Inoue, Takayuki Hama, Scott Shenker
OSDI9
2010 A robust seamless communication architecture for next-generation mobile terminals on multi-CPU SoCs
abstract
We propose a robust seamless communication architecture that enables legacy mobile terminal software designed for single-CPU processors to be run on multi-CPU processors without any software modifications. This architecture features two new technologies: proxy processes, which help achieve the design of its user-level system-call hooking and a robust design method, which reduces bandwidth variation by systematic parameter optimization. Our evaluations confirmed that this architecture achieves fundamental features with satisfactory performance, that we have succeeded in getting actual mobile terminal software to run on three CPUs without modifying the software, and that the robust design method reduces bandwidth variation by 21%.
Hiroaki Inoue, Junji Sakai, Masato Edahiro
ACM Trans. Embed. Comput. Syst.1
2009 Dynamic security domain scaling on embedded symmetric multiprocessors
abstract
We propose a method for dynamic security-domain scaling on SMPs that offers both highly scalable performance and high security for future high-end embedded systems. Its most important feature is its highly efficient use of processor resources, accomplished by dynamically changing the number of processors within a security-domain (i.e., dynamically yielding processors to other security-domains) in response to application load requirements. Two new technologies make this scaling possible without any virtualization software: (1) self-transition management and (2) unified virtual address mapping. Evaluations show that this domain control provides highly scalable performance and incurs almost no performance overhead in security-domains. The increase in OSs in binary code size is less than 1.5%, and the time required for individual state transitions is on the order of a single millisecond. This scaling is the first in the world to make possible the dynamic changing of the number of processors within a security-domain on an ARM SMP.
Hiroaki Inoue, Tsuyoshi Abe, Kazuhisa Ishizaka, Junji Sakai, Masato Edahiro
ACM Trans. Design Autom. Electr. Syst.1
2008 VAST: Virtualization-Assisted Concurrent Autonomous Self-Test
abstract
Virtualization-Assisted concurrent, autonomous Self-Test, or VAST, enables a multi-/many-core system to test itself, concurrently during normal operation, without any user-visible downtime. Such on-line self-test is required for large-scale robust systems with built-in support for circuit failure prediction, failure detection, diagnosis, and self-healing. The main idea behind VAST is hardware and software co-design of on-line self-test features in a multi-/many-core system through integration of: 1. multi-/many-core architecture, 2. virtualization software, and, 3. special self-test techniques such as BIST (Built-In Self-Test) or CASP (Concurrent Autonomous chip self-test using Stored Patterns). As a result, optimized trade-offs in system design complexity, system performance and power impact, and test thoroughness are possible. Experimental results from an actual multi-core system demonstrate that: 1. VAST is practical and effective; and, 2. Special VAST-supported self-test policies enable extremely thorough on-line self-test with very small performance impact.
Hiroaki Inoue, Yanjing Li, Subhasish Mitra
ITC1
2008 FIDES: An advanced chip multiprocessor platform for secure next generation mobile terminals
abstract
We propose a secure platform on a chip multiprocessor, FIDES, in order to enable next generation mobile terminals to execute downloaded native applications for Linux. Its most important feature is the higher security based on multigrained separation mechanisms. Four new technologies support the FIDES platform: bus filter logic, XIP kernels, policy separation, and dynamic access control. With these technologies, the FIDES platform can tolerate both application-level and kernel-level bugs on an actual download subsystem. Thus, the best-suited platform to secure next generation mobile terminals is FIDES.
Hiroaki Inoue, Junji Sakai, Sunao Torii, Masato Edahiro
ACM Trans. Embed. Comput. Syst.1
2008 Processor virtualization for secure mobile terminals
abstract
We propose a processor virtualization architecture, VIRTUS, to provide a dedicated domain for preinstalled applications and virtualized domains for downloaded native applications. With it, security-oriented next-generation mobile terminals can provide any number of domains for native applications. VIRTUS features three new technologies, namely, VMM asymmetrization, dynamic interdomain communication (IDC), and virtualization-assist logic, and it is first in the world to virtualize an ARM-based multiprocessor. Evaluations have shown that VMM asymmetrization results in significantly less performance degradation and LOC increase than do other VMMs. Further, dynamic IDC overhead is low enough, and virtualization-assist logic can be implemented in a sufficiently small area.
Hiroaki Inoue, Junji Sakai, Masato Edahiro
ACM Trans. Design Autom. Electr. Syst.1
2007 Towards scalable and secure execution platform for embedded systems
abstract
Reliability of embedded systems can be enhanced by multicore and partitioning approaches. Physical partitioning based on AMP multicore achieves runtime stability of multiple applications in a system and prevents the whole system shutdown as well even when a malicious code creeps in. Combined with logical partitioning by processor visualization and SMP technologies, the multicore architecture could realize more flexible and more scalable platform for future embedded systems.
Hiroaki Inoue, Masato Edahiro, Junji Sakai
ASP-DAC1
2006 VIRTUS: a new processor virtualization architecture for security-oriented next-generation mobile terminals
abstract
We propose a new processor virtualization architecture, VIRTUS, to provide a dedicated domain for pre-installed applications and virtualized domains for downloaded native applications. With it, security-oriented next-generation mobile terminals can provide any number of domains for native applications. VIRTUS features three new technologies: VMM asymmetrization, dynamic inter-domain communication and virtualization-assist logic, and it is first in the world to virtualize an ARM-based multiprocessor.
Hiroaki Inoue, Akihisa Ikeno, Masaki Kondo, Junji Sakai, Masato Edahiro
DAC1
1988 An 8 mm length nonblocking 4×4 optical switch array
abstract
A novel single-mode-single-slip-structure S/sup 3/ optical switch using carrier-induced refractive index change is proposed as a unit cell for a small polarization-independent nonblocking N*N optical switch array. Sixteen S/sup 3/ optical switches have been integrated into a nonblocking 4*4 optical switch array on an InP substrate. The 8-mm-length InGaAsP/InP 4*4 optical array has shown satisfactory switching characteristics and is suitable for larger scale integration of optical switch arrays and also for integration with other active optical devices such as laser diodes.>
Hiroaki Inoue, Hitoshi Nakamura, Kenichi Morosawa, Yoshimitsu Sasaki, Toshio Katsuyama, Naoki Chinone
IEEE J. Sel. Areas Commun.1