EDBT 2026 Demo / reviewers in the wild / expert
Hideho Arakida
dblp:50/6720
· DBLP profile ↗
7ranked-venue papers
0as first author
0since 2021 · last 2012
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6Software engineering, systems software and programming languages · 2Graphics, computer vision, multimedia, augmented reality and games · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
3 papers |
Memory systems · 72% Processor architecture and microarchitecture · 15% Energy-efficient computing · 7% |
Topics — the 11 heaviest of 11, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Memory systems
on-chip memory |
0.2 | 2 | 2008 | Comparative evaluation of memory models for chip multiprocessors · ACM Trans. Archit. Code Optim. 2008 Comparing memory systems for chip multiprocessors · ISCA 2007 |
Processor architecture and microarchitecture
chip multiprocessor |
0.1 | 2 | 2008 | Comparing memory systems for chip multiprocessors · ISCA 2007 Comparative evaluation of memory models for chip multiprocessors · ACM Trans. Archit. Code Optim. 2008 |
Memory systems
cache coherence |
0.1 | 1 | 2008 | Comparative evaluation of memory models for chip multiprocessors · ACM Trans. Archit. Code Optim. 2008 |
Memory systems › memory access optimization
memory streaming |
0.1 | 1 | 2008 | Comparative evaluation of memory models for chip multiprocessors · ACM Trans. Archit. Code Optim. 2008 |
Memory systems › memory management
software-managed memory |
0.1 | 1 | 2008 | Comparative evaluation of memory models for chip multiprocessors · ACM Trans. Archit. Code Optim. 2008 |
Memory systems › on-chip memory
on-chip memory design |
0.1 | 1 | 2007 | Comparing memory systems for chip multiprocessors · ISCA 2007 |
Parallel and multicore computing › parallel programming models
stream programming |
0.0 | 1 | 2007 | Comparing memory systems for chip multiprocessors · ISCA 2007 |
Energy-efficient computing › voltage scaling
adaptive voltage scaling |
0.0 | 1 | 1998 | Design Methodology of Ultra Low-Power MPEG4 Codec Core Exploiting Voltage Scaling Techniques · DAC 1998 |
Integrated circuit design
low-power circuit design |
0.0 | 1 | 1998 | Design Methodology of Ultra Low-Power MPEG4 Codec Core Exploiting Voltage Scaling Techniques · DAC 1998 |
Energy-efficient computing
voltage scaling |
0.0 | 1 | 1998 | Design Methodology of Ultra Low-Power MPEG4 Codec Core Exploiting Voltage Scaling Techniques · DAC 1998 |
Energy-efficient computing
power management |
0.0 | 1 | 1998 | Design Methodology of Ultra Low-Power MPEG4 Codec Core Exploiting Voltage Scaling Techniques · DAC 1998 |
Methods — techniques the papers use, named apart from their topics
simulation · 0.2performance evaluation · 0.1performance comparison · 0.1voltage scaling · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2012 | Visconti2 - a heterogeneous multi-core SoC for image-recognition applications
Masato Uchiyama, Hideho Arakida, Yasuki Tanabe, Tsukasa Ike, Takanori Tamai, Moriyasu Banno |
Hot Chips Symposium | 2 |
| 2009 | Design and implementation of scalable, transparent threads for multi-core media processorabstractIn this paper, we propose a scalable and transparent parallelization scheme using threads for multi-core processor. The performance achieved by our scheme is scalable to the number of cores, and the application program is not affected by the actual number of cores. For the performance efficiency, we designed the threads so that they do not suspend and that they do not start their execution until the data necessary for them are available. We implemented our design using three modules: the dependency controller, which controls dependencies among threads, the thread pool, which manages the ready threads, and the thread dispatcher, which fetches threads from the pool and executes them on the core. Our design and implementation provide efficient thread scheduling with low overhead. Moreover, by hiding the actual number of cores, it realizes transparency. We confirmed the transparency and scalability of our scheme by applying it to the H.264 decoder program. With this scheme, modification of application program is not necessary even if the number of cores changes due to disparate requirements. This feature makes the developing time shorter and contributes to the reduction of the developing cost. Takeshi Kodaka, Shunsuke Sasaki, Takahiro Tokuyoshi, Ryuichiro Ohyama, Nobuhiro Nonogaki, Koji Kitayama, Yasuyuki Ueda, Hideho Arakida, Yuji Okuda, Toshiki Kizu, Yoshiro Tsuboi, Nobu Matsumoto |
DATE | 9 |
| 2008 | Comparative evaluation of memory models for chip multiprocessorsabstractThere are two competing models for the on-chip memory in Chip Multiprocessor (CMP) systems: hardware-managed coherent caches and software-managed streaming memory . This paper performs a direct comparison of the two models under the same set of assumptions about technology, area, and computational capabilities. The goal is to quantify how and when they differ in terms of performance, energy consumption, bandwidth requirements, and latency tolerance for general-purpose CMPs. We demonstrate that for data-parallel applications on systems with up to 16 cores, the cache-based and streaming models perform and scale equally well. For certain applications with little data reuse, streaming scales better due to better bandwidth use and macroscopic software prefetching. However, the introduction of techniques such as hardware prefetching and nonallocating stores to the cache-based model eliminates the streaming advantage. Overall, our results indicate that there is not sufficient advantage in building streaming memory systems where all on-chip memory structures are explicitly managed. On the other hand, we show that streaming at the programming model level is particularly beneficial, even with the cache-based model, as it enhances locality and creates opportunities for bandwidth optimizations. Moreover, we observe that stream programming is actually easier with the cache-based model because the hardware guarantees correct, best-effort execution even when the programmer cannot fully regularize an application's code. Jacob Leverich, Hideho Arakida, Alex Solomatnikov, Amin Firoozshahian, Mark Horowitz, Christoforos E. Kozyrakis |
ACM Trans. Archit. Code Optim. | 2 |
| 2007 | Comparing memory systems for chip multiprocessorsabstractThere are two basic models for the on-chip memory in CMP systems:hardware-managed coherent caches and software-managed streaming memory. This paper performs a direct comparison of the two modelsunder the same set of assumptions about technology, area, and computational capabilities. The goal is to quantify how and when they differ in terms of performance, energy consumption, bandwidth requirements, and latency tolerance for general-purpose CMPs. We demonstrate that for data-parallel applications, the cache-based and streaming models perform and scale equally well. For certain applications with little data reuse, streaming scales better due to better bandwidth use and macroscopic software prefetching. However, the introduction of techniques such as hardware prefetching and non-allocating stores to the cache-based model eliminates the streaming advantage. Overall, our results indicate that there is not sufficient advantage in building streaming memory systems where all on-chip memory structures are explicitly managed. On the other hand, we show that streaming at the programming model level is particularly beneficial, even with the cache-based model, as it enhances locality and creates opportunities for bandwidth optimizations. Moreover, we observe that stream programming is actually easier with the cache-based model because the hardware guarantees correct, best-effort execution even when the programmer cannot fully regularize an application's code. Jacob Leverich, Hideho Arakida, Alex Solomatnikov, Amin Firoozshahian, Mark Horowitz, Christoforos E. Kozyrakis |
ISCA | 2 |
| 2001 | A Single-Chip Low-Power Mpeg-4 Audiovisual Lsi Using Embedded Dram TechnologyabstractA single-chip MPEG-4 audiovisual LSI based on the proposed scalable multiprocessor architecture has been developed for IMT-2000 multimedia applications. The LSI consists of three 16-bit multimedia-extended RISC processors and dedicated hardware accelerators, so as to achieve both low power consumption and high cost-effectiveness. It handles the MPEG- 4 video SP@L1 codec with the QCIF image at 15 frames per second, the AMR speech codec, and the ITU-T H.223 multiplexing at 60MHz consuming only 80mW, which is 33% of the previous design. The MPEG-4 audiovisual LSI was fabricated in 0.18us CMOS technology with quad metal using the embedded DRAM. Masafumi Takahashi, Tsuyoshi Nishikawa, Hideho Arakida, Tohru Furuyama |
ICME | 3 |
| 2000 | A scalable MPEG-4 video codec architecture for IMT-2000 multimedia applicationsabstractA scalable MPEG-4 video codec architecture is proposed to achieve low power consumption and high cost-effectiveness for IMT-2000 multimedia applications. The MPEG-4 video codec consists of a 16-bit multimedia-extended RISC processor and dedicated hardware accelerators, which bring about both low power consumption and programmability. The proposed architecture is extended and applied for the development of two MPEG-4 LSIs. One is an MPEG-4 video codec LSI, which performs an MPEG-4 video encoding and decoding at 15 frames per second with quarter common intermediate format. The other is an MPEG-4 audiovisual LSI, containing three 16-bit RISC processors and a 16-Mbit embedded DRAM, executes the major functions of 3GPP 3G-324M video telephony for IMT-2000 applications. By introducing the optimization of the embedded DRAM configuration, clock gating technique, and low power motion estimation, the MPEG-4 audiovisual LSI consumes only 240 mW when it activates MPEG-4 video SP@L1 codec, the AMR speech codec, and the H.223 annex B multiplex at 60 MHz clock rate. Masafumi Takahashi, Tsuyoshi Nishikawa, Hideho Arakida, Noriaki Machida, Hideaki Yamamoto, Toshihide Fujiyoshi, Yoko Matsumoto, Osamu Yamagishi, Tatsuo Samata, Atsushi Asano, Toshihiro Terazawa, Kenji Ohmori, Junya Shirakura, Yoshinori Watanabe, Hiroki Nakamura, Shigenobu Minami, Tohru Furuyama |
ISCAS | 3 |
| 1998 | Design Methodology of Ultra Low-Power MPEG4 Codec Core Exploiting Voltage Scaling TechniquesabstractThis paper describes a fully automated low-power design methodology in which three different voltage-scaling techniques are combined together. Supply voltage is scaled globally, selectively, and adaptively while keeping the performance. This methodology enabled us to design an MPEG4 codec core with 58% less power than the original in three week turn-around-time. Kimiyoshi Usami, Mutsunori Igarashi, Takashi Ishikawa, Masahiro Kanazawa, Masafumi Takahashi, Mototsugu Hamada, Hideho Arakida, Toshihiro Terazawa, Tadahiro Kuroda |
DAC | 7 |