Hideho Arakida

dblp:50/6720 · DBLP profile ↗
← Back
7ranked-venue papers
0as first author
0since 2021 · last 2012
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6Software engineering, systems software and programming languages · 2Graphics, computer vision, multimedia, augmented reality and games · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
3 papers
Memory systems · 72% Processor architecture and microarchitecture · 15% Energy-efficient computing · 7%

Topics — the 11 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Memory systems
on-chip memory
0.222008
Comparative evaluation of memory models for chip multiprocessors · ACM Trans. Archit. Code Optim. 2008
Comparing memory systems for chip multiprocessors · ISCA 2007
Processor architecture and microarchitecture
chip multiprocessor
0.122008
Comparing memory systems for chip multiprocessors · ISCA 2007
Comparative evaluation of memory models for chip multiprocessors · ACM Trans. Archit. Code Optim. 2008
Memory systems
cache coherence
0.112008
Comparative evaluation of memory models for chip multiprocessors · ACM Trans. Archit. Code Optim. 2008
Memory systems › memory access optimization
memory streaming
0.112008
Comparative evaluation of memory models for chip multiprocessors · ACM Trans. Archit. Code Optim. 2008
Memory systems › memory management
software-managed memory
0.112008
Comparative evaluation of memory models for chip multiprocessors · ACM Trans. Archit. Code Optim. 2008
Memory systems › on-chip memory
on-chip memory design
0.112007
Comparing memory systems for chip multiprocessors · ISCA 2007
Parallel and multicore computing › parallel programming models
stream programming
0.012007
Comparing memory systems for chip multiprocessors · ISCA 2007
Energy-efficient computing › voltage scaling
adaptive voltage scaling
0.011998
Design Methodology of Ultra Low-Power MPEG4 Codec Core Exploiting Voltage Scaling Techniques · DAC 1998
Integrated circuit design
low-power circuit design
0.011998
Design Methodology of Ultra Low-Power MPEG4 Codec Core Exploiting Voltage Scaling Techniques · DAC 1998
Energy-efficient computing
voltage scaling
0.011998
Design Methodology of Ultra Low-Power MPEG4 Codec Core Exploiting Voltage Scaling Techniques · DAC 1998
Energy-efficient computing
power management
0.011998
Design Methodology of Ultra Low-Power MPEG4 Codec Core Exploiting Voltage Scaling Techniques · DAC 1998

Methods — techniques the papers use, named apart from their topics

simulation · 0.2performance evaluation · 0.1performance comparison · 0.1voltage scaling · 0.0
YearPublicationVenuePosition
2012 Visconti2 - a heterogeneous multi-core SoC for image-recognition applications
Masato Uchiyama, Hideho Arakida, Yasuki Tanabe, Tsukasa Ike, Takanori Tamai, Moriyasu Banno
Hot Chips Symposium2
2009 Design and implementation of scalable, transparent threads for multi-core media processor
abstract
In this paper, we propose a scalable and transparent parallelization scheme using threads for multi-core processor. The performance achieved by our scheme is scalable to the number of cores, and the application program is not affected by the actual number of cores. For the performance efficiency, we designed the threads so that they do not suspend and that they do not start their execution until the data necessary for them are available. We implemented our design using three modules: the dependency controller, which controls dependencies among threads, the thread pool, which manages the ready threads, and the thread dispatcher, which fetches threads from the pool and executes them on the core. Our design and implementation provide efficient thread scheduling with low overhead. Moreover, by hiding the actual number of cores, it realizes transparency. We confirmed the transparency and scalability of our scheme by applying it to the H.264 decoder program. With this scheme, modification of application program is not necessary even if the number of cores changes due to disparate requirements. This feature makes the developing time shorter and contributes to the reduction of the developing cost.
Takeshi Kodaka, Shunsuke Sasaki, Takahiro Tokuyoshi, Ryuichiro Ohyama, Nobuhiro Nonogaki, Koji Kitayama, Yasuyuki Ueda, Hideho Arakida, Yuji Okuda, Toshiki Kizu, Yoshiro Tsuboi, Nobu Matsumoto
DATE9
2008 Comparative evaluation of memory models for chip multiprocessors
abstract
There are two competing models for the on-chip memory in Chip Multiprocessor (CMP) systems: hardware-managed coherent caches and software-managed streaming memory . This paper performs a direct comparison of the two models under the same set of assumptions about technology, area, and computational capabilities. The goal is to quantify how and when they differ in terms of performance, energy consumption, bandwidth requirements, and latency tolerance for general-purpose CMPs. We demonstrate that for data-parallel applications on systems with up to 16 cores, the cache-based and streaming models perform and scale equally well. For certain applications with little data reuse, streaming scales better due to better bandwidth use and macroscopic software prefetching. However, the introduction of techniques such as hardware prefetching and nonallocating stores to the cache-based model eliminates the streaming advantage. Overall, our results indicate that there is not sufficient advantage in building streaming memory systems where all on-chip memory structures are explicitly managed. On the other hand, we show that streaming at the programming model level is particularly beneficial, even with the cache-based model, as it enhances locality and creates opportunities for bandwidth optimizations. Moreover, we observe that stream programming is actually easier with the cache-based model because the hardware guarantees correct, best-effort execution even when the programmer cannot fully regularize an application's code.
Jacob Leverich, Hideho Arakida, Alex Solomatnikov, Amin Firoozshahian, Mark Horowitz, Christoforos E. Kozyrakis
ACM Trans. Archit. Code Optim.2
2007 Comparing memory systems for chip multiprocessors
abstract
There are two basic models for the on-chip memory in CMP systems:hardware-managed coherent caches and software-managed streaming memory. This paper performs a direct comparison of the two modelsunder the same set of assumptions about technology, area, and computational capabilities. The goal is to quantify how and when they differ in terms of performance, energy consumption, bandwidth requirements, and latency tolerance for general-purpose CMPs. We demonstrate that for data-parallel applications, the cache-based and streaming models perform and scale equally well. For certain applications with little data reuse, streaming scales better due to better bandwidth use and macroscopic software prefetching. However, the introduction of techniques such as hardware prefetching and non-allocating stores to the cache-based model eliminates the streaming advantage. Overall, our results indicate that there is not sufficient advantage in building streaming memory systems where all on-chip memory structures are explicitly managed. On the other hand, we show that streaming at the programming model level is particularly beneficial, even with the cache-based model, as it enhances locality and creates opportunities for bandwidth optimizations. Moreover, we observe that stream programming is actually easier with the cache-based model because the hardware guarantees correct, best-effort execution even when the programmer cannot fully regularize an application's code.
Jacob Leverich, Hideho Arakida, Alex Solomatnikov, Amin Firoozshahian, Mark Horowitz, Christoforos E. Kozyrakis
ISCA2
2001 A Single-Chip Low-Power Mpeg-4 Audiovisual Lsi Using Embedded Dram Technology
abstract
A single-chip MPEG-4 audiovisual LSI based on the proposed scalable multiprocessor architecture has been developed for IMT-2000 multimedia applications. The LSI consists of three 16-bit multimedia-extended RISC processors and dedicated hardware accelerators, so as to achieve both low power consumption and high cost-effectiveness. It handles the MPEG- 4 video SP@L1 codec with the QCIF image at 15 frames per second, the AMR speech codec, and the ITU-T H.223 multiplexing at 60MHz consuming only 80mW, which is 33% of the previous design. The MPEG-4 audiovisual LSI was fabricated in 0.18us CMOS technology with quad metal using the embedded DRAM.
Masafumi Takahashi, Tsuyoshi Nishikawa, Hideho Arakida, Tohru Furuyama
ICME3
2000 A scalable MPEG-4 video codec architecture for IMT-2000 multimedia applications
abstract
A scalable MPEG-4 video codec architecture is proposed to achieve low power consumption and high cost-effectiveness for IMT-2000 multimedia applications. The MPEG-4 video codec consists of a 16-bit multimedia-extended RISC processor and dedicated hardware accelerators, which bring about both low power consumption and programmability. The proposed architecture is extended and applied for the development of two MPEG-4 LSIs. One is an MPEG-4 video codec LSI, which performs an MPEG-4 video encoding and decoding at 15 frames per second with quarter common intermediate format. The other is an MPEG-4 audiovisual LSI, containing three 16-bit RISC processors and a 16-Mbit embedded DRAM, executes the major functions of 3GPP 3G-324M video telephony for IMT-2000 applications. By introducing the optimization of the embedded DRAM configuration, clock gating technique, and low power motion estimation, the MPEG-4 audiovisual LSI consumes only 240 mW when it activates MPEG-4 video SP@L1 codec, the AMR speech codec, and the H.223 annex B multiplex at 60 MHz clock rate.
Masafumi Takahashi, Tsuyoshi Nishikawa, Hideho Arakida, Noriaki Machida, Hideaki Yamamoto, Toshihide Fujiyoshi, Yoko Matsumoto, Osamu Yamagishi, Tatsuo Samata, Atsushi Asano, Toshihiro Terazawa, Kenji Ohmori, Junya Shirakura, Yoshinori Watanabe, Hiroki Nakamura, Shigenobu Minami, Tohru Furuyama
ISCAS3
1998 Design Methodology of Ultra Low-Power MPEG4 Codec Core Exploiting Voltage Scaling Techniques
abstract
This paper describes a fully automated low-power design methodology in which three different voltage-scaling techniques are combined together. Supply voltage is scaled globally, selectively, and adaptively while keeping the performance. This methodology enabled us to design an MPEG4 codec core with 58% less power than the original in three week turn-around-time.
Kimiyoshi Usami, Mutsunori Igarashi, Takashi Ishikawa, Masahiro Kanazawa, Masafumi Takahashi, Mototsugu Hamada, Hideho Arakida, Toshihiro Terazawa, Tadahiro Kuroda
DAC7