Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Masahiro Sowa

dblp:15/6694 · DBLP profile ↗
← Back
14ranked-venue papers
3as first author
0since 2021 · last 2009
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 2 first-authorSoftware engineering, systems software and programming languages · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Processor architecture and microarchitecture · 50% Memory systems · 38% Parallel and multicore computing · 12%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Processor architecture and microarchitecture
dataflow architecture
0.011982
A Data Flow Computer Architecture with Program and Token Memories · IEEE Trans. Computers 1982
Memory systems
memory architecture
0.011982
A Data Flow Computer Architecture with Program and Token Memories · IEEE Trans. Computers 1982
Processor architecture and microarchitecture › microprocessor design › processor core design
functional units
0.011982
A Data Flow Computer Architecture with Program and Token Memories · IEEE Trans. Computers 1982
Parallel and multicore computing
parallel computing
0.011982
A Data Flow Computer Architecture with Program and Token Memories · IEEE Trans. Computers 1982
YearPublicationVenuePosition
2009 Efficient compilation for queue size constrained queue processors
Arquimedes Canedo, Ben A. Abderazek, Masahiro Sowa
Parallel Comput.3
2008 A new code generation algorithm for 2-offset producer order queue computation model
Arquimedes Canedo, Ben A. Abderazek, Masahiro Sowa
Comput. Lang. Syst. Struct.3
2008 The QC-2 parallel Queue processor architecture
Ben A. Abderazek, Arquimedes Canedo, Tsutomu Yoshinaga, Masahiro Sowa
J. Parallel Distributed Comput.4
2008 Dual-execution mode processor architecture
Md. Musfiquzzaman Akanda, Ben A. Abderazek, Masahiro Sowa
J. Supercomput.3
2007 An Efficient Code Generation Algorithm for Code Size Reduction Using 1-Offset P-Code Queue Computation Model
Arquimedes Canedo, Ben A. Abderazek, Masahiro Sowa
EUC3
2007 New Code Generation Algorithm for QueueCore - An Embedded Processor with High ILP
abstract
Modern architectures rely on exploiting parallelism found at the instruction level to achieve high performance. Aggressive ILP compilers expose high amounts of instruction level parallelism where, in some cases, the number of architected registers is not enough to hold the results of potential parallel instructions. This paper presents a new code generation scheme for the QueueCore, a 32-bit queue-based architecture capable of executing high amounts of ILP. QueueCore's instructions implicitly read their operands and write results. Compiling for the QueueCore requires that all instructions have at most one explicit operand represented as an offset calculated at compile-time. Additionally, the instructions must be scheduled in level-order manner. The proposed algorithm successfully restricts all instructions to have at most one offset reference, it computes the offset values, and makes a level-order scheduling of the program. To evaluate the effectiveness of the new code generation scheme we developed a queue compiler and compiled a set of benchmark programs. Our results show that the code has more parallelism than optimized RISC code by factors ranging from 1.12 to 2.30. QueueCore's instruction set allows us to generate code about 40%-18% denser than optimized RISC code.
Arquimedes Canedo, Ben A. Abderazek, Masahiro Sowa
PDCAT3
2007 Queue Register File Optimization Algorithm for QueueCore Processor
abstract
The queue computation model offers an attractive alternative for high-performance embedded computing given its characteristics of short instructions and high instruction level parallelism. A queue-based processor uses a FIFO queue to read and write operands through hardware pointers located at the head and tail of the queue. Queue length is the number of elements stored between the head and the tail pointers during computations. We have found that 95% of the statements in integer applications require a queue length of less than 32 words. The remaining 5% requires larger queue length sizes up to 230 queue words. In this paper we propose a compiler technique to optimize the queue utilization for the hungry statements that require a large amount of queue. We show that for SPEC CINT95 benchmarks, our technique optimizes the queue length without decreasing parallelism. However, our optimization has a penalty of a slight increase in code size.
Arquimedes Canedo, Ben A. Abderazek, Masahiro Sowa
SBAC-PAD3
2006 High-Level Modeling and FPGA Prototyping of Produced Order Parallel Queue Processor Core
Ben A. Abderazek, Tsutomu Yoshinaga, Masahiro Sowa
J. Supercomput.3
2005 Modular Design Structure and High-Level Prototyping for Novel Embedded Processor Core
Ben A. Abderazek, Sotaro Kawata, Tsutomu Yoshinaga, Masahiro Sowa
EUC4
2005 An Efficient Dynamic Switching Mechanism (DSM) for Hybrid Processor Architecture
Md. Musfiquzzaman Akanda, Ben A. Abderazek, Sotaro Kawata, Masahiro Sowa
EUC4
2005 Parallel Queue Processor Architecture Based on Produced Order Computation Model
Masahiro Sowa, Ben A. Abderazek, Tsutomu Yoshinaga
J. Supercomput.1
2003 On the Design of a Register Queue Based Processor Architecture (FaRM-rq)
Ben A. Abderazek, Soichi Shigeta, Tsutomu Yoshinaga, Masahiro Sowa
ISPA4
1987 A Method for Speeding up Serial Processing in Dataflow Computers by Means of a Program Counter
abstract
The introduction of a program counter and an accumulator into a dataflow computer can improve its speed during serial processing. This paper describes a possible implementation together with a simple performance analysis. In dataflow processing, more than one processor can be used to execute parts of a single program in parallel. Performance is good, but some parts of many programs are inherently serial (sequential). Serial execution requires that explicit expressions be used for transferring data between instructions and for link manipulation, which makes serial processing in a dataflow computer slower than in a computer specially designed for the task. However, it is possible to arrange for the serial processing to be carried out by only one processor and for almost all of the instructions to be arranged in the correct order for execution. Link manipulation can then be achieved implicitly by using the program counter of a von Numann computer and one of the data items can be passed through an accumulator in the processor. It seems likely that the introduction of these features into the architecture of a dataflow computer would result in a substantial speeding up of serial processing. Analysis anticipates that the speed can be improved by between 20 and 50%.
Masahiro Sowa
Comput. J.1
1982 A Data Flow Computer Architecture with Program and Token Memories
abstract
This paper presents a new data flow computer architecture using two types of memory: program memory and token memory. The program memory (PM) stores a data flow program or graph indicating data dependencies among instructions. The PM need not be changed during execution of one program and is duplicated for each processing or functional unit. The token memory (TM) stores mainly currently needed data values and plays a role similar to accumulators or registers in conventional computers. The size of the TM is expected to be relatively small because the data values occupy space in the TM for a very short time. The small scale TM in our architecture makes it possible to reduce the complexity of the switch mechanism between functional units and memories, and the size of the control circuit.
Masahiro Sowa, Tadao Murata
IEEE Trans. Computers1