VLDB 2026 Research / reviewers in the wild / expert
Masahiro Sowa
dblp:15/6694
· DBLP profile ↗
14ranked-venue papers
3as first author
0since 2021 · last 2009
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 2 first-authorSoftware engineering, systems software and programming languages · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Processor architecture and microarchitecture · 50% Memory systems · 38% Parallel and multicore computing · 12% |
Topics — the 4 heaviest of 4, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Processor architecture and microarchitecture
dataflow architecture |
0.0 | 1 | 1982 | A Data Flow Computer Architecture with Program and Token Memories · IEEE Trans. Computers 1982 |
Memory systems
memory architecture |
0.0 | 1 | 1982 | A Data Flow Computer Architecture with Program and Token Memories · IEEE Trans. Computers 1982 |
Processor architecture and microarchitecture › microprocessor design › processor core design
functional units |
0.0 | 1 | 1982 | A Data Flow Computer Architecture with Program and Token Memories · IEEE Trans. Computers 1982 |
Parallel and multicore computing
parallel computing |
0.0 | 1 | 1982 | A Data Flow Computer Architecture with Program and Token Memories · IEEE Trans. Computers 1982 |
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2009 | Efficient compilation for queue size constrained queue processors
Arquimedes Canedo, Ben A. Abderazek, Masahiro Sowa |
Parallel Comput. | 3 |
| 2008 | A new code generation algorithm for 2-offset producer order queue computation model
Arquimedes Canedo, Ben A. Abderazek, Masahiro Sowa |
Comput. Lang. Syst. Struct. | 3 |
| 2008 | The QC-2 parallel Queue processor architecture
Ben A. Abderazek, Arquimedes Canedo, Tsutomu Yoshinaga, Masahiro Sowa |
J. Parallel Distributed Comput. | 4 |
| 2008 | Dual-execution mode processor architecture
Md. Musfiquzzaman Akanda, Ben A. Abderazek, Masahiro Sowa |
J. Supercomput. | 3 |
| 2007 | An Efficient Code Generation Algorithm for Code Size Reduction Using 1-Offset P-Code Queue Computation Model
Arquimedes Canedo, Ben A. Abderazek, Masahiro Sowa |
EUC | 3 |
| 2007 | New Code Generation Algorithm for QueueCore - An Embedded Processor with High ILPabstractModern architectures rely on exploiting parallelism found at the instruction level to achieve high performance. Aggressive ILP compilers expose high amounts of instruction level parallelism where, in some cases, the number of architected registers is not enough to hold the results of potential parallel instructions. This paper presents a new code generation scheme for the QueueCore, a 32-bit queue-based architecture capable of executing high amounts of ILP. QueueCore's instructions implicitly read their operands and write results. Compiling for the QueueCore requires that all instructions have at most one explicit operand represented as an offset calculated at compile-time. Additionally, the instructions must be scheduled in level-order manner. The proposed algorithm successfully restricts all instructions to have at most one offset reference, it computes the offset values, and makes a level-order scheduling of the program. To evaluate the effectiveness of the new code generation scheme we developed a queue compiler and compiled a set of benchmark programs. Our results show that the code has more parallelism than optimized RISC code by factors ranging from 1.12 to 2.30. QueueCore's instruction set allows us to generate code about 40%-18% denser than optimized RISC code. Arquimedes Canedo, Ben A. Abderazek, Masahiro Sowa |
PDCAT | 3 |
| 2007 | Queue Register File Optimization Algorithm for QueueCore ProcessorabstractThe queue computation model offers an attractive alternative for high-performance embedded computing given its characteristics of short instructions and high instruction level parallelism. A queue-based processor uses a FIFO queue to read and write operands through hardware pointers located at the head and tail of the queue. Queue length is the number of elements stored between the head and the tail pointers during computations. We have found that 95% of the statements in integer applications require a queue length of less than 32 words. The remaining 5% requires larger queue length sizes up to 230 queue words. In this paper we propose a compiler technique to optimize the queue utilization for the hungry statements that require a large amount of queue. We show that for SPEC CINT95 benchmarks, our technique optimizes the queue length without decreasing parallelism. However, our optimization has a penalty of a slight increase in code size. Arquimedes Canedo, Ben A. Abderazek, Masahiro Sowa |
SBAC-PAD | 3 |
| 2006 | High-Level Modeling and FPGA Prototyping of Produced Order Parallel Queue Processor Core
Ben A. Abderazek, Tsutomu Yoshinaga, Masahiro Sowa |
J. Supercomput. | 3 |
| 2005 | Modular Design Structure and High-Level Prototyping for Novel Embedded Processor Core
Ben A. Abderazek, Sotaro Kawata, Tsutomu Yoshinaga, Masahiro Sowa |
EUC | 4 |
| 2005 | An Efficient Dynamic Switching Mechanism (DSM) for Hybrid Processor Architecture
Md. Musfiquzzaman Akanda, Ben A. Abderazek, Sotaro Kawata, Masahiro Sowa |
EUC | 4 |
| 2005 | Parallel Queue Processor Architecture Based on Produced Order Computation Model
Masahiro Sowa, Ben A. Abderazek, Tsutomu Yoshinaga |
J. Supercomput. | 1 |
| 2003 | On the Design of a Register Queue Based Processor Architecture (FaRM-rq)
Ben A. Abderazek, Soichi Shigeta, Tsutomu Yoshinaga, Masahiro Sowa |
ISPA | 4 |
| 1987 | A Method for Speeding up Serial Processing in Dataflow Computers by Means of a Program CounterabstractThe introduction of a program counter and an accumulator into a dataflow computer can improve its speed during serial processing. This paper describes a possible implementation together with a simple performance analysis. In dataflow processing, more than one processor can be used to execute parts of a single program in parallel. Performance is good, but some parts of many programs are inherently serial (sequential). Serial execution requires that explicit expressions be used for transferring data between instructions and for link manipulation, which makes serial processing in a dataflow computer slower than in a computer specially designed for the task. However, it is possible to arrange for the serial processing to be carried out by only one processor and for almost all of the instructions to be arranged in the correct order for execution. Link manipulation can then be achieved implicitly by using the program counter of a von Numann computer and one of the data items can be passed through an accumulator in the processor. It seems likely that the introduction of these features into the architecture of a dataflow computer would result in a substantial speeding up of serial processing. Analysis anticipates that the speed can be improved by between 20 and 50%. Masahiro Sowa |
Comput. J. | 1 |
| 1982 | A Data Flow Computer Architecture with Program and Token MemoriesabstractThis paper presents a new data flow computer architecture using two types of memory: program memory and token memory. The program memory (PM) stores a data flow program or graph indicating data dependencies among instructions. The PM need not be changed during execution of one program and is duplicated for each processing or functional unit. The token memory (TM) stores mainly currently needed data values and plays a role similar to accumulators or registers in conventional computers. The size of the TM is expected to be relatively small because the data values occupy space in the TM for a very short time. The small scale TM in our architecture makes it possible to reduce the complexity of the switch mechanism between functional units and memories, and the size of the control circuit. Masahiro Sowa, Tadao Murata |
IEEE Trans. Computers | 1 |