Shinji Tomita

dblp:84/6370 · DBLP profile ↗
← Back
17ranked-venue papers
3as first author
0since 2021 · last 2012
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 11 · 2 first-authorSoftware engineering, systems software and programming languages · 4 · 2 first-authorArtificial intelligence and machine learning · 2 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2Human-computer interaction and ubiquitous computing · 2Applied, interdisciplinary, general and emerging computing · 2

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
3 papers
Computer animation and physical simulation · 48% Geometric modeling and processing · 48% Rendering · 4%
Computer architecture, parallel and distributed computing, and storage systems
7 papers
Processor architecture and microarchitecture · 82% Memory systems · 15% Parallel and multicore computing · 2%

Topics — the 20 heaviest of 21, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Geometric modeling and processing › shape modeling › shape editing
mesh editing
0.112008
Vertex-preserving Cutting of Elastic Objects · VR 2008
Computer animation and physical simulation › deformable body simulation
soft tissue simulation
0.112008
Vertex-preserving Cutting of Elastic Objects · VR 2008
Processor architecture and microarchitecture › instruction scheduling
dynamic instruction scheduling
0.012001
A high-speed dynamic instruction scheduling scheme for superscalar processors · MICRO 2001
Processor architecture and microarchitecture
instruction scheduling
0.012001
A high-speed dynamic instruction scheduling scheme for superscalar processors · MICRO 2001
Processor architecture and microarchitecture
instruction-level parallelism
0.031989
SIMP (Single Instruction stream/Multiple Instruction Pipelining): A Novel High-Speed Single-Processor Architecture · ISCA 1989
A Computer with Low-Level Parallelism QA-2: Its Applications to 3-D Graphics and Prolog/Lisp Machines · ISCA 1986
A User-Microprogrammable, Local Host Computer With Low-Level Parallelism · ISCA 1983
Memory systems
cache coherence
0.011993
A distributed shared memory multiprocessor ASURA: memory and cache architecture · SC 1993
Memory systems › shared memory
distributed shared memory
0.011993
A distributed shared memory multiprocessor ASURA: memory and cache architecture · SC 1993
Processor architecture and microarchitecture
multiprocessor architecture
0.011993
A distributed shared memory multiprocessor ASURA: memory and cache architecture · SC 1993
Processor architecture and microarchitecture
superscalar processor
0.012001
A high-speed dynamic instruction scheduling scheme for superscalar processors · MICRO 2001
Processor architecture and microarchitecture
out-of-order execution
0.011989
SIMP (Single Instruction stream/Multiple Instruction Pipelining): A Novel High-Speed Single-Processor Architecture · ISCA 1989
Processor architecture and microarchitecture
special-purpose processor
0.011986
A Computer with Low-Level Parallelism QA-2: Its Applications to 3-D Graphics and Prolog/Lisp Machines · ISCA 1986
Rendering
hidden surface removal
0.011984
A parallel processor system for three-dimensional color graphics · SIGGRAPH 1984
Rendering › hidden surface removal
scan-line algorithms
0.011984
A parallel processor system for three-dimensional color graphics · SIGGRAPH 1984
Processor architecture and microarchitecture › microprogramming
microprogrammable processor
0.011983
A User-Microprogrammable, Local Host Computer With Low-Level Parallelism · ISCA 1983
Processor architecture and microarchitecture
branch prediction
0.011989
SIMP (Single Instruction stream/Multiple Instruction Pipelining): A Novel High-Speed Single-Processor Architecture · ISCA 1989
Processor architecture and microarchitecture
microprogramming
0.011980
A Dynamically Microprogammable Computer with Low-Level Parallelism · IEEE Trans. Computers 1980
Parallel and multicore computing
parallel computing
0.011980
A Dynamically Microprogammable Computer with Low-Level Parallelism · IEEE Trans. Computers 1980
High-performance computing
scientific computing systems
0.011983
A User-Microprogrammable, Local Host Computer With Low-Level Parallelism · ISCA 1983
Processor architecture and microarchitecture
instruction set architecture
0.011980
A Dynamically Microprogammable Computer with Low-Level Parallelism · IEEE Trans. Computers 1980
Parallel and multicore computing › parallel architecture
MIMD architecture
0.011980
A Dynamically Microprogammable Computer with Low-Level Parallelism · IEEE Trans. Computers 1980

Methods — techniques the papers use, named apart from their topics

topological update scheme · 0.1finite element method · 0.1dynamic scheduling scheme · 0.0simulation · 0.0hierarchical directory · 0.0compiler optimization · 0.0ALU chaining · 0.0tomasulo's algorithm · 0.0speculative execution · 0.0raster-scan display · 0.0hierarchical multiprocessor · 0.0microprogramming · 0.0
YearPublicationVenuePosition
2012 Prototype Implementation of a GPU-based Interactive Coupled Fluid-Structure Simulation
abstract
This paper reports a prototype implementation of 3D Fluid-Structure Simulation. In this implementation, heartbeat fluid flows inside a rectangular tube where the operator can interactively change the width of the tube. In order to achieve interactive simulation speed, some parts of the simulation is accelerated by GPU. For further improvement of the simulation speed, two techniques have been examined: the first one reduces the communication between CPU and GPU, and the second one offloads a whole fluid flow simulation onto GPU.
Ryota Henmi, Yusuke Nishimura, Hiroaki Suzuki, Shinji Fukuma, Shin-ichiro Mori, Akinori Yamaguchi, Shinji Tomita
SNPD7
2011 A Fine-Grained Runtime Power/Performance Optimization Method for Processors with Adaptive Pipeline Depth
Jun Yao 0001, Shinobu Miwa, Hajime Shimada, Shinji Tomita
J. Comput. Sci. Technol.4
2008 Vertex-preserving Cutting of Elastic Objects
abstract
This paper proposes vertex-preserving cutting methods on finite element models for interactive soft tissue simulation. Unlike existing methods, we aim to shape variety of incisions using only initial vertices of tetrahedral meshes. Neither tetrahedral decomposition nor vertex creation is used. The number of vertices is preserved. This avoids increase of computation cost as well as allows fast update of physical status of finite element models. To preserve 3D shape and sharp feature of initial meshes through on-the-fly mesh modification, constraints are introduced to the topological update scheme. In our model, the size of stiffness matrix is constant. Our framework efficiently simulates several varieties of smooth incisions with sufficient quality for surgical simulation, and also achieves interactive performance in complex meshes with thousands of elements.
Megumi Nakao, Kotaro Minato, Naoto Kume, Shin-ichiro Mori, Shinji Tomita
VR5
2003 Bilateral Tradings with and without Strategic Thinking
Shinji Tomita, Akira Namatame
MABS1
2002 Development of an Education Tool for Computer System
abstract
An education tool has been developed to explain internal behavior and structure of computer graphically. It is written in Java Language and designed for students to understand how computer works in the classroom lecture of information science. This paper describes some characteristics of our education tool and improvement of its facilities, especially embedded mail handling module for information exchange between teacher and students.
Yoshiro Imai, Shinji Tomita, Haruo Niimi, Hitoshi Inomo, Wataru Shiraki
ICCE2
2001 A high-speed dynamic instruction scheduling scheme for superscalar processors
Masahiro Goshima, Kengo Nishino, Toshiaki Kitamura, Yasuhiko Nakashima, Shinji Tomita, Shin-ichiro Mori
MICRO5
1996 Amon2: A Parallel Wire Routing Algorithm on a Torus Network Parallel Computer
Hesham Keshk, Shin-ichiro Mori, Hiroshi Nakashima, Shinji Tomita
International Conference on Supercomputing4
1995 A proposal of self-cleanup cache
Shin-ichiro Mori, Masahiro Goshima, Hiroshi Nakashima, Shinji Tomita
PACT4
1995 Amon: A Parallel Slice Algorithm for Wire Routing
abstract
In thk paper a new parallel global and detailed wire routing algorithm called "Amen" is introduced.Both of them are done in parallel using different processor elements (PEs).We introduce a new way of dividing the multilayer grid into layers, and dividing layers into slices.Each PE, in the detailed routing, has a responsibility for one or more slice.This way of division obtains a high degree of parallelism.A new detailed routing algorithm which routes nets under the condition that these net paths do not prevent other nets from being routed later is also introduced.Amen gives high routing quality by using fewer vias and shorter wire length than the other algorithms.The results show that this algorithm obtains higher connection ratio than maze running algorithm.
Hesham Keshk, Shin-ichiro Mori, Hiroshi Nakashima, Shinji Tomita
International Conference on Supercomputing4
1993 A distributed shared memory multiprocessor ASURA: memory and cache architecture
abstract
ASURA is a large scale, chwier-based, distributed, shared memory, multiprocessor being developed at Kyoto University and Kubota Corporation.Up to 128 clusters are interconnected to form an ASURA system of up to 102d processors.The basic concept of the ASURA design is to take advantage of the hierarchical structure of the system.Implementing this concept, a large shared cache is piaced between each cluster and the inter-ciuster network.The shared cache and the shared memories distributed among the clusters form part of ASURA's hierarchical memory architecture, providing various unique features to ASURA.In this paper, the hierarchicalmemory architecture of ASURA and its unique cache coherence scheme, including a proposal of a new hierarchical directory scheme, are described wzth some simulation results.Permission tocopy without fes all or pan of Ibis .n81e&al is granted, pmvidd that tbe copies me not made or dishiluted fw dim commenial advantage, the ACMwpyrisbtmalice and tbe title of the F.Iblication and its date appesr, snd notice is given that cqying is by permission of tbe Association forComputing Mschimy.ToCOPY ciherwi%csto republih.requires a fee Snlvcu Specilic pmnissian.
Shin-ichiro Mori, Hideki Saito 0001, Masahiro Goshima, Mamoru Yanagihara, Takashi Tanaka, David Fraser, Kazuki Joe, Hiroyuki Nitta, Shinji Tomita
SC9
1992 Benchmarking a vector-processor prototype based on multithreaded streaming/FIFO vector (MSFV) architecture
abstract
This paper presents the benchmark results on a vector-processor prototype based on the MSFV (multithreaded streaming/FIFO vector) architecture. The MSFV architecture is single-chip oriented, and thus its main object is to save the off-chip memory bandwidth by exploiting the register bandwidth instead. The register bandwidth is exploited by the synergism of FIFO register, chaining, streaming, and multithreading. This paper tries to identify the strength and weakness of those architectural features. The results for basic vector operations and Livermore Fortran Kernels are reported in terms of normalized FLOPC (floating-point operations per clock cycle) and compared to previously-reported results on the Cray X-MP, Y-MP, Fujitsu VP-200, Hitachi S-810/20, NEC SX-2, and SX-3. These comparisons show that, for many basic vector operations, the execution rate of the MSFV prototype results in worst due to its saving the memory bandwidth. However, for Livermore Fortran Kernels, the MSFV prototype results in worst due to its saving the memory bandwidth. However, for Livermore Fortran Kernels, the MSFV prototype outperforms the VP-200 by 2.11 times (geometric mean) and one processor of the X-MP by 1.22 times (geometric mean) in terms of FLOPC. Also, it is 0.67 times (geometric mean) faster than the S-810/20, and 0.76 times (geometric mean) faster than the SX-2. The paper concludes that the MSFV architecture is successful in saving the memory bandwidth.
Tetsuo Hironaka, Takashi Hashimoto, Keizo Okazaki, Kazuaki J. Murakami, Shinji Tomita
ICS5
1989 The Kyushu University reconfigurable parallel processor: design of memory and intercommunicaiton architectures
abstract
The reconfigurable parallel processor system under development at Kyushu University is an MIMD-type multiprocessor which consists of N processing-elements (currently N is 128) fully connected by S N × N crossbar networks (currently S is 1). Each PE (Processing Element) employs a Fujitsu SPARC MB86900/10 chip-set, a Weitek WTL1164/65 chip-set, an MMU (Memory Management Unit) with 64K bytes of cache, 4M bytes of memory, and an MCU (Message Communication Unit). The modular 128 × 128 crossbar network is implemented by arranging 256 identical 8 × 8 crossbar LSI-modules in a 16 × 16 matrix form. The full 128-PE configuration achieves supercomputer levels of performance by providing 1.28 GIPS and 205 MFLOPS of computing power, 512M bytes of memory, and 2.56G bytes/s of inter-PE communication bandwidth. At the same time, it exploits unique reconfigurability in the memory and intercommunication architectures. By utilizing these two types of reconfigurability, we believe that the system can be effectively tailored to a wide spectrum of applications such as numerical computation, image processing, computer graphics, artificial intelligence, neurocomputing, and so on.
Kazuaki J. Murakami, Shin-ichiro Mori, Akira Fukuda, Toshinori Sueyoshi, Shinji Tomita
ICS5
1989 SIMP (Single Instruction stream/Multiple Instruction Pipelining): A Novel High-Speed Single-Processor Architecture
abstract
SIMP is a novel multiple instruction-pipeline parallel architecture. It is targeted for enhancing the performance of SISD processors drastically by exploiting both temporal and spatial parallelisms, and for keeping program compatibility as well. Degree of performance enhancement achieved by SIMP depends on; i) how to supply multiple instructions continuously, and ii) how to resolve data and control dependencies effectively. We have devised the outstanding techniques for instruction fetch and dependency resolution. The instruction fetch mechanism employs unique schemes of; i) prefetching multiple instructions with the help of branch prediction, ii) squashing instructions selectively, and iii) providing multiple conditional modes as a result. The dependency resolution mechanism permits out-of-order execution of sequential instruction stream. Our out-of-order execution model is based on Tomasulo's algorithm which has been used in single instruction-pipeline processors. However, it is greatly extended and accommodated to multiple instruction pipelining with; i) detecting and identifying multiple dependencies simultaneously, ii) alleviating the effects of control dependencies with both eager execution and advance execution, and iii) ensuring a precise machine state against branches and interrupts. By taking advantage of these techniques, SIMP is one of the most promising architectures toward the coming generation of high-speed single processors.
Kazuaki J. Murakami, Naohiko Irie, Morihiro Kuga, Shinji Tomita
ISCA4
1986 A Computer with Low-Level Parallelism QA-2: Its Applications to 3-D Graphics and Prolog/Lisp Machines
abstract
We proposed a computer with low-level parallelism as one of the basic computer architectures and built a large scale experimental system called QA-2. By low-level parallelism, we mean that a long-word instruction controls simultaneously many ALUs, busses, registers and memories in a mode of fine-grained parallelism. The QA-2 employs a 256-bit instruction by which four different ALU operations, four memory accesses to different/continuous locations and one powerful sequence control are all specified and performed in parallel. If many simultaneously executable operations are detected and embedded in one instruction at compile time, this type of computer can provide a high-degree of performance for a wide variety of applications. This paper describes the architectural benefits and limitations of low-level parallelism in performing 3-D color image generation and interpreting Prolog/Lisp programs. The hardware organization with four ALUs, which are actually implemented in the QA-2, is verified to be adequate. In fact, nearly three out of four ALUs can work in parallel. Any architecture with more than four ALUs can not achieve a significant degree of performance enhancement. This paper also shows the degree of performance improvement achieved by the techniques such as ALU chaining and highly-structured sequence control mechanisms. As compared with the IBM 370 architecture, the QA-2 can generate 3-D color images in 1/5 of dynamic instruction steps. The compiler version of Prolog machine on the QA-2 is as fast (45K LIPS) as the ICOT's PSI. From all results, we expect that the QA-2 is a high-performance computer which will be utilized in the future personal computing environment.
Shinji Tomita, Kiyoshi Shibayama, Toshiyuki Nakata 0001, Shinji Yuasa, Hiroshi Hagiwara
ISCA1
1984 A parallel processor system for three-dimensional color graphics
abstract
This paper describes the hardware architecture and the employed algorithm of a parallel processor system for three-dimensional color graphics. The design goal of the system is to generate realistic images of three-dimensional environments on a raster-scan video display in real-time. In order to achieve this goal, the system is constructed as a two-level hierarchical multi-processor system which is particularly suited to incorporate scan-line algorithm for hidden surface elimination. The system consists of several Scan-Line Processors (SLPs), each of which controls several slave PiXel Processors (PXPs). The SLP prepares the specific data structure relevant to each scan line, while the PXP manipulates every pixel data in its own territory. Internal hardware structures of the SLP and the PXP are quite different, being designed for their dedicated tasks.
Haruo Niimi, Yoshiro Imai, Masayoshi Murakami, Shinji Tomita, Hiroshi Hagiwara
SIGGRAPH4
1983 A User-Microprogrammable, Local Host Computer With Low-Level Parallelism
abstract
This paper describes the architecture of a dynamically microprogrammable computer with low-level parallelism, called QA-2, which is designed as a high-performance, local host computer for laboratory use. The architectural principle of the QA-2 is the marriage of high-speed, parallel processing capability offered by four powerful Arithmetic and Logic Units (ALUs) with architectural flexibility provided by large scale, dynamic user-microprogramming. By changing its writable control storage dynamically, the QA-2 can be tailored to a wide spectrum of research-oriented applications covering high-level language processing and real-time processing.
Shinji Tomita, Kiyoshi Shibayama, Toshiaki Kitamura, Toshiyuki Nakata 0001, Hiroshi Hagiwara
ISCA1
1980 A Dynamically Microprogammable Computer with Low-Level Parallelism
abstract
A new microprogrammable computer with low-level parallelism was built and has been utilized as a research vehicle for solving different classes of research-oriented applications such as real-time processings on static/dynamic images, pictures and signals, and emulations of both existing and virtual machines including high (intermediate) level language machines. The design goal of the machine was to achieve a high degree of processing enhancement in research- oriented applications by means of a low-level parallel processing organization combined with dynamically microprogrammable control. The machine has the capability to process multiple data streams, performing parallel operations with four 16-bit ALU's. These ALU's are independently controlled by the different fields of a 160-bit horizontal-type microinstruction, and have simultaneous access to 15 working registers. This microprogrammed MIMD organization is expected to provide a greater degree of flexibility for low-level parallel processing. In addition, not only does the machine contain powerful ALU's and a large number of registers, but also it employs flexible control structures and a hierarchical organization of control storage. All of these combine to yield extensive microprogramming capability which the user can effectively tailor to a wide spectrum of applications.
Hiroshi Hagiwara, Shinji Tomita, Shigeru Oyanagi, Kiyoshi Shibayama
IEEE Trans. Computers2