EDBT 2026 Demo / reviewers in the wild / expert
Shinji Tomita
dblp:84/6370
· DBLP profile ↗
17ranked-venue papers
3as first author
0since 2021 · last 2012
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 11 · 2 first-authorSoftware engineering, systems software and programming languages · 4 · 2 first-authorArtificial intelligence and machine learning · 2 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2Human-computer interaction and ubiquitous computing · 2Applied, interdisciplinary, general and emerging computing · 2
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer graphics and multimedia
3 papers |
Computer animation and physical simulation · 48% Geometric modeling and processing · 48% Rendering · 4% | |
| Computer architecture, parallel and distributed computing, and storage systems
7 papers |
Processor architecture and microarchitecture · 82% Memory systems · 15% Parallel and multicore computing · 2% |
Topics — the 20 heaviest of 21, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Geometric modeling and processing › shape modeling › shape editing
mesh editing |
0.1 | 1 | 2008 | Vertex-preserving Cutting of Elastic Objects · VR 2008 |
Computer animation and physical simulation › deformable body simulation
soft tissue simulation |
0.1 | 1 | 2008 | Vertex-preserving Cutting of Elastic Objects · VR 2008 |
Processor architecture and microarchitecture › instruction scheduling
dynamic instruction scheduling |
0.0 | 1 | 2001 | A high-speed dynamic instruction scheduling scheme for superscalar processors · MICRO 2001 |
Processor architecture and microarchitecture
instruction scheduling |
0.0 | 1 | 2001 | A high-speed dynamic instruction scheduling scheme for superscalar processors · MICRO 2001 |
Processor architecture and microarchitecture
instruction-level parallelism |
0.0 | 3 | 1989 | SIMP (Single Instruction stream/Multiple Instruction Pipelining): A Novel High-Speed Single-Processor Architecture · ISCA 1989 A Computer with Low-Level Parallelism QA-2: Its Applications to 3-D Graphics and Prolog/Lisp Machines · ISCA 1986 A User-Microprogrammable, Local Host Computer With Low-Level Parallelism · ISCA 1983 |
Memory systems
cache coherence |
0.0 | 1 | 1993 | A distributed shared memory multiprocessor ASURA: memory and cache architecture · SC 1993 |
Memory systems › shared memory
distributed shared memory |
0.0 | 1 | 1993 | A distributed shared memory multiprocessor ASURA: memory and cache architecture · SC 1993 |
Processor architecture and microarchitecture
multiprocessor architecture |
0.0 | 1 | 1993 | A distributed shared memory multiprocessor ASURA: memory and cache architecture · SC 1993 |
Processor architecture and microarchitecture
superscalar processor |
0.0 | 1 | 2001 | A high-speed dynamic instruction scheduling scheme for superscalar processors · MICRO 2001 |
Processor architecture and microarchitecture
out-of-order execution |
0.0 | 1 | 1989 | SIMP (Single Instruction stream/Multiple Instruction Pipelining): A Novel High-Speed Single-Processor Architecture · ISCA 1989 |
Processor architecture and microarchitecture
special-purpose processor |
0.0 | 1 | 1986 | A Computer with Low-Level Parallelism QA-2: Its Applications to 3-D Graphics and Prolog/Lisp Machines · ISCA 1986 |
Rendering
hidden surface removal |
0.0 | 1 | 1984 | A parallel processor system for three-dimensional color graphics · SIGGRAPH 1984 |
Rendering › hidden surface removal
scan-line algorithms |
0.0 | 1 | 1984 | A parallel processor system for three-dimensional color graphics · SIGGRAPH 1984 |
Processor architecture and microarchitecture › microprogramming
microprogrammable processor |
0.0 | 1 | 1983 | A User-Microprogrammable, Local Host Computer With Low-Level Parallelism · ISCA 1983 |
Processor architecture and microarchitecture
branch prediction |
0.0 | 1 | 1989 | SIMP (Single Instruction stream/Multiple Instruction Pipelining): A Novel High-Speed Single-Processor Architecture · ISCA 1989 |
Processor architecture and microarchitecture
microprogramming |
0.0 | 1 | 1980 | A Dynamically Microprogammable Computer with Low-Level Parallelism · IEEE Trans. Computers 1980 |
Parallel and multicore computing
parallel computing |
0.0 | 1 | 1980 | A Dynamically Microprogammable Computer with Low-Level Parallelism · IEEE Trans. Computers 1980 |
High-performance computing
scientific computing systems |
0.0 | 1 | 1983 | A User-Microprogrammable, Local Host Computer With Low-Level Parallelism · ISCA 1983 |
Processor architecture and microarchitecture
instruction set architecture |
0.0 | 1 | 1980 | A Dynamically Microprogammable Computer with Low-Level Parallelism · IEEE Trans. Computers 1980 |
Parallel and multicore computing › parallel architecture
MIMD architecture |
0.0 | 1 | 1980 | A Dynamically Microprogammable Computer with Low-Level Parallelism · IEEE Trans. Computers 1980 |
Methods — techniques the papers use, named apart from their topics
topological update scheme · 0.1finite element method · 0.1dynamic scheduling scheme · 0.0simulation · 0.0hierarchical directory · 0.0compiler optimization · 0.0ALU chaining · 0.0tomasulo's algorithm · 0.0speculative execution · 0.0raster-scan display · 0.0hierarchical multiprocessor · 0.0microprogramming · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2012 | Prototype Implementation of a GPU-based Interactive Coupled Fluid-Structure SimulationabstractThis paper reports a prototype implementation of 3D Fluid-Structure Simulation. In this implementation, heartbeat fluid flows inside a rectangular tube where the operator can interactively change the width of the tube. In order to achieve interactive simulation speed, some parts of the simulation is accelerated by GPU. For further improvement of the simulation speed, two techniques have been examined: the first one reduces the communication between CPU and GPU, and the second one offloads a whole fluid flow simulation onto GPU. Ryota Henmi, Yusuke Nishimura, Hiroaki Suzuki, Shinji Fukuma, Shin-ichiro Mori, Akinori Yamaguchi, Shinji Tomita |
SNPD | 7 |
| 2011 | A Fine-Grained Runtime Power/Performance Optimization Method for Processors with Adaptive Pipeline Depth
Jun Yao 0001, Shinobu Miwa, Hajime Shimada, Shinji Tomita |
J. Comput. Sci. Technol. | 4 |
| 2008 | Vertex-preserving Cutting of Elastic ObjectsabstractThis paper proposes vertex-preserving cutting methods on finite element models for interactive soft tissue simulation. Unlike existing methods, we aim to shape variety of incisions using only initial vertices of tetrahedral meshes. Neither tetrahedral decomposition nor vertex creation is used. The number of vertices is preserved. This avoids increase of computation cost as well as allows fast update of physical status of finite element models. To preserve 3D shape and sharp feature of initial meshes through on-the-fly mesh modification, constraints are introduced to the topological update scheme. In our model, the size of stiffness matrix is constant. Our framework efficiently simulates several varieties of smooth incisions with sufficient quality for surgical simulation, and also achieves interactive performance in complex meshes with thousands of elements. Megumi Nakao, Kotaro Minato, Naoto Kume, Shin-ichiro Mori, Shinji Tomita |
VR | 5 |
| 2003 | Bilateral Tradings with and without Strategic Thinking
Shinji Tomita, Akira Namatame |
MABS | 1 |
| 2002 | Development of an Education Tool for Computer SystemabstractAn education tool has been developed to explain internal behavior and structure of computer graphically. It is written in Java Language and designed for students to understand how computer works in the classroom lecture of information science. This paper describes some characteristics of our education tool and improvement of its facilities, especially embedded mail handling module for information exchange between teacher and students. Yoshiro Imai, Shinji Tomita, Haruo Niimi, Hitoshi Inomo, Wataru Shiraki |
ICCE | 2 |
| 2001 | A high-speed dynamic instruction scheduling scheme for superscalar processors
Masahiro Goshima, Kengo Nishino, Toshiaki Kitamura, Yasuhiko Nakashima, Shinji Tomita, Shin-ichiro Mori |
MICRO | 5 |
| 1996 | Amon2: A Parallel Wire Routing Algorithm on a Torus Network Parallel Computer
Hesham Keshk, Shin-ichiro Mori, Hiroshi Nakashima, Shinji Tomita |
International Conference on Supercomputing | 4 |
| 1995 | A proposal of self-cleanup cache
Shin-ichiro Mori, Masahiro Goshima, Hiroshi Nakashima, Shinji Tomita |
PACT | 4 |
| 1995 | Amon: A Parallel Slice Algorithm for Wire RoutingabstractIn thk paper a new parallel global and detailed wire routing algorithm called "Amen" is introduced.Both of them are done in parallel using different processor elements (PEs).We introduce a new way of dividing the multilayer grid into layers, and dividing layers into slices.Each PE, in the detailed routing, has a responsibility for one or more slice.This way of division obtains a high degree of parallelism.A new detailed routing algorithm which routes nets under the condition that these net paths do not prevent other nets from being routed later is also introduced.Amen gives high routing quality by using fewer vias and shorter wire length than the other algorithms.The results show that this algorithm obtains higher connection ratio than maze running algorithm. Hesham Keshk, Shin-ichiro Mori, Hiroshi Nakashima, Shinji Tomita |
International Conference on Supercomputing | 4 |
| 1993 | A distributed shared memory multiprocessor ASURA: memory and cache architectureabstractASURA is a large scale, chwier-based, distributed, shared memory, multiprocessor being developed at Kyoto University and Kubota Corporation.Up to 128 clusters are interconnected to form an ASURA system of up to 102d processors.The basic concept of the ASURA design is to take advantage of the hierarchical structure of the system.Implementing this concept, a large shared cache is piaced between each cluster and the inter-ciuster network.The shared cache and the shared memories distributed among the clusters form part of ASURA's hierarchical memory architecture, providing various unique features to ASURA.In this paper, the hierarchicalmemory architecture of ASURA and its unique cache coherence scheme, including a proposal of a new hierarchical directory scheme, are described wzth some simulation results.Permission tocopy without fes all or pan of Ibis .n81e&al is granted, pmvidd that tbe copies me not made or dishiluted fw dim commenial advantage, the ACMwpyrisbtmalice and tbe title of the F.Iblication and its date appesr, snd notice is given that cqying is by permission of tbe Association forComputing Mschimy.ToCOPY ciherwi%csto republih.requires a fee Snlvcu Specilic pmnissian. Shin-ichiro Mori, Hideki Saito 0001, Masahiro Goshima, Mamoru Yanagihara, Takashi Tanaka, David Fraser, Kazuki Joe, Hiroyuki Nitta, Shinji Tomita |
SC | 9 |
| 1992 | Benchmarking a vector-processor prototype based on multithreaded streaming/FIFO vector (MSFV) architectureabstractThis paper presents the benchmark results on a vector-processor prototype based on the MSFV (multithreaded streaming/FIFO vector) architecture. The MSFV architecture is single-chip oriented, and thus its main object is to save the off-chip memory bandwidth by exploiting the register bandwidth instead. The register bandwidth is exploited by the synergism of FIFO register, chaining, streaming, and multithreading. This paper tries to identify the strength and weakness of those architectural features. The results for basic vector operations and Livermore Fortran Kernels are reported in terms of normalized FLOPC (floating-point operations per clock cycle) and compared to previously-reported results on the Cray X-MP, Y-MP, Fujitsu VP-200, Hitachi S-810/20, NEC SX-2, and SX-3. These comparisons show that, for many basic vector operations, the execution rate of the MSFV prototype results in worst due to its saving the memory bandwidth. However, for Livermore Fortran Kernels, the MSFV prototype results in worst due to its saving the memory bandwidth. However, for Livermore Fortran Kernels, the MSFV prototype outperforms the VP-200 by 2.11 times (geometric mean) and one processor of the X-MP by 1.22 times (geometric mean) in terms of FLOPC. Also, it is 0.67 times (geometric mean) faster than the S-810/20, and 0.76 times (geometric mean) faster than the SX-2. The paper concludes that the MSFV architecture is successful in saving the memory bandwidth. Tetsuo Hironaka, Takashi Hashimoto, Keizo Okazaki, Kazuaki J. Murakami, Shinji Tomita |
ICS | 5 |
| 1989 | The Kyushu University reconfigurable parallel processor: design of memory and intercommunicaiton architecturesabstractThe reconfigurable parallel processor system under development at Kyushu University is an MIMD-type multiprocessor which consists of N processing-elements (currently N is 128) fully connected by S N × N crossbar networks (currently S is 1). Each PE (Processing Element) employs a Fujitsu SPARC MB86900/10 chip-set, a Weitek WTL1164/65 chip-set, an MMU (Memory Management Unit) with 64K bytes of cache, 4M bytes of memory, and an MCU (Message Communication Unit). The modular 128 × 128 crossbar network is implemented by arranging 256 identical 8 × 8 crossbar LSI-modules in a 16 × 16 matrix form. The full 128-PE configuration achieves supercomputer levels of performance by providing 1.28 GIPS and 205 MFLOPS of computing power, 512M bytes of memory, and 2.56G bytes/s of inter-PE communication bandwidth. At the same time, it exploits unique reconfigurability in the memory and intercommunication architectures. By utilizing these two types of reconfigurability, we believe that the system can be effectively tailored to a wide spectrum of applications such as numerical computation, image processing, computer graphics, artificial intelligence, neurocomputing, and so on. Kazuaki J. Murakami, Shin-ichiro Mori, Akira Fukuda, Toshinori Sueyoshi, Shinji Tomita |
ICS | 5 |
| 1989 | SIMP (Single Instruction stream/Multiple Instruction Pipelining): A Novel High-Speed Single-Processor ArchitectureabstractSIMP is a novel multiple instruction-pipeline parallel architecture. It is targeted for enhancing the performance of SISD processors drastically by exploiting both temporal and spatial parallelisms, and for keeping program compatibility as well. Degree of performance enhancement achieved by SIMP depends on; i) how to supply multiple instructions continuously, and ii) how to resolve data and control dependencies effectively. We have devised the outstanding techniques for instruction fetch and dependency resolution. The instruction fetch mechanism employs unique schemes of; i) prefetching multiple instructions with the help of branch prediction, ii) squashing instructions selectively, and iii) providing multiple conditional modes as a result. The dependency resolution mechanism permits out-of-order execution of sequential instruction stream. Our out-of-order execution model is based on Tomasulo's algorithm which has been used in single instruction-pipeline processors. However, it is greatly extended and accommodated to multiple instruction pipelining with; i) detecting and identifying multiple dependencies simultaneously, ii) alleviating the effects of control dependencies with both eager execution and advance execution, and iii) ensuring a precise machine state against branches and interrupts. By taking advantage of these techniques, SIMP is one of the most promising architectures toward the coming generation of high-speed single processors. Kazuaki J. Murakami, Naohiko Irie, Morihiro Kuga, Shinji Tomita |
ISCA | 4 |
| 1986 | A Computer with Low-Level Parallelism QA-2: Its Applications to 3-D Graphics and Prolog/Lisp MachinesabstractWe proposed a computer with low-level parallelism as one of the basic computer architectures and built a large scale experimental system called QA-2. By low-level parallelism, we mean that a long-word instruction controls simultaneously many ALUs, busses, registers and memories in a mode of fine-grained parallelism. The QA-2 employs a 256-bit instruction by which four different ALU operations, four memory accesses to different/continuous locations and one powerful sequence control are all specified and performed in parallel. If many simultaneously executable operations are detected and embedded in one instruction at compile time, this type of computer can provide a high-degree of performance for a wide variety of applications. This paper describes the architectural benefits and limitations of low-level parallelism in performing 3-D color image generation and interpreting Prolog/Lisp programs. The hardware organization with four ALUs, which are actually implemented in the QA-2, is verified to be adequate. In fact, nearly three out of four ALUs can work in parallel. Any architecture with more than four ALUs can not achieve a significant degree of performance enhancement. This paper also shows the degree of performance improvement achieved by the techniques such as ALU chaining and highly-structured sequence control mechanisms. As compared with the IBM 370 architecture, the QA-2 can generate 3-D color images in 1/5 of dynamic instruction steps. The compiler version of Prolog machine on the QA-2 is as fast (45K LIPS) as the ICOT's PSI. From all results, we expect that the QA-2 is a high-performance computer which will be utilized in the future personal computing environment. Shinji Tomita, Kiyoshi Shibayama, Toshiyuki Nakata 0001, Shinji Yuasa, Hiroshi Hagiwara |
ISCA | 1 |
| 1984 | A parallel processor system for three-dimensional color graphicsabstractThis paper describes the hardware architecture and the employed algorithm of a parallel processor system for three-dimensional color graphics. The design goal of the system is to generate realistic images of three-dimensional environments on a raster-scan video display in real-time. In order to achieve this goal, the system is constructed as a two-level hierarchical multi-processor system which is particularly suited to incorporate scan-line algorithm for hidden surface elimination. The system consists of several Scan-Line Processors (SLPs), each of which controls several slave PiXel Processors (PXPs). The SLP prepares the specific data structure relevant to each scan line, while the PXP manipulates every pixel data in its own territory. Internal hardware structures of the SLP and the PXP are quite different, being designed for their dedicated tasks. Haruo Niimi, Yoshiro Imai, Masayoshi Murakami, Shinji Tomita, Hiroshi Hagiwara |
SIGGRAPH | 4 |
| 1983 | A User-Microprogrammable, Local Host Computer With Low-Level ParallelismabstractThis paper describes the architecture of a dynamically microprogrammable computer with low-level parallelism, called QA-2, which is designed as a high-performance, local host computer for laboratory use. The architectural principle of the QA-2 is the marriage of high-speed, parallel processing capability offered by four powerful Arithmetic and Logic Units (ALUs) with architectural flexibility provided by large scale, dynamic user-microprogramming. By changing its writable control storage dynamically, the QA-2 can be tailored to a wide spectrum of research-oriented applications covering high-level language processing and real-time processing. Shinji Tomita, Kiyoshi Shibayama, Toshiaki Kitamura, Toshiyuki Nakata 0001, Hiroshi Hagiwara |
ISCA | 1 |
| 1980 | A Dynamically Microprogammable Computer with Low-Level ParallelismabstractA new microprogrammable computer with low-level parallelism was built and has been utilized as a research vehicle for solving different classes of research-oriented applications such as real-time processings on static/dynamic images, pictures and signals, and emulations of both existing and virtual machines including high (intermediate) level language machines. The design goal of the machine was to achieve a high degree of processing enhancement in research- oriented applications by means of a low-level parallel processing organization combined with dynamically microprogrammable control. The machine has the capability to process multiple data streams, performing parallel operations with four 16-bit ALU's. These ALU's are independently controlled by the different fields of a 160-bit horizontal-type microinstruction, and have simultaneous access to 15 working registers. This microprogrammed MIMD organization is expected to provide a greater degree of flexibility for low-level parallel processing. In addition, not only does the machine contain powerful ALU's and a large number of registers, but also it employs flexible control structures and a hierarchical organization of control storage. All of these combine to yield extensive microprogramming capability which the user can effectively tailor to a wide spectrum of applications. Hiroshi Hagiwara, Shinji Tomita, Shigeru Oyanagi, Kiyoshi Shibayama |
IEEE Trans. Computers | 2 |