Kiyokuni Kawachiya

dblp:43/3549 · DBLP profile ↗
← Back
21ranked-venue papers
5as first author
0since 2021 · last 2020
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 10 · 2 first-authorSystems, architecture and hardware · 6 · 1 first-authorComputer networks · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2Artificial intelligence and machine learning · 1Databases, data management, data science and information retrieval · 1Graphics, computer vision, multimedia, augmented reality and games · 1Human-computer interaction and ubiquitous computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
4 papers
Distributed systems · 38% Parallel and multicore computing · 22% GPUs and heterogeneous computing · 15%
Artificial intelligence
1 paper
Efficient and distributed learning · 100%
Software engineering, system software, and programming languages
6 papers
Runtime systems and virtual machines · 51% Concurrent programming · 14% Operating systems · 12%

Topics — the 21 heaviest of 25, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Distributed systems
fault tolerance
0.622019
Failure Recovery in Resilient X10 · ACM Trans. Program. Lang. Syst. 2019
Resilient X10: efficient failure-aware programming · PPoPP 2014
Parallel and multicore computing
parallel programming models
0.622019
Failure Recovery in Resilient X10 · ACM Trans. Program. Lang. Syst. 2019
Resilient X10: efficient failure-aware programming · PPoPP 2014
Machine learning › Efficient and distributed learning
memory-efficient training
0.412019
Profiling based out-of-core hybrid method for large neural networks: poster · PPoPP 2019
Machine learning › Efficient and distributed learning › memory-efficient training
out-of-core training
0.412019
Profiling based out-of-core hybrid method for large neural networks: poster · PPoPP 2019
Distributed systems › fault tolerance
failure recovery
0.412019
Failure Recovery in Resilient X10 · ACM Trans. Program. Lang. Syst. 2019
GPUs and heterogeneous computing
GPU computing
0.412019
Profiling based out-of-core hybrid method for large neural networks: poster · PPoPP 2019
Storage systems
out-of-core computation
0.412019
Profiling based out-of-core hybrid method for large neural networks: poster · PPoPP 2019
Runtime systems and virtual machines › virtual machine implementation
java virtual machine
0.132010
A study of Java's non-Java memory · OOPSLA 2010
Analysis and reduction of memory inefficiencies in Java strings · OOPSLA 2008
Lock reservation: Java locks can mostly do without atomic operations · OOPSLA 2002
Runtime systems and virtual machines › dynamic compilation
just-in-time compilation
0.122006
Replay compilation: improving debuggability of a just-in-time compiler · OOPSLA 2006
Effectiveness of cross-platform optimizations for a java just-in-time compiler · OOPSLA 2003
Runtime systems and virtual machines
garbage collection
0.112008
Analysis and reduction of memory inefficiencies in Java strings · OOPSLA 2008
Operating systems › resource management
memory management
0.112008
Analysis and reduction of memory inefficiencies in Java strings · OOPSLA 2008
Debugging and program repair › software debugging
compiler debugging
0.112006
Replay compilation: improving debuggability of a just-in-time compiler · OOPSLA 2006
Concurrent programming
synchronization
0.122002
Lock reservation: Java locks can mostly do without atomic operations · OOPSLA 2002
A Study of Locking Objects with Bimodal Fields · OOPSLA 1999
Program analysis
data flow analysis
0.012003
Effectiveness of cross-platform optimizations for a java just-in-time compiler · OOPSLA 2003
Compilers and program optimization › interprocedural optimization
inlining
0.012003
Effectiveness of cross-platform optimizations for a java just-in-time compiler · OOPSLA 2003
Concurrent programming › synchronization
locking
0.012002
Lock reservation: Java locks can mostly do without atomic operations · OOPSLA 2002
Performance modeling and evaluation
workload characterization
0.012010
A study of Java's non-Java memory · OOPSLA 2010
Content delivery and video streaming
video-on-demand
0.011996
A Portable Communication System for Video-On-Demand Applications Using the Existing Infrastructure · INFOCOM 1996
Runtime systems and virtual machines › managed runtime
java runtime
0.011999
A Study of Locking Objects with Bimodal Fields · OOPSLA 1999
Ubiquitous computing and smart environments › mobile computing
mobile web browsing
0.011998
NaviPoint: An Input Device for Mobile Information Browsing · CHI 1998
Transport protocols and congestion control
flow control
0.011996
A Portable Communication System for Video-On-Demand Applications Using the Existing Infrastructure · INFOCOM 1996

Methods — techniques the papers use, named apart from their topics

profiling · 0.8happens-before invariance · 0.4formal semantics · 0.4visualization · 0.2memory statistics gathering · 0.2heap analysis · 0.1garbage collection · 0.1state saving · 0.1replay compilation · 0.1partial redundancy elimination · 0.0exception check elimination · 0.0thread locality exploitation · 0.0user evaluation · 0.0input device design · 0.0user-level library implementation · 0.0
YearPublicationVenuePosition
2020 Acceleration of Large Deep Learning Training with Hybrid GPU Memory Management of Swapping and Re-computing
abstract
Deep learning has achieved overwhelmingly better accuracy than existing methods in various fields. To further improve deep learning, a deeper and larger neural network (NN) model is indispensable, but graphics processing unit (GPU) memory is not large enough to train such NN models. One promising method to reduce GPU memory consumption is data swapping, which eases the burden on GPU memory by swapping out intermediate data from GPU memory to central processing unit (CPU) memory while the data are not necessary. However, this method introduces communication overhead in transferring data between GPU and CPU memory. Another method is recomputation, which discards intermediate data once and then computes them again when necessary. Unlike the data-swapping method, the re-computation method introduces additional computation but does not require CPU-GPU communication. Therefore, it may reduce CPU-GPU communication by introducing it effectively in the data-swapping method. In this paper, we developed a faster training method for large NN models by combining the re-computation and the data-swapping methods. This method edits the graph of TensorFlow automatically. We developed heuristics about which part of the graph should be recomputed. The heuristics divides the graph into sub-graphs and applies the re-computation in the decreasing order of the amount of swapping data in the sub-graphs. This heuristics is intended to reduce maximal communications by re-computing minimal sub-graphs. Our hybrid method improves performance by up to 15.5% in image size 8000×8000 of ResNet50, 15.3% in image size 7500×7500 of ResNet152, and 12.4% in image size 1000×1000 of DeepLabV3+ compared with the existing data-swapping method.
Haruki Imai, Tung D. Le, Yasushi Negishi, Kiyokuni Kawachiya
IEEE BigData4
2019 Automatic GPU memory management for large neural models in TensorFlow
abstract
Deep learning models are becoming larger and will not fit in the limited memory of accelerators such as GPUs for training. Though many methods have been proposed to solve this problem, they are rather ad-hoc in nature and difficult to extend and integrate with other techniques. In this paper, we tackle the problem in a formal way to provide a strong foundation for supporting large models. We propose a method of formally rewriting the computational graph of a model where swap-out and swap-in operations are inserted to temporarily store intermediate results on CPU memory. By introducing a categorized topological ordering for simulating graph execution, the memory consumption of a model can be easily analyzed by using operation distances in the ordering. As a result, the problem of fitting a large model into a memory-limited accelerator is reduced to the problem of reducing operation distances in a categorized topological ordering. We then show how to formally derive swap-out and swap-in operations from an existing graph and present rules to optimize the graph. Finally, we propose a simulation-based auto-tuning to automatically find suitable graph-rewriting parameters for the best performance. We developed a module in TensorFlow, called LMS, by which we successfully trained ResNet-50 with a 4.9x larger mini-batch size and 3D U-Net with a 5.6x larger image resolution.
Tung D. Le, Haruki Imai, Yasushi Negishi, Kiyokuni Kawachiya
ISMM4
2019 High Resolution Medical Image Segmentation Using Data-Swapping Method
Haruki Imai, Samuel Matzek, Tung D. Le, Yasushi Negishi, Kiyokuni Kawachiya
MICCAI (3)5
2019 Profiling based out-of-core hybrid method for large neural networks: poster
abstract
Neural networks (NNs) have archived high accuracy in many fields of machine learning such as image recognition. Since computations of NNs are heavy tasks, GPUs have been widely used to accelerate them. However, the problem sizes of NNs that can be computed are limited by GPU memory capacity.
Haruki Imai, Tung D. Le, Yasushi Negishi, Kiyokuni Kawachiya, Ryo Matsumiya, Toshio Endo
PPoPP5
2019 Failure Recovery in Resilient X10
abstract
Cloud computing has made the resources needed to execute large-scale in-memory distributed computations widely available. Specialized programming models, e.g., MapReduce, have emerged to offer transparent fault tolerance and fault recovery for specific computational patterns, but they sacrifice generality. In contrast, the Resilient X10 programming language adds failure containment and failure awareness to a general purpose, distributed programming language. A Resilient X10 application spans over a number of places. Its formal semantics precisely specify how it continues executing after a place failure. Thanks to failure awareness, the X10 programmer can in principle build redundancy into an application to recover from failures. In practice, however, correctness is elusive, as redundancy and recovery are often complex programming tasks. This article further develops Resilient X10 to shift the focus from failure awareness to failure recovery, from both a theoretical and a practical standpoint. We rigorously define the distinction between recoverable and catastrophic failures. We revisit the happens-before invariance principle and its implementation. We shift most of the burden of redundancy and recovery from the programmer to the runtime system and standard library. We make it easy to protect critical data from failure using resilient stores and harness elasticity—dynamic place creation—to persist not just the data but also its spatial distribution. We demonstrate the flexibility and practical usefulness of Resilient X10 by building several representative high-performance in-memory parallel application kernels and frameworks. These codes are 10× to 25× larger than previous Resilient X10 benchmarks. For each application kernel, the average runtime overhead of resiliency is less than 7%. By comparing application kernels written in the Resilient X10 and Spark programming models, we demonstrate that Resilient X10’s more general programming model can enable significantly better application performance for resilient in-memory distributed computations.
David Grove, Sara S. Hamouda, Benjamin Herta, Arun Iyengar, Kiyokuni Kawachiya, Josh Milthorpe, Vijay A. Saraswat, Avraham Shinnar, Mikio Takeuchi, Olivier Tardieu
ACM Trans. Program. Lang. Syst.5
2018 Involving CPUs into Multi-GPU Deep Learning
abstract
The most important part of deep learning, training the neural network, often requires the processing of a large amount of data and can takes days to complete. Data parallelism is widely used for training deep neural networks on multiple GPUs in a single machine thanks to its simplicity. However, its scalability is bound by the number of data transfers, mainly for exchanging and accumulating gradients among the GPUs. In this paper, we present a novel approach to data parallel training called CPU-GPU data parallel (CGDP) training that utilizes free CPU time on the host to speed up the training in the GPUs. We also present a cost model for analyzing and comparing the performances of both the typical data parallel training and the CPU-GPU data parallel training. Using the cost model, we formally show why our approach is better than the typical one and clarify the remaining issues. Finally, we explain how we optimized CPU-GPU data parallel training by introducing chunks of layers and present a runtime algorithm that automatically finds a good configuration for the training. The algorithm is effective for very deep neural networks, which are the current trend in deep learning. Experimental results showed that we achieved speedups of $1.21$, $1.04$, $1.21$ and $1.07$ for four state-of-the-art neural networks: AlexNet, GoogLeNet-v1, VGGNet-16, and ResNet-152, respectively. Weak scaling efficiency greater than $90$ was achieved for all networks across four GPUs.
Tung D. Le, Taro Sekiyama, Yasushi Negishi, Haruki Imai, Kiyokuni Kawachiya
ICPE5
2014 Resilient X10: efficient failure-aware programming
abstract
Scale-out programs run on multiple processes in a cluster. In scale-out systems, processes can fail. Computations using traditional libraries such as MPI fail when any component process fails. The advent of Map Reduce, Resilient Data Sets and MillWheel has shown dramatic improvements in productivity are possible when a high-level programming framework handles scale-out and resilience automatically.
David Cunningham, David Grove, Benjamin Herta, Arun Iyengar, Kiyokuni Kawachiya, Hiroki Murata, Vijay A. Saraswat, Mikio Takeuchi, Olivier Tardieu
PPoPP5
2014 String deduplication for Java-based middleware in virtualized environments
abstract
To increase the memory efficiency in physical servers is a significant concern for increasing the number of virtual machines (VM) in them. When similar web application service runs in each guest VM, many string data with the same values are created in every guest VMs. These duplications of string data are redundant from the viewpoint of memory efficiency in the host OS. This paper proposes two approaches to reduce the duplication in Java string in a single Java VM (JVM) and across JVMs. The first approach is to share string objects cross JVMs by using a read-only memory-mapped file. The other approach is to selectively unify string objects created at runtime in the web applications. This paper evaluates our approach by using the Apache DayTrader and the DaCapo benchmark suite. Our prototype implementation chieved 7% to 12% reduction in the total size of the objects allocated over the lifetime of the programs. In addition, we observed the performance of DayTrader was maintained even under a situation of high density guest VMs in a KVM host machine.
Michihiro Horie, Kazunori Ogata, Kiyokuni Kawachiya, Tamiya Onodera
VEE3
2010 A study of Java's non-Java memory
abstract
A Java application sometimes raises an out-of-memory ex-ception. This is usually because it has exhausted the Java heap. However, a Java application can raise an out-of-memory exception when it exhausts the memory used by Java that is not in the Java heap. We call this area non-Java memory. For example, an out-of-memory exception in the non-Java memory can happen when the JVM attempts to load too many classes. Although it is relatively rare to ex-haust the non-Java memory compared to exhausting the Java heap, a Java application can consume a considerable amount of non-Java memory.This paper presents a quantitative analysis of non-Java memory. To the best of our knowledge, this is the first in-depth analysis of the non-Java memory. To do this we cre-ated a tool called Memory Analyzer for Redundant, Unused, and String Areas (MARUSA), which gathers memory statis-tics from both the OS and the Java virtual machine, break-ing down and visualizing the non-Java memory usage.We studied the use of non-Java memory for a wide range of Java applications, including the DaCapo benchmarks and Apache DayTrader. Our study is based on the IBM J9 Java Virtual Machine for Linux. Although some of our results may be specific to this combination, we believe that most of our observations are applicable to other platforms as well.
Kazunori Ogata, Dai Mikurube, Kiyokuni Kawachiya, Scott Trent, Tamiya Onodera
OOPSLA3
2010 Efficient runtime tracking of allocation sites in Java
abstract
Tracking the allocation site of every object at runtime is useful for reliable, optimized Java. To be used in production environments, the tracking must be accurate with minimal speed loss. Previous approaches suffer from performance degradation due to the additional field added to each object or track the allocation sites only probabilistically. We propose two novel approaches to track the allocation sites of every object in Java with only a 1.0% slow-down on average. Our first approach, the Allocation-Site-as-a-Hash-code (ASH) Tracker, encodes the allocation site ID of an object into the hash code field of its header by regarding the ID as part of the hash code. ASH Tracker avoids an excessive increase in hash code collisions by dynamically shrinking the bit-length of the ID as more and more objects are allocated at that site. For those Java VMs without the hash code field, our second approach, the Allocation-Site-via-a-Class-pointer (ASC) Tracker, makes the class pointer field in an object header refer to the allocation site structure of the object, which in turn points to the actual class structure. ASC Tracker mitigates the indirection overhead by constant-class-field duplication and allocation-site equality checks. While a previous approach of adding a 4-byte field caused up to 14.4% and an average 5% slowdown, both ASH and ASC Trackers incur at most a 2.0% and an average 1.0% loss. We demonstrate the usefulness of our low-overhead trackers by an allocation-site-aware memory leak detector and allocation-site-based pretenuring in generational GC. Our pretenuring achieved on average 1.8% and up to 11.8% speedups in SPECjvm2008.
Rei Odaira, Kazunori Ogata, Kiyokuni Kawachiya, Tamiya Onodera, Toshio Nakatani
VEE3
2008 Analysis and reduction of memory inefficiencies in Java strings
abstract
This paper describes a novel approach to reduce the memory consumption of Java programs, by focusing on their "string memory inefficiencies". In recent Java applications, string data occupies a large amount of the heap area. For example, about 40% of the live heap area is used for string data when a production J2EE application server is running. By investigating the string data in the live heap, we identified two types of memory inefficiencies -- "duplication" and "unused literals". In the heap, there are many string objects that have the same values. There also exist many string literals whose values are not actually used by the application. Since these inefficiencies exist as live objects, they cannot be eliminated by existing garbage collection techniques, which only remove dead objects. Quantitative analysis of Java heaps in real applications revealed that more than 50% of the string data in the live heap is wasted by these inefficiencies. To reduce the string memory inefficiencies, this paper proposes two techniques at the Java virtual machine level, "StringGC" for eliminating duplicated strings at the time of garbage collection, and "Lazy Body Creation" for delaying part of the literal instantiation until the literal's value is actually used. We also present an interesting technique at the Java program level, which we call "BundleConverter", for preventing unused message literals from being instantiated. Prototype implementations on a production Java virtual machine have achieved about 18% reduction of the live heap in the production application server. The proposed techniques could also reduce the live heap of standard Java benchmarks by 11.6% on average, without noticeable performance degradation.
Kiyokuni Kawachiya, Kazunori Ogata, Tamiya Onodera
OOPSLA1
2007 Libra: a library operating system for a jvm in a virtualized execution environment
abstract
If the operating system could be specialized for every application, many applications would run faster. For example, Java virtual machines (JVMs) provide their own threading model and memory protection, so general-purpose operating system implementations of these abstractions are redundant. However, traditional means of transforming existing systems into specialized systems are difficult to adopt because they require replacing the entire operating system. This paper describes Libra, an execution environment specialized for IBM's J9 JVM. Libra does not replace the entire operating system. Instead, Libra and J9 form a single statically-linked image that runs in a hypervisor partition. Libra provides the services necessary to achieve good performance for the Java workloads of interest but relies on an instance of Linux in another hypervisor partition to provide a networking stack, a filesystem, and other services. The expense of remote calls is offset by the fact that Libra's services can be customized for a particular workload; for example, on the Nutch search engine, we show that two simple customizations improve application throughput by a factor of 2.7.
Glenn Ammons, Jonathan Appavoo, Maria A. Butrico, Dilma Da Silva, David Grove, Kiyokuni Kawachiya, Orran Krieger, Bryan S. Rosenburg, Eric Van Hensbergen, Robert W. Wisniewski
VEE6
2007 Cloneable JVM: a new approach to start isolated java applications faster
abstract
Java has been successful particularly for writing applications in the server environment. However, isolation of multiple applications hasnot been efficiently achieved in Java. Many customers require that their applications are guarded by independent OS processes, but starting a Java application with a new process results in a long sequence of initializations being repeated each time. To date, there has been no way to quickly start a new Java application as an isolated OS process. In this paper, we propose a new isolation approach called Cloneable JVM to eliminate this startup overhead in Java. The key idea is to createa new Java application by copying, or cloning, the already-initialized image of the primary JVM process. Since the clone is already initialized, it can begin actual operations immediately as a new isolated process. This cloning abstraction can support new scenarios for Java, such as user isolation and transaction isolation. We implemented a prototype of the Cloneable JVM by modifying a production JVM on Linux, which provides a new API for cloning constructed on the Isolate API defined in JSR 121. Using this cloning API, several Java applications, including a large production J2EE application server, we remodified to demonstrate the isolation scenarios. Evaluations using these prototypes showed that new ready-to-serve Java applications can start up as a new process in less than 5 seconds, which is 4 to 170 times faster than starting these applications from scratch.
Kiyokuni Kawachiya, Kazunori Ogata, Daniel Silva 0001, Tamiya Onodera, Hideaki Komatsu, Toshio Nakatani
VEE1
2006 Replay compilation: improving debuggability of a just-in-time compiler
abstract
The performance of Java has been tremendously improved by the advance of Just-in-Time (JIT) compilation technologies. However, debugging such a dynamic compiler is much harder than a static compiler. Recompiling the problematic method to produce a diagnostic output does not necessarily work as expected, because the compilation of a method depends on runtime information at the time of compilation.In this paper, we propose a new approach, called replay JIT compilation, which can reproduce the same compilation remotely by using two compilers, the state-saving compiler and the replaying compiler. The state-saving compiler is used in a normal run, and, while compiling a method, records into a log all of the input for the compiler. The replaying compiler is then used in a debugging run with the system dump, to recompile a method with the options for diagnostic output. We reduced the overhead to save the input by using the system dump and by categorizing the input based on how its value changes. In our experiment, the increase of the compilation time for saving the input was only 1%, and the size of the additional memory needed for saving the input was only 10% of the compiler-generated code.
Kazunori Ogata, Tamiya Onodera, Kiyokuni Kawachiya, Hideaki Komatsu, Toshio Nakatani
OOPSLA3
2004 Lock Reservation for Java Reconsidered
Tamiya Onodera, Kiyokuni Kawachiya, Akira Koseki
ECOOP2
2003 Effectiveness of cross-platform optimizations for a java just-in-time compiler
abstract
This paper describes the system overview of our Java Just-In-Time (JIT) compiler, which is the basis for the latest production version of IBM Java JIT compiler that supports a diversity of processor architectures including both 32-bit and 64-bit modes, CISC, RISC, and VLIW architectures. In particular, we focus on the design and evaluation of the cross-platform optimizations that are common across different architectures. We studied the effectiveness of each optimization by selectively disabling it in our JIT compiler on three different platforms: IA-32, IA-64, and PowerPC. Our detailed measurements allowed us to rank the optimizations in terms of the greatest performance improvements with the smallest compilation times. The identified set includes method inlining only for tiny methods, exception check eliminations using forward dataflow analysis and partial redundancy elimination, scalar replacement for instance and class fields using dataflow analysis, optimizations for type inclusion checks, and the elimination of merge points in the control flow graphs. These optimizations can achieve 90% of the peak performance for two industry-standard benchmark programs on these platforms with only 34% of the compilation time compared to the case for using all of the optimizations.
Kazuaki Ishizaki, Mikio Takeuchi, Kiyokuni Kawachiya, Toshio Suganuma, Osamu Gohda, Tatsushi Inagaki, Akira Koseki, Kazunori Ogata, Motohiro Kawahito, Toshiaki Yasue, Takeshi Ogasawara, Tamiya Onodera, Hideaki Komatsu, Toshio Nakatani
OOPSLA3
2002 Lock reservation: Java locks can mostly do without atomic operations
abstract
Because of the built-in support for multi-threaded programming, Java programs perform many lock operations. Although the overhead has been significantly reduced in the recent virtual machines, One or more atomic operations are required for acquiring and releasing an object's lock even in the fastest cases.This paper presents a novel algorithm called lock reservation. It exploits thread locality of Java locks, which claims that the locking sequence of a Java lock contains a very long repetition of a specific thread. The algorithm allows locks to be reserved for threads. When a thread attempts to acquire a lock, it can do without any atomic operation if the lock is reserved for the thread. Otherwise, it cancels the reservation and falls back to a conventional locking algorithm.We have evaluated an implementation of lock reservation in IBM's production virtual machine and compiler. The results show that it achieved performance improvements up to 53% in real Java programs.
Kiyokuni Kawachiya, Akira Koseki, Tamiya Onodera
OOPSLA1
1999 A Study of Locking Objects with Bimodal Fields
abstract
Object locking can be efficiently implemented by bimodal use of a field reserved in an object. The field is used as a lightweight lock in one mode, while it holds a reference to a heavyweight lock in the other mode. A bimodal locking algorithm recently proposed for Java achieves the highest performance in the absence of contention, and is still fast enough when contention occurs.
Tamiya Onodera, Kiyokuni Kawachiya
OOPSLA2
1998 NaviPoint: An Input Device for Mobile Information Browsing
abstract
Article Free Access Share on NaviPoint: an input device for mobile information browsing Authors: Kiyokuni Kawachiya IBM Research, Tokyo Research Laboratory, 1623-14, Shimotsuruma, Yamato, Kanagawa 242-8502, Japan IBM Research, Tokyo Research Laboratory, 1623-14, Shimotsuruma, Yamato, Kanagawa 242-8502, JapanView Profile , Hiroshi Ishikawa IBM Research, Tokyo Research Laboratory, 1623-14, Shimotsuruma, Yamato, Kanagawa 242-8502, Japan IBM Research, Tokyo Research Laboratory, 1623-14, Shimotsuruma, Yamato, Kanagawa 242-8502, JapanView Profile Authors Info & Claims CHI '98: Proceedings of the SIGCHI Conference on Human Factors in Computing SystemsJanuary 1998 Pages 1–8https://doi.org/10.1145/274644.274645Published:01 January 1998Publication History 14citation985DownloadsMetricsTotal Citations14Total Downloads985Last 12 Months13Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF
Kiyokuni Kawachiya, Hiroshi Ishikawa 0008
CHI1
1996 A Portable Communication System for Video-On-Demand Applications Using the Existing Infrastructure
abstract
Video on demand (VOD) systems are becoming popular. They have special requirements for receiving VOD data steadily, but common general communication protocols cannot meet their requirements. Several new protocols have been designed to fulfil this service requirement, but it takes a long time for a protocol to be supported from end to end. We propose a VOD communication system that is implemented as an user-level library on top of a common existing protocol. This protocol is easy to port, and the level of connectivity is almost the same as in the underlying existing protocol. We describe the important points related to its implementation, such as the low overhead of implementation and the flow control mechanism. We implemented a prototype system, measured its performance, and concluded that implementing a communication system as a user-level library enables us to implement a VOD portable communication system with a practical level of performance.
Yasushi Negishi, Kiyokuni Kawachiya, Kazuya Tago
INFOCOM2
1995 Evaluation of QoS-Control Servers on Real-Time Mach
Kiyokuni Kawachiya, Masanobu Ogata, Nobuhiko Nishio, Hideyuki Tokuda
NOSSDAV1