Mary Inaba

dblp:83/1237 · DBLP profile ↗
← Back
30ranked-venue papers
4as first author
3since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 1 first-author · 3 since 2021Systems, architecture and hardware · 11Theory of computation · 9 · 3 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 since 2021Computer networks · 1Software engineering, systems software and programming languages · 1Databases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2024 Parallel Clause Sharing Strategy Based on Graph Structure of SAT Problem
Yoichiro Iida, Tomohiro Sonobe, Mary Inaba
SAT3
2023 Understand Restart of SAT Solver Using Search Similarity Index (Student Abstract)
abstract
SAT solvers are widely used to solve many industrial problems because of their high performance, which is achieved by various heuristic methods. Understanding why these methods are effective is essential to improving them. One approach to this is analyzing them using qualitative measurements. In our previous study, we proposed search similarity index (SSI), a metric to quantify the similarity between searches. SSI significantly improved the performance of the parallel SAT solver. Here, we apply SSI to analyze the effect of restart, a key SAT solver technique. Experiments using SSI reveal the correlation between the difficulty of instances and the search change effect by restart, and the reason behind the effectiveness of the state-of-the-art restart method is also explained.
Yoichiro Iida, Tomohiro Sonobe, Mary Inaba
AAAI3
2022 Diversification of Parallel Search of Portfolio SAT Solver by Search Similarity Index
Yoichiro Iida, Tomohiro Sonobe, Mary Inaba
PRICAI (1)3
2018 Skyline Computation for Low-Latency Image-Activated Cell Identification
abstract
High-throughput label-free single cell screening technology has been studied for noninvasive analysis of various kinds of cells. We tackle the cell identification task in the cell sorting system as a continuous skyline computation. Skyline Computation is a method for extracting interesting entries from a large population with multiple attributes. Jointed rooted-tree (JR-tree) is continuous skyline computation algorithm that manages entries using a rooted-tree structure. JR-tree delays extend the tree to deeper levels to accelerate tree construction and traversal. In this study, we proposed the JR-tree-based parallel skyline computation accelerator. We implemented it on a field-programmable gate array (FPGA). We evaluated our proposed software and hardware algorithms against an existing software algorithm using synthetic and real-world datasets.
Kenichi Koizumi, Kei Hiraki, Mary Inaba
AAAI3
2018 Constructing Hierarchical Bayesian Networks With Pooling
abstract
Inspired by the Bayesian brain hypothesis and deep learning, we develop a Bayesian autoencoder, a method of constructing recognition systems using a Bayesian network. We construct hierarchical Bayesian networks based on feature extraction and implement pooling to achieve invariance within a Bayesian network framework. The constructed networks propagate information bidirectionally between layers. We expect they will be able to achieve brain-like recognition using local features and global information such as their environments.
Kaneharu Nishino, Mary Inaba
AAAI2
2018 Continuous Skyline Computation Accelerator with Parallelizing Dominance Relation Calculations: (Abstract Only)
abstract
Skyline Computation is a method for extracting interesting entries from a large population with multiple attributes. These entries, called skyline or Pareto optimal entries, are known to have extreme characteristics that cannot be found by using outlier detection methods. Skyline computation is an important task for characterizing large amounts of data and selecting interesting entries with extreme features. When the population changes dynamically, the task of calculating a sequence of skyline sets is called a continuous skyline computation. This task is known to be difficult for the following reasons: (1) information must be kept for non-skyline entries, since they may join the skyline in the future; (2) the appearance or disappearance of even a single entry can change the skyline drastically; and (3) it is difficult to adopt a geometric acceleration algorithm for skyline computation tasks with high-dimensional datasets. A new algorithm, called jointed rooted-tree (JR-tree), has been developed that manages entries using a rooted-tree structure. JR-tree delays extend the tree to deeper levels to accelerate tree construction and traversal. In this study, we propose the JR-tree based continuous skyline computation acceleration algorithm. Our hardware algorithm parallelizes the calculations of dominance relation between a target entry and the skyline entries. We implemented our hardware algorithm on an FPGA and showed that high-speed tree construction and traversal can be realized. Comparing our FPGA-based implementation with an Intel CPU running state-of-the-art software algorithms, it was found to reduce the query processing time for synthetic and real-world datasets. Our hardware implementation is 1.7x to 35x faster than the software implementations.
Kenichi Koizumi, Kei Hiraki, Mary Inaba
FPGA3
2017 BJR-Tree: Fast Skyline Computation Algorithm for Serendipitous Searching Problems
abstract
High-throughput label-free single cell screening technology has been studied for noninvasive analysis of various kinds of cells. Selecting the prominent cells with extreme features from a large number of cells is an important and interesting problem. We call this problem the serendipitous searching problem (SSP), where it is important to find entries located near the rind of the population in multi-dimensional feature space. We tackle the SSP as a continuous skyline computation. The skyline computation is originally used to extract interesting entries from a database with multi-attributions. The skyline points are continuously updated at the time existing entries disappear and new entries arrive. In this paper, we propose a balanced jointed rooted tree (BJR-tree) algorithm and a nondominated relation cache (ND-cache) for continuous skyline computation. The BJR-tree expresses the dominance relation as an arc and stores the "dominated" relations. The ND-cache complements the BJR-tree by reducing the recalculation of the dominance relations. The execution times of the BJR-tree and existing continuous skyline computation algorithms are compared on randomly constructed synthetic datasets with multiple temporal and spatial features. The BJR-tree is then evaluated on actual measured information of blood cells. On the two and eight-dimensional synthetic datasets, the BJR-tree computed the continuous skylines approximately 3x and 70x faster than the LookOut, respectively. On actual datasets, the BJR-tree is approximately 2.4x faster than the LookOut.
Kenichi Koizumi, Peter Eades, Kei Hiraki, Mary Inaba
DSAA4
2017 Boost SAT Solver with Hybrid Branching Heuristic
abstract
Most state-of-the-art satisfiability (SAT) solvers are capable of solving large application instances with efficient branching heuristics. The VSIDS heuristic is widely used because of its robustness. This paper focuses on the inherent ties in VSIDS and proposes a new branching heuristic called TBVSIDS, which attempts to break the ties with the consideration of the interplay between the branching heuristic and learned clauses. However, a branching heuristic cannot cover all problems, and its performance improves when combined with an appropriate configuration. Therefore, we also propose a hybrid model of branching heuristics based on random forest. The efficiencies of TBVSIDS and hybrid branching heuristics are evaluated on benchmarks in SAT Competitions. By constructing a model that reduces the overfitting problem, we hope to realize a hybrid branching heuristic that is widely applicable to other solvers.
Seongsoo Moon, Mary Inaba
SOCS2
2016 Bayesian AutoEncoder: Generation of Bayesian Networks with Hidden Nodes for Features
abstract
We propose Bayesian AutoEncoder (BAE) in order to construct a recognition system which uses feedback information. BAE constructs a generative model of input data as a Bayes Net. The network trained by BAE obtains its hidden variables as the features of given data. It can execute inference for each variable through belief propagation, using both feedforward and feedback information. We confirmed that BAE can construct small networks with one hidden layer and extract features as hidden variables from 3x3 and 5x5 pixel input data.
Kaneharu Nishino, Mary Inaba
AAAI2
2015 Feature Extraction Based on Generating Bayesian Network
Kaneharu Nishino, Mary Inaba
ICONIP (4)2
2015 Efficient implementation of continuous skyline computation on a multi-core processor
abstract
The skyline operator has been proposed as a method for extracting highly-utility samples from a large database. A set of the extracted samples is called `skyline'. The theme of the MEMOCODE 2015 Design Contest is to accelerate continuous skyline computation, skyline computing for a streaming dataset, on any platform. In this paper, we present our method that achieved the best performance in the contest. We describe our data structure, algorithms, and optimization methods for the contest reference code in the multi-core processor. We have accelerated our solution in the two aspects of efficient algorithms and code optimizations. The task of the contest is to compute the skyline at each time-step for 800,000 entries with a seven-dimensional vector value and the activation time and the deactivation time. We use one commodity computer and the average runtime of our solution is 407 milliseconds.
Kenichi Koizumi, Mary Inaba, Kei Hiraki
MEMOCODE2
2014 Community Branching for Parallel Portfolio SAT Solvers
Tomohiro Sonobe, Shuya Kondoh, Mary Inaba
SAT3
2012 Unified memory optimizing architecture: memory subsystem control with a unified predictor
abstract
Data prefetching, advanced cache replacement policy, and memory access scheduling are incorporated in modern processors. Typically, each technique holds recently accessed locations independently and controls the memory subsystem based on the prediction of future memory access. Unfortunately, these specific optimizations often increase the implementation cost, decrease the system performance, and reduce scalability of the processor chip.
Yasuo Ishii, Mary Inaba, Kei Hiraki
ICS2
2009 Access map pattern matching for data cache prefetch
abstract
A novel data prefetching method -- access map pattern matching (AMPM) -- that uses "memory access map" is proposed. The AMPM prefetching concentrate hardware resources on collecting the access footprint of the frequently accessed area which we called "hot zones". 2-bit state is associated with each cache lines of hot zone. A set of these states is called "memory access map". Prefetch requests are generated from the pattern matching of the memory access map. The pattern matching detects multiple memory access patterns in parallel and generates more prefetch requests than conventional prefetchers. The evaluation result shows that the AMPM prefetcher improves performance by 42.0% in FP benchmarks.
Yasuo Ishii, Mary Inaba, Kei Hiraki
ICS2
2008 MCAMP: communication optimization on massively parallel machines with hierarchical scratch-pad memory
abstract
Massively parallel machines that integrate a large number of simple processors and small scratch-pad memories (SPMs) into a single chip can achieve a high peak performance per watt of power. In these machines, communication optimizations are important because the communication bandwidth tends to be a bottleneck. Previously proposed communication optimizations using copy candidates, which have been shown to be effective, detect frequently reused array regions by compile-time analysis and copy the regions to scratch-pad memories nearer to the processors. However, they have been proposed for uniprocessor systems or small parallel machines with one or more layers of scratch-pad memories, and the analysis time increases when they are applied to massively parallel machines. In this paper, we propose Multilayer Copy-candidate Analysis for Massively Parallel machines (MCAMP), a communication optimization method for massively parallel machines. MCAMP re-formalizes the framework used in earlier works and improves the scalability of the analysis by assuming the homogeneity of the target systems. We implemented an MCAMP optimizer, which takes an input program that consists of perfectly nested loops containing array references and computation codes, and generates optimized communication. We measured the performance of the output programs of the MCAMP optimizer by executing them on a real massively parallel machine GRAPE-DR using a software tool chain that we also implemented. We showed that MCAMP can achieve optimal data transfer patterns and comparable performance to that of hand-optimized codes with a short analysis time.
Hiroshige Hayashizaki, Yutaka Sugawara, Mary Inaba, Kei Hiraki
PACT3
2008 CVC: The C to RTL compiler for callback-based verification model
abstract
Model-based verification is extensively begin used for accelerating the development of embedded system. However, by this approach, a model and actual RTL are required to be implemented separately, which increases the time required to ensure the equivalence of virtual models and actual hardware. To reduce the costs incurred in separate implementations, we propose to directly generate RTL from verification model used in CoMET, which is a callback-based verification environment. We design and implement CVC, a compiler used for generating RTL, using a callback-based verification model described in a subset of the C language; we impose a restriction on CVC to describe the callback efficiently. Our method enables developers to implement the complete RTL without any compromises in the RTL performance just after the verification of the callback-based model is completed.
Yasuhiro Ito, Yutaka Sugawara, Mary Inaba, Kei Hiraki
FPL3
2008 Effect of Parallel TCP Stream Equalizer on Real Long Fat-pipe Network
abstract
With the rapid progress of high-performance cluster applications, data transfer between clusters in distant locations becomes more important. But, it is difficult to transfer data using parallel TCP streams on long distance high bandwidth network. In this paper, we microscopically observe parallel TCP streams on 10 Gbps network using our network analyzer, propose, implement, and evaluate "Stream Equalizer'' which relaxes self-congestion and balances throughput among streams.We evaluate it using a real wide-area network over the Pacific Ocean.The network analyzer and the Stream Equalizer are implemented on FPGA-based programmable high-speed network testbed TGNLE-1.
Yutaka Sugawara, Takeshi Yoshino, Hiroshi Tezuka, Mary Inaba, Kei Hiraki
NCA4
2008 Performance optimization of TCP/IP over 10 gigabit ethernet by precise instrumentation
abstract
End-to-end communications on 10 Gigabit Ethernet (10 GbE) WAN became popular. However, there are difficulties that need to be solved before utilizing Long Fat-pipe Networks (LFNs) by using TCP. We observed that the followings caused performance depression: short-term bursty data transfer, mismatch between TCP and hardware support, and excess CPU load. In this research, we have established systematic methodologies to optimize TCP on LFNs. In order to pinpoint causes of performance depression, we analyzed real networks precisely by using our hardware-based wire-rate analyzer with 100-ns time-resolution. We took the following actions on the basis of the observations: (1) utilizing hardware-based pacing to avoid unnecessary packet losses due to collisions at bottlenecks, (2) modifying TCP to adapt packet coalescing mechanism, (3) modifying programs to reduce memory copies. We have achieved a constant through-put of 9.08 Gbps on a 500 ms RTT network for 5 h. Our approach has overcome the difficulties on single-end 10 GbE LFNs.
Takeshi Yoshino, Yutaka Sugawara, Katsushi Inagami, Junji Tamatsukuri, Mary Inaba, Kei Hiraki
SC5
2007 GRAPE-DR: 2-Pflops massively-parallel computer with 512-core, 512-Gflops processor chips for scientific computing
abstract
We describe the GRAPE-DR (Greatly Reduced Array of Processor Elements with Data Reduction) system, which will consist of 4096 processor chips each with 512 cores operating at the clock frequency of 500 MHz. The peak speed of a processor chip is 512Gflops (single precision) or 256 Gflops (double precision).
Junichiro Makino, Kei Hiraki, Mary Inaba
SC3
2005 High-speed and Memory Efficient TCP Stream Scanning using FPGA
abstract
In this paper, we propose methods to enable high-speed and memory efficient TCP stream level string matching using FPGA. Packet loss and inconsistent retransmissions are handled without dropping packets. Received packets are processed in their arriving order to reduce the buffering memory size. Consistency of retransmission packets is checked using hash value comparison. We evaluate the proposed system using Xilinx XC2VP100-5 FPGA. A 40Gbps network is supported by the proposed system with 140MB memory usage under a realistic traffic pattern. In addition, the proposed system realizes 39.3Gbps packet-processing throughput for a 1017 characters rule, and 1.85Gbps throughput for a 16375 characters rule.
Yutaka Sugawara, Mary Inaba, Kei Hiraki
FPL2
2004 Over 10Gbps String Matching Mechanism for Multi-stream Packet Scanning Systems
Yutaka Sugawara, Mary Inaba, Kei Hiraki
FPL2
2004 Theoretical Analysis of Performances of TCP/IP Congestion Control Algorithm with Different Distances
Tsuyoshi Ito, Mary Inaba
NETWORKING2
2004 Inter-Layer Coordination for Parallel TCP Streams on Long Fat Pipe Networks
abstract
As the network speed grows, inter-layer coordination becomes more important. This paper shows 3 inter-layer coordination methods; (1) "Comet-TCP"; cooperation of data-link layer and transport layer using hardware, (2) "Transmission Rate Controlled TCP (TRC-TCP)"; cooperation of data-link layer and transport layer using software, and (3) "Dulling Edges of Cooperative Parallel streams (DECP)"; cooperation of transport layer and application layer. We show the experimental results of file transfer at Bandwidth Challenge in SC2003; one and a half round trip from Japan to U.S., 15,000 miles, which has 350 ms RTT and 8.2 Gbps bandwidth. Comet-TCP hardware solution attained max 7.56 Gbps using a pair of 16 IA servers, which is 92% of available bandwidth and DECP software attained max 7.01 Gbps using a pair of 32 IA servers.
Hiroyuki Kamezawa, Makoto Nakamura, Junji Tamatsukuri, Nao Aoshima, Mary Inaba, Kei Hiraki
SC5
2002 Data Reservoir: utilization of multi-gigabit backbone network for data-intensive research
abstract
We propose data sharing facility for data intensive scientific research, "Data Reservoir"; which is optimized to transfer huge amount of data files between distant places fully Utilizing multi-gigabit backbone network. In addition, "Data Reservoir" can be used as an ordinary UNIX server in local network without any modification of server softwares. We use low-level protocol and hierarchical striping to realize (1) separation of bulk data transfer and local accesses by cashing, (2) file-system transparency, I.e. interoperable whatever in higher layer than disk driver, including file system. (3) scalability for network and storage. This paper shows our design, implementation using iSCSI protocol [1] and their performances for both 1Gbps model in the real network and 10Gbps model in our laboratory.
Kei Hiraki, Mary Inaba, Junji Tamatsukuri, Ryutaro Kurusu, Yukichi Ikuta, Hisashi Koga, Akira Zinzaki
SC2
1998 Voronoi Diagrams by Divergences with Additive Weights
abstract
No abstract available.
Kunihiko Sadakane, Hiroshi Imai, Kensuke Onishi, Mary Inaba, Fumihiko Takeuchi, Keiko Imai
SCG4
1998 Geometric Clustering Models in Feature Space
Mary Inaba, Hiroshi Imai
Discovery Science1
1997 Application of an Effective Geometric Clustering Method to the Color Quantization Problem
abstract
No abstract available.
Mary Inaba, Hiroshi Imai, Motoki Nakade, Tatsurou Sekiguchi
SCG1
1996 Experimental Results of Randomized Clustering Algorithm
abstract
No abstract available.
Mary Inaba, Hiroshi Imai, Naoki Katoh
SCG1
1996 A Package for Triangulations
abstract
No abstract available.
Tsuyoshi Ono, Yoshiaki Kyoda, Tomonari Masada, Kazuyoshi Hayase, Tetsuo Shibuya, Motoki Nakade, Mary Inaba, Hiroshi Imai, Keiko Imai, David Avis
SCG7
1994 Applications of Weighted Voronoi Diagrams and Randomization to Variance-Based k-Clustering (Extended Abstract)
abstract
In this paper we consider thek-clustering problem for a set S of n points i=(xi) in thed-dimensional space with variance-based errors as clustering criteria, motivated from the color quantization problem of computing a color lookup table for frame buffer display. As the inter-cluster criterion to minimize, the sum on intra-cluster errors over every cluster is used, and as the intra-cluster criterion of a cluster Sj,
Mary Inaba, Naoki Katoh, Hiroshi Imai
SCG1