Haruo Yokota

dblp:y/HaruoYokota · DBLP profile ↗
← Back
68ranked-venue papers
6as first author
8since 2021 · last 2025
0000-0001-9788-0443ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 41 · 4 first-author · 3 since 2021Artificial intelligence and machine learning · 18 · 5 since 2021Security and privacy · 13 · 1 first-authorSoftware engineering, systems software and programming languages · 12 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 11 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4Systems, architecture and hardware · 3 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Computer networks · 1
YearPublicationVenuePosition
2025 Improving the Efficiency of Interactive Sequential Pattern Mining by Closed Pattern Discovery
Yui Aoyagi, Hieu Hanh Le, Ryosuke Matsuo, Tomoyoshi Yamazaki, Kenji Araki, Haruo Yokota, Masato Oguchi
ADMA (4)6
2025 Extracting and Visualizing Frequent Medical Instruction Patterns with Statistical Insights from Multi-Institutional Electronic Medical Record Data
abstract
Despite the widespread adoption of electronic medical records (EMR), the variations in format and terminology across institutions hinder inter-institutional comparisons and feature extraction. This paper proposes a demonstration of extracting frequent disease-specific instruction sequences and efficiently visualizing them with statistical insights, e.g. statistical trends, and abnormal inspection result rates from real multi-institutional EMR data. The utility of the developed visualization tool was presented for decision support and improving clinical processes.
Miwa Sugitani, Ryosuke Matsuo, Tomoyoshi Yamazaki, Kenji Araki, Masato Oguchi, Haruo Yokota, Hieu Hanh Le
CBMS6
2024 Read-safe snapshots: An abort/wait-free serializable read method for read-only transactions on mixed OLTP/OLAP workloads
abstract
This paper proposes Read-Safe Snapshots (RSS), a concurrency control method that ensures reading the latest serializable version on multiversion concurrency control (MVCC) for read-only transactions without creating any serializability anomaly, thereby enhancing the transaction processing throughput under mixed workloads of online transactional processing (OLTP) and online analytical processing (OLAP). Ensuring serializability for data consistency between OLTP and OLAP is vital to prevent OLAP from obtaining nonserializable results. Existing serializability methods achieve this consistency by making OLTP or OLAP transactions aborts or waits, but these can lead to throughput degradation when implemented for large read sets in read-only OLAP transactions under mixed workloads of the recent real-time analysis applications. To deal with this problem, we present an RSS construction algorithm that does not affect the conventional OLTP performance and simultaneously avoids producing additional aborts and waits. Moreover, the RSS construction method can be easily applied to the read-only replica of a multinode system as well as a single-node system because no validation for serializability is required. Our experimental findings showed that RSS could prevent read-only OLAP transactions from creating anomaly cycles under a multinode environment of master-copy replication, which led to the achievement of serializability with the low overhead of about 15% compared to baseline OLTP/OLAP throughputs under snapshot isolation (SI). The OLTP throughput under our proposed method in a mixed OLTP/OLAP workload was about 45% better than SafeSnapshots, a serializable snapshot isolation (SSI) equipped with a read-only optimization method, and did not degrade the OLAP throughput.
Takamitsu Shioi, Takashi Kambayashi, Suguru Arakawa, Ryoji Kurosawa, Satoshi Hikida, Haruo Yokota
Inf. Syst.6
2023 Analysis of Transitions in Differences between Frequent Medical-order Sequences for COVID-19
abstract
With the increasing use of electronic medical records, medical support from analysis of the accumulated medical information is expected. Currently, new treatment methods and drugs are being developed for the treatment of new diseases, but the transition history of medical orders has yet to be visualized for diseases such as COVID-19. In this paper, we use sequential pattern mining to extract frequent medical orders and then apply the longest common subsequence variant (LCSV) and merged sequence variant (MSV) to analyze the differences in treatment patterns at different times. We also propose three types of sliding window (time interval window, sequence number window, and time-sequence number window) to analyze the transition history of medical orders. As an example, we applied these methods to Japanese electronic medical records covering the first to the fifth waves of COVID-19 and analyzed the differences in medical-order patterns for the five infection waves and the transition history of medical orders. We then visualized the difference with MSV. The results showed that the proposed method can successfully visualize the differences in medical orders between infection waves, and the transition history of medical orders can be revealed. The validity of the results was confirmed by the medical staff involved.
Zitai Zhao, Yuki Yasumitsu, Hieu Hanh Le, Tomoyoshi Yamazaki, Kenji Araki, Haruo Yokota
CBMS6
2023 Methods for Analyzing Medical-Order Sequence Variants in Sequential Pattern Mining for Electronic Medical Record Systems
abstract
Electronic medical record systems have been adopted by many large hospitals worldwide, enabling the recorded data to be analyzed by various computer-based techniques to gain a better understanding of hospital-based disease treatments. Among such techniques, sequential pattern mining, already widely used for data mining and knowledge discovery in other application domains, has shown great potential for discovering frequent patterns in sequences of disease treatments. However, studies have yet to evaluate the use of medical-order sequence variants , where a “frequent pattern” can include some limited variations to the pattern, or have considered the factors that lead to these variants. Such a study would be meaningful for medical tasks such as improving the quality of a particular treatment method, comparing treatments with multiple hospitals, recommending the best-suited treatment for each patient, and optimizing the running costs in hospitals. This article proposes methods for evaluating medical-order sequence variants and understanding variant factors based on a statistical approach. We consider the safety and efficiency of sequences and related information about the variants, such as gender, age, and test results from hospitals. Our proposal has been demonstrated as effective by experimentally evaluating an electronic medical record system’s real dataset and obtaining feedback from medical workers. The experimental results indicate that the medical treatment history and specimen test results after hospitalization are significant in identifying the factors that lead to variants.
Hieu Hanh Le, Tatsuhiro Yamada, Yuichi Honda, Takatoshi Sakamoto, Ryosuke Matsuo, Tomoyoshi Yamazaki, Kenji Araki, Haruo Yokota
ACM Trans. Comput. Heal.8
2022 Comparison of Sequence Variants and the Application in Electronic Medical Records
Hieu Hanh Le, Ryosuke Matsuo, Tomoyoshi Yamazaki, Kenji Araki, Haruo Yokota
DEXA (2)6
2021 Sequential Pattern Mining of Large Combinable Items with Values for a Set-of-items Recommendation
abstract
Next-item recommendation solutions based on sequential pattern mining have been widely used in empirical studies. However, current solutions do not consider recommendations involving large combinations of items with varied values. For example, inspecting many specimens is key to understanding a patient's current health status and to checking a medical prescription's effectiveness. Typically, a specimen inspection may involve dozens of inspection items selected from more than a thousand possible items, with each item being associated with a measured value. The values themselves will differ for different items. Based on the pattern of previous item values, recommending the next specimen inspection from a huge number of candidate ones, requires that a combination of many inspection items must be processed efficiently. This paper presents a method for vectorizing a combination of items and item values, and identifying clusters of item-set types. From the item-set types that best suit the target input, a set of items is recommended using both frequency and uniqueness. The method was tested experimentally, using real data from a university hospital's electronic medical record system. The results showed that the proposed method can successfully recommend specific item types with the highest precision and recall ratios. The validity of the results was confirmed by the hospital's medical staff.
Hieu Hanh Le, Yutaka Horino, Tomoyoshi Yamazaki, Kenji Araki, Haruo Yokota
CBMS5
2021 Lightweight Dynamic Redundancy Control with Adaptive Encoding for Server-based Storage
abstract
With the recent performance improvements in commodity hardware, low-cost commodity server-based storage has become a practical alternative to dedicated-storage appliances. Because of the high failure rate of commodity servers, data redundancy across multiple servers is required in a server-based storage system. However, the extra storage capacity for this redundancy significantly increases the system cost. Although erasure coding (EC) is a promising method to reduce the amount of redundant data, it requires distributing and encoding data among servers. There remains a need to reduce the performance impact of these processes involving much network traffic and processing overhead. Especially, the performance impact becomes significant for random-intensive applications. In this article, we propose a new lightweight redundancy control for server-based storage. Our proposed method uses a new local filesystem-based approach that avoids distributing data by adding data redundancy to locally stored user data. Our method switches the redundancy method of user data between replication and EC according to workloads to improve capacity efficiency while achieving higher performance. Our experiments show up to 230% better online-transaction-processing performance for our method compared with CephFS, a widely used alternative system. We also confirmed that our proposed method prevents unexpected performance degradation while achieving better capacity efficiency.
Takayuki Fukatani, Hieu Hanh Le, Haruo Yokota
ACM Trans. Storage3
2019 Centralized Trust Scheme for Cluster Routing of Wireless Sensor Networks
abstract
With the increasing prevalence of Internet-of-Things (IoT) application, Wireless Sensor Networks (WSNs), as a layer in the IoT hierarchy which collects data from the environment, have received wide attention. An efficient category of routing protocols for WSN is cluster routing. However, secured messages routing has long been a concern, as traditional network security practices are inappropriate for WSNs due to their inherent hardware limitations. Hence, this paper introduces and explains a novel centralized statistics-based trust scheme with consideration on power budget, within the context of WSN cluster routing, aiming at detecting unreliable or potentially malicious sensor nodes. The proposed scheme is theoretically able to secure confidentiality, integrity and authenticity of messages in a WSN cluster routing setup, while the simulation results indicate that the computed trust values can correctly identify the node that causes message loss, which helps to improve messages availability.
Nesrine Berjab, Hieu Hanh Le, Haruo Yokota
IEEE BigData4
2019 Energy Efficient Data Placement and Buffer Management for Multiple Replication
Satoshi Hikida, Hieu Hanh Le, Haruo Yokota
DEXA (2)3
2019 Analyzing Sequence Pattern Variants in Sequential Pattern Mining and Its Application to Electronic Medical Record Systems
Hieu Hanh Le, Tatsuhiro Yamada, Yuichi Honda, Masaaki Kayahara, Muneo Kushima, Kenji Araki, Haruo Yokota
DEXA (2)7
2019 Differentially private sequential pattern mining considering time interval for electronic medical record systems
abstract
Electronic medical record (EMR) systems have now been widely adopted to support medical workers. There also has been much interest in the machine-based generation of clinical pathways that can utilize sequential pattern mining (SPM) to extract them from historical EMR systems. However, the existing methods do not protect individual privacy, even though they involve sensitive medical data. To ensure the privacy of individual data, this paper describes two algorithms that deploy differential privacy by adding noise during calculations in the SPM considering time interval for guaranteeing privacy. The proposals can limit the amount of added noise by adding noise to the frequency calculations of only a part of candidate closed sequences. Experiments on real medical datasets show that our proposal can ensure the robust and high utility of mining process even with minimum privacy budget and amount of added noise.
Hieu Hanh Le, Muneo Kushima, Kenji Araki, Haruo Yokota
IDEAS4
2019 Effects of Mining Parameters on the Performance of the Sequence Pattern Variants Analyzing Method Applied to Electronic Medical Record Systems
abstract
Sequential pattern mining (SPM) is widely used for data mining and knowledge discovery in various application domains. Recently, we have proposed an analyzing method to evaluate the sequence pattern variant (SPV) that is the original sequence containing frequent patterns including variants. Such a study is meaningful for medical tasks such as improving the quality of a disease's treatment method. This paper aims to evaluate the effectiveness of the proposed analyzing method in more detail when it was applied to Electronic Medical Record Systems. Using a real dataset, it is observed that the analyzing method is successful in statistically discovering the meaningful indicators that are leading to the difference between comparative SPVs, such as complicated risk, severity risk of the disease, the length of stay in the hospital and the total medical cost. Moreover, it is observed that the length of stay and the medical cost can gain more benefit from increasing the significance level parameter used in comparing the SPVs.
Hieu Hanh Le, Tatsuhiro Yamada, Yuichi Honda, Masaaki Kayahara, Muneo Kushima, Kenji Araki, Haruo Yokota
iiWAS7
2019 Lightweight Dynamic Redundancy Control for Server-Based Storage
abstract
The recent performance improvements in commodity hardware have made commodity server-based storage a practical alternative to dedicated-storage appliances. Because of the low reliability of commodity servers, data redundancy across multiple servers is required for high availability of a server-based storage system. However, the extra storage capacity required to enable this redundancy increases the system cost significantly. Although erasure coding (EC) is a promising approach to reducing the amount of redundant data, it is only available in systems using distributed storage. There remains the need to reduce the performance overhead of using distributed storage, which involves much network traffic and metadata processing. In this paper, we propose a lightweight redundancy-control method called "dynamic redundancy control" for server-based storage systems. Our method adds additional file metadata for EC to the local filesystem, enabling the system to take advantage of EC with reduced network traffic and metadata processing. In addition, our method dynamically controls the data redundancy in user data between replication and EC to improve the capacity efficiency while mitigating performance degradation. Our experiments show that our method achieves up to 76% less processing time for a metadata-intensive workload, up to 150% higher read performance, and up to 147% better online-transaction-processing performance than CephFS, a widely used alternative system. In addition, our method successfully improves capacity efficiency while mitigating performance degradation.
Takayuki Fukatani, Hieu Hanh Le, Haruo Yokota
SRDS3
2018 Hierarchical Abnormal-Node Detection Using Fuzzy Logic for ECA Rule-Based Wireless Sensor Networks
abstract
The Internet of things (IoT) is a distributed, networked system composed of many embedded sensor devices. Unfortunately, these devices are resource constrained and susceptible to malicious data-integrity attacks and failures, leading to unreliability and sometimes to major failure of parts of the entire system. Intrusion detection and failure handling are essential requirements for IoT security. Nevertheless, as far as we know, the area of data-integrity detection for IoT has yet to receive much attention. Most previous intrusion-detection methods proposed for IoT, particularly for wireless sensor networks (WSNs), focus only on specific types of network attacks. Moreover, these approaches usually rely on using precise values to specify abnormality thresholds. However, sensor readings are often imprecise and crisp threshold values are inappropriate. To guarantee a lightweight, dependable monitoring system, we propose a novel hierarchical framework for detecting abnormal nodes in WSNs. The proposed approach uses fuzzy logic in event-condition-action (ECA) rule-based WSNs to detect malicious nodes, while also considering failed nodes. The spatiotemporal semantics of heterogeneous sensor readings are considered in the decision process to distinguish malicious data from other anomalies. Following our experiments with the proposed framework, we stress the significance of considering the sensor correlations to achieve detection accuracy, which has been neglected in previous studies. Our experiments using real-world sensor data demonstrate that our approach can provide high detection accuracy with low false-alarm rates. We also show that our approach performs well when compared to two well-known classification algorithms.
Nesrine Berjab, Hieu Hanh Le, Chia-Mu Yu, Sy-Yen Kuo, Haruo Yokota
PRDC5
2017 A Concurrency Control Protocol that Selects Accessible Replicated Pages to Avoid Latch Collisions for B-Trees in Manycore Environments
Tomohiro Yoshihara, Haruo Yokota
DEXA (2)2
2017 Key Management in Internet of Things via Kronecker Product
abstract
As the number of everyday objects connected to the Internet grows rapidly, securing these connected devices is a big security challenge. Key establishment in Internet of Things (IoT) becomes a challenging problem when considering the resource constrained sensor nodes. In spite of the fact that many clever solutions have been proposed, no practical and suitable scheme has emerged, especially for the extremely large amount of sensor nodes in the wireless sensor network (WSNs) in the future. In this paper, we propose a new key establishment scheme for IoT. The scheme is achieved by Kronecker product and satisfies the following conditions. 1) Substantially decreases the amount of data needs to be stored in a sensor node, 2) efficiently compute the pairwise key, 3) no communication is needed during the computation of the keys. The security evaluation is performed and we also present an in depth analysis of our scheme in terms of computation cost, communication cost and storage cost.
I-Chen Tsai, Chia-Mu Yu, Haruo Yokota, Sy-Yen Kuo
PRDC3
2017 QUILTS: Multidimensional Data Partitioning Framework Based on Query-Aware and Skew-Tolerant Space-Filling Curves
abstract
Recently, massive data management plays an increasingly important role in data analytics because data access is a major bottleneck. Data skipping is a promising technique to reduce the number of data accesses. Data skipping partitions data into pages and accesses only pages that contain data to be retrieved by a query. Therefore, effective data partitioning is required to minimize the number of page accesses. However, it is an NP-hard problem to obtain optimal data partitioning given query pattern and data distribution.
Shoji Nishimura, Haruo Yokota
SIGMOD Conference2
2016 JARS: Join-Aware Distributed RDF Storage
abstract
The enormous increase of data in RDF format calls for efficient storage and retrieval approaches. Being a highly connected data, RDF generates massive amounts of intermediate results during query processing. Many of the current RDF storage approaches involve large amounts of inter-node data movement even for simple selective query patterns. We propose JARS, a join-aware distributed RDF storage system with a dual-hash partitioning strategy coupled with two layered distributed clustered indexing and a rule-based query-execution approach. JARS eliminates the inter-node communication for star patterns and mitigates the communication cost for chain pattern SPARQL queries. Our experiments indicate that JARS achieves significant performance enhancement over the state-of-the-art RDF storage systems.
Anjali Rajith, Shoji Nishimura, Haruo Yokota
IDEAS3
2016 Sequential pattern mining on electronic medical records with handling time intervals and the efficacy of medicines
abstract
It is useful to employ electronic medical records to improve medical studies. Based on their experience, medical workers conventionally prepare clinical pathways as guidelines for the typical flow for the medical treatment of each disease. In this study, we propose an approach for verifying existing clinical pathways and recommend variants or new pathways by analyzing historical records. We propose a method based on the application of sequential pattern mining to record logs with handling time intervals between treatments. We also focus on the efficacy of medicines instead of their names because various medicines have the same efficacy and they change dynamically. We evaluated the proposed method using actual logs and the results demonstrated that the proposed method is effective.
Keishiro Uragaki, Tomoyuki Hosaka, Yoshitaka Arahori, Muneo Kushima, Tomoyoshi Yamazaki, Kenji Araki, Haruo Yokota
ISCC7
2016 A semantics and image retrieval system for hierarchical image databases
Shreelekha Pandey, Pritee Khanna, Haruo Yokota
Inf. Process. Manag.3
2016 Clustering of hierarchical image database to reduce inter-and intra-semantic gaps in visual space for finding specific image semantics
Shreelekha Pandey, Pritee Khanna, Haruo Yokota
J. Vis. Commun. Image Represent.3
2015 An Efficient Gear-Shifting Power-Proportional Distributed File System
Hieu Hanh Le, Satoshi Hikida, Haruo Yokota
DEXA (2)3
2015 An effective use of adaptive combination of visual features to retrieve image semantics from a hierarchical image database
Shreelekha Pandey, Pritee Khanna, Haruo Yokota
J. Vis. Commun. Image Represent.3
2014 CARIC-DA: Core Affinity with a Range Index for Cache-Conscious Data Access in a Multicore Environment
Fang Xi, Takeshi Mishima, Haruo Yokota
DASFAA (1)3
2014 A fragmented data-declustering strategy for high skew tolerance and efficient failure recovery
abstract
Data declustering is a common technique to improve data I/O performance by retrieving data in parallel from multiple storage nodes. Data-declustering methods with replicated data also increase system availability, reliability and skew tolerance. Current replicated declustering schemes fall into two types. In the first, each storage node has its data fully replicated on one other storage node. In the second, each storage node has its replica data fragmented and distributed over several storage nodes. The different schemes have different trade-offs between skew tolerance and data reliability. In this paper, we introduce Fragmented Chained Declustering (FCD) to provide better load balancing and efficient failed-data restoration, with the same or lower data loss rates than previous solutions. Experimental results show the load-balancing and failure recovery efficiency of FCD.
Nam Dang, Haruo Yokota
IDEAS3
2014 A Prototype of Power Saving Storage Method RAPoSDA
abstract
Because of the development of information technologies and the pervasiveness of cloud services, Reducing power consumption in large scale storage systems becomes a very important issue. Recently, we have proposed a method to solve the problem called Replica Assisted Power Saving Disk Array (RAPoSDA) and verified its effectiveness. RAPoSDA utilizes a primary backup configuration on both cache memories and disk drives to ensure system reliability and it dynamically controls the timing and targeting of disk access based on individual disk rotation states. Until now, we have evaluated the effectiveness of RAPoSDA by using simulation program which we had developed. The evaluation shows that RAPoSDA achieves significantly reducing power consumption. However, we have not evaluated it effectiveness on the practical environment. In this paper, we introduce the prototype system of RAPoSDA to evaluate the power saving capability of RAPoSDA.
Satoshi Hikida, Hiroki Oguri, Haruo Yokota
SRDS3
2013 Efficient gear-shifting for a power-proportional distributed data-placement method
abstract
Energy-aware commodity-based distributed file systems for efficient Big Data processing are increasingly moving towards power-proportional designs. However, current data placement methods for such systems have not given careful consideration to the effect of gear-shifting during operations. If the system wants to shift to a higher gear, it must reallocate the updated datasets that were modified in a lower gear when a subset of the nodes was powered off, but without disrupting the servicing of requests from clients. Inefficient gear-shifting that requires a large amount of data reallocation greatly degrades the system performance. This paper proposes a data placement method known as Accordion, which uses data replication to arrange the data layout comprehensively and provide efficient gear-shifting. Compared with current approaches, Accordion reduces the amount of data transferred, which significantly shortens the period required to reallocate the updated data during gear-shifting. The effect of this reduction is larger with higher gears so Accordion is suitable for smooth gear-shifting in multigear systems. Moreover, the times when the active nodes serve the requests are well distributed so Accordion is capable of higher scalability than existing methods based on the I/O throughput performance. Accordion does not require any strict constraint on the number of nodes in the system therefore our proposed method is expected to work well in practical environments. Extensive empirical experiments using actual machines with an Accordion prototype based on the Hadoop Distributed File System demonstrated that our proposed method significantly reduced the period required to transfer updated data, i.e., by 66% compared with an existing method.
Hieu Hanh Le, Satoshi Hikida, Haruo Yokota
IEEE BigData3
2013 NameNode and DataNode Coupling for a Power-Proportional Hadoop Distributed File System
Hieu Hanh Le, Satoshi Hikida, Haruo Yokota
DASFAA (2)3
2013 Finding Image Semantics from a Hierarchical Image Database Based on Adaptively Combined Visual Features
Pritee Khanna, Shreelekha Pandey, Haruo Yokota
DEXA (1)3
2013 A File Recommendation Method Based on Task Workflow Patterns Using File-Access Logs
Takayuki Kawabata, Fumiaki Itoh, Yousuke Watanabe, Haruo Yokota
DEXA (2)5
2013 An Evaluation of Similarity Search Methods Blending Structures and Keywords in XML Documents
abstract
For the past few years, hundreds of document-formats based on XML have appeared. Office documents are typical examples of XML documents. Besides, demands for searching documents become increasing and complicated since we need not only keyword search but also similarity search. In our previous work, we proposed LAX+, an algorithm for measuring a similarity value between XML trees. However, there is a problem that LAX+ performs a rigid matching at leaf-nodes of XML trees. In this paper, we propose two methods: KLAX and LAX&KEY. To measure a precise similarity value between leaf-nodes, KLAX improves LAX+ by-checking the number of common keywords in the leaf-nodes. LAX&KEY separately measures a similarity value between XML trees by LAX+ and a similarity value of common keywords in XML trees, and then combines them. In our experiments with docx, xlsx, and pptx files, the proposed methods yield better results in precision and recall.
Apichaya Auvattanasombat, Yousuke Watanabe, Haruo Yokota
iiWAS3
2012 Data layout management for energy-saving key-value storage using a write off-loading technique
abstract
The objective of our research has been to achieve energy-saving key-value store system. We introduce architecture of our distributed key value store and control framework for energy saving. To reduce power consumption, the control framework should reduce the time required before node power-down. Since this time depends heavily on the amount of data migration, we employ an approach which reduce that amount. We describe and evaluate three control strategies that can be executed on this control framework, each consisting of a data allocation algorithm and data control layout algorithm for use when nodes are to be powered down. With a prototype system developed for evaluation purposes, the time required before node power down is reduced by use of our proposed data allocation algorithm, by means of its taking into account the possible presence of Write off-loading nodes.
Masaki Kan, Dai Kobayashi, Haruo Yokota
CloudCom3
2012 A Power Saving Storage Method That Considers Individual Disk Rotation
Satoshi Hikida, Hieu Hanh Le, Haruo Yokota
DASFAA (2)3
2012 Style-based similarity search for office XML documents
abstract
Recent office documents follow an XML archive format, so they consist of multiple XML files. XML files in office documents include information about page structures and styles such as font, color and position. But, existing text-based search engines do not focus on structure and style of documents. By utilizing them, we can achieve similarity search for office documents based on structures and styles. We propose SOS, a similarity search method based on structures and styles of office documents. To compute a similarity value between office documents, we have to compute similarity values between multiple pairs of XML files in the documents. We also propose LAX+, which is an algorithm to calculate a similarity value for a pair of XML files, by extending existing XML leaf node clustering algorithm. In our experiments, we use docx, xlsx and pptx files and evaluate SOS and LAX+ by precision and recall.
Yousuke Watanabe, Hidetaka Kamigaito, Haruo Yokota
iiWAS3
2011 Data Allocation Based on XML Query Patterns to Reduce Power Consumption
abstract
The amount of electrical power consumed in data centers is increasing, so reducing power consumption is an important issue for cloud computing. An effective approach to reducing power consumption is suitable data allocation in data storage. Previous studies of this approach were mainly based on data-access frequency to decide the allocation of data. We propose a data allocation method based on query-pattern analysis. Because many Internet applications use the XML format to transfer and store data, we study XML queries here. Fortunately, a tree structure in XML is useful to distinguish the chunk of information to be accessed by a query. Our proposed algorithm, XARrP, an XML data Allocation algorithm for Reducing Power consumption, treats queries in XPath format with special symbols and attributes. XARP is separated into four steps: XPath processing, class mining, association, and data allocation. The XPath process retrieves multiple XML classes from query logs, even though some of them may not be expressed exactly. The class mining process analyzes query patterns to discover frequent classes. Infrequent classes are separated into several class sets in the association process. Finally, the frequent classes and the infrequent class sets are stored in different disks to reduce power consumption. Assuming no special hardware but disk drives spinning down after a specific time period, evaluation results show that power consumption is reduced by 32.5% compared with applying a naive striping method and reduced by 10.8% when compared with a method applying a previous XML cache algorithm. The performance expressed as number of transactions processed per unit of power of the proposed method is 2.5 times better than that of striping.
Xuehua Jiang, Yousuke Watanabe, Haruo Yokota
DASC3
2011 An Evaluation of Power-Proportional Data Placement for Hadoop Distributed File Systems
abstract
Power-saving storage in data centers is now gaining much interest because of the increase in power requirements. In particular, with the expansion of services requiring distributed-processing frameworks, power proportionality in power-aware file systems is attracting great attention from academia and industry. The concept of power proportionality is that a system should perform work in proportion to the energy it consumes. To provide this important characteristic, data placement methods that enable a system to operate in multiple gears, each containing a different number of active nodes, have been proposed. Of these, RABBIT, with its power proportional data placement, is a novel method that is implemented over a Hadoop Distributed File System (HDFS) and has been shown to be successful in guaranteeing power proportionality for read only tasks. However, the data layout used in RABBIT does not consider write access occurring when the system operates in low gear and there are inactive nodes. In this paper, to discover an appropriate approach for power-efficient distributed file systems, we evaluate the performance of the RABBIT and PARAID methods on read-only tasks and the cost to system performance of write access in low gear. The skewed data placement in PARAID, which was one of the first methods to suggest power proportionality in disk-based storage systems, is supposed to offer good performance when dealing with a frequently updated dataset.
Hieu Hanh Le, Satoshi Hikida, Haruo Yokota
DASC3
2011 A File Search Method Based on Intertask Relationships Derived from Access Frequency and RMC Operations on Files
Kenichi Otagiri, Yousuke Watanabe, Haruo Yokota
DEXA (1)4
2011 Research history generation using maximum margin clustering of research papers based on metainformation
abstract
Our research aim is the automatic generation of a researcher's research history from research articles published on the internet. Research history generation based on the k-Means clustering algorithm has been proposed in previous work. However, the performance of the k-Means algorithm is unsatisfactory. We propose a method based on Maximum Margin Clustering (MMC). MMC is a new clustering algorithm based on Support Vector Machines (SVM). It is known that MMC is better than existing clustering algorithms such as k-Means. In this paper, we describe how to convert articles into vectors using metainformation about them and how to decide an initial setting for MMC automatically. We demonstrate by experiment that the purity of a method based on MMC is about 0.58 and its entropy is about 0.415. This result is better than that achieved in previous work (purity: 0.35, entropy: 0.47).
Daichi Kato, Haruo Yokota, Taiichi Hashimoto
iiWAS3
2011 Relationship extraction methods based on co-occurrence in web pages and files
abstract
Every day, information on the Web becomes increasingly enriched. Web access is now very useful in many aspects of daily life, particularly for writing documents and programs. In fact, it has become quite usual to edit files while referring to information on the Web. During the file-editing process, we usually visit so many Web pages that we cannot remember all of the relevant ones. Later, if we want to revisit the same Web pages to modify some part of a file, it can be very hard to track down the Web pages originally referred to. In this paper, we propose methods for finding relationships between files and Web pages based on the co-occurrence of data in Web-access logs and file-access logs. These relationships are very useful for revisiting Web pages related to target files. To analyze co-occurrence in these two types of access logs, there are two approaches for merging the logs, involving a trade-off between accuracy and execution time. We call them the Pre-Merge and Post-Merge methods, and we have evaluated these two methods using actual access logs.
Yousuke Watanabe, Haruo Yokota
iiWAS3
2010 Compound Treatment of Chained Declustered Replicas Using a Parallel Btree for High Scalability and Availability
Akitsugu Watanabe, Haruo Yokota
DEXA (2)3
2010 Searching Keyword-lacking Files based on Latent Interfile Relationships
Tetsutaro Watanabe, Takashi Kobayashi 0001, Haruo Yokota
ICSOFT (1)3
2010 Comparing Hadoop and Fat-Btree Based Access Method for Small File I/O Applications
Haruo Yokota
WAIM2
2010 FileSearchCube: A File Grouping Tool Combining Multiple Types of Interfile-Relationships
Yousuke Watanabe, Kenichi Otagiri, Haruo Yokota
WAIM3
2009 A Low-Storage-Consumption XML Labeling Method for Efficient Structural Information Extraction
Wenxin Liang, Akihiro Takahashi, Haruo Yokota
DEXA3
2009 Performance and Reliability of a Revocation Method Utilizing Encrypted Backup Data
abstract
When multiple users access a network storage system for cloud computing, security becomes a key factor in the service, as well as performance and reliability. The "encrypt-on-disk'' scheme effectively protects transmitted and stored data in network storage. However, this scheme has the problem of revocation for shared files. Active revocation is safe but has denial periods to allow immediate reencryption, while lazy revocation has no denial period but is unsafe during the delay. We propose intelligent storage nodes capable of handling active revocation in storage without the denial period by adopting a primary-backup configuration. This approach provides a good combination of security and availability by replication. However, the reencryption process negatively affects the update performance. Delaying the reencryption process and disk write on the backup node improves performance with no ill effect on security and a small decrease of MTTDL for the simple primary-backup configuration. We evaluate the performance of the proposed approaches by experiments, and the reliability by estimation.
Kazuki Takayama, Haruo Yokota
PRDC2
2009 FENECIA: failure endurable nested-transaction based execution of composite Web services with incorporated state analysis
Neila Ben Lakhal, Takashi Kobayashi 0001, Haruo Yokota
VLDB J.3
2008 Superimposed Code-Based Indexing Method for Extracting MCTs from XML Documents
Wenxin Liang, Takeshi Miki, Haruo Yokota
DEXA3
2008 A concurrency control protocol for parallel B-tree structures without latch-coupling for explosively growing digital content
abstract
While shared-nothing parallel infrastructures provide fast processing of explosively growing digital content, managing data efficiently across multiple nodes is important. The value-range partitioning method with parallel B-tree structures in a shared-nothing environment is an efficient approach for handling large amounts of data. To handle large amounts of data, it is also important to provide an efficient concurrency control protocol for the parallel B-tree. Many studies have proposed concurrency control protocols for B-trees, which use latch-coupling. None of these studies has considered that latch-coupling contains a performance bottleneck of sending of messages between processing elements (PEs) in distributed environments because latch-coupling is efficient for a B-tree on a single machine. The only protocol without latch-coupling is the B-link algorithm, but it is difficult to use the B-link algorithm directly on an entire parallel B-tree structure because it is necessary to guarantee the consistency of the side pointers. We propose a new concurrency control protocol named LCFB that requires no latch-coupling in optimistic processes. LCFB reduces the amount of communication between PEs during a B-tree traversal. To detect access path errors in the LCFB protocol caused by removal of latch-coupling, we assign boundary values to each index page. Because a page split may cause page deletion in a Fat-Btree, we also propose an effective method for handling page deletions without latch-coupling. We then combine LCFB with the B-link algorithm within each PE to reduce the cost of Structure Modification Operations (SMOs) in a PE, as a solution to the difficulty of consistency management for the side pointers in a parallel B-tree structure. To compare the performance of the proposed protocol with conventional protocols MARK-OPT, INC-OPT, and ARIES/IM, we implemented them on an autonomous disk system with a Fat-Btree structure. Experimental results in various environments indicate that the system throughput of the proposed protocols is always superior to those of the other protocols, especially in large-scale configurations, and LCFB with the B-link algorithm is effective at higher update ratios.
Tomohiro Yoshihara, Dai Kobayashi, Haruo Yokota
EDBT3
2008 Exploiting Path Information for Syntax-Based XML Subtree Matching in RDBs
abstract
In this paper, we propose two methods exploiting path information, direct-parent based method and full-path based method for syntax-based XML subtree matching in RDBs. In each proposed method, we discuss two ways of using the path information. The one is utilizing the path information after matching the leaf nodes. The other is using the path information together with the PCDATA value of leaf node as the join object. We perform experiments using the real bibliography XML documents stored in RDBs to evaluate the execution time, precision and recall of subtree matching. The experimental results indicate that both the two proposed path-based methods can effectively improve the precision and recall of subtree matching comparing with the original SLAX algorithm.
Wenxin Liang, Haruo Yokota
WAIM2
2007 Dynamic language model adaptation using presentation slides for lecture speech recognition
abstract
We propose a dynamic language model adaptation method that uses the temporal information from lecture slides for lecture speech recognition. The proposed method consists of two steps. First, the language model is adapted with the text information extracted from all the slides of a given lecture. Next, the text information of a given slide is extracted based on temporal information and used for local adaptation. Hence, the language model, used to recognize speech associated with the given slide changes dynamically from one slide to the next. We evaluated the proposed method with the speech data from four Japanese lecture courses. Our experiments show the effectiveness of our
Hiroki Yamazaki, Koji Iwano, Koichi Shinoda, Sadaoki Furui, Haruo Yokota
INTERSPEECH5
2006 Treatment of Laser Pointer and Speech Information in Lecture Scene Retrieval
abstract
We have previously proposed a unified presentation contents search mechanism named UPRISE (unified presentation slide retrieval by impression search engine), and have also proposed a method to use laser pointer information in lecture scene retrieval. In this paper, we discuss the treatment of the laser pointer and speech information, and propose two methods to filter the laser pointer information using keyword occurrence in slides and speech. We also propose weighting schemata with filtered laser pointer information using slide text and speech information. We evaluate our approach by using actual lecture videos and presentation slides
Wataru Nakano, Takashi Kobayashi 0001, Yutaka Katsuyama, Satoshi Naoi, Haruo Yokota
ISM5
2006 Activity Scheduling inWeb-Service Based Workflow Management for Balancing Load and Handling Failures
abstract
In a workflow management system, appropriate allocation of its activities greatly contributes to the improvement of its efficiency. We have proposed OXTHAS, a loadbalancing method of scheduling the activities in workflow management systems using Web-services. The OXTHAS makes the activities to be assigned to appropriate executors based on the estimation of their processing capacity using execution histories in workflow engines. In this paper, we propose a failure-aware re-scheduling method for the OXTHAS using process time-outs under network and system failures. To allocate activities appropriately, the estimated processing capacity and time-out duration are re-calculated with consideration of the penalty for a failure when a process time-out occurs. We then evaluate the effectiveness of the proposed methods through simulations.
Hideyuki Katoh, Takashi Kobayashi 0001, Haruo Yokota
MDM3
2006 An Efficient Commit Protocol Exploiting Primary-Backup Placement in a Distributed Storage System
abstract
Advanced data engineering applications require a large-scale storage system that is both scalable and dependable. In such a system, an atomic commit protocol becomes imperative to ensure the consistency and atomicity of transactions. In this paper we present a new commit protocol, BA-1.5PC, which is well tailored to such distributed storage environments as autonomous disks that use a primary-backup storage schema. The protocol achieves an efficient commit process while also guaranteeing a high dependability by combining several approaches: (1) a low-overhead log mechanism that eliminates blocking disk I/Os, (2) removing the voting phase from commit processing to gain a faster commit process, and (3) a primary-backup assisted recovery strategy to enhance dependability in the presence of possible failures, so that a master failure in the decision phase are not block prepared cohorts of a transaction. Experiments were carried out on a trial version of an autonomous disks system to verify its efficiency. The results indicate that this protocol significantly outperforms several well-known commit protocols in terms of transaction throughput
Xiangyong Ouyang, Tomohiro Yoshihara, Haruo Yokota
PRDC3
2005 VLEI code: An Efficient Labeling Method for Handling XML Documents in an RDB
abstract
A number of XML labeling methods have been proposed to store XML documents in relational databases. However, they have a vulnerable point, in insertion operations. We propose the variable length endless insertable (VLEI) code and apply it to XML labeling to reduce the cost of insertion operations. Results of our experiments indicate that a combination of the VLEI code and Dewey order is effective for handling skewed insertions.
Kazuhito Kobayashi, Wenxin Liang, Dai Kobayashi, Akitsugu Watanabe, Haruo Yokota
ICDE5
2005 Adaptive Lapped Declustering: A Highly Available Data-Placement Method Balancing Access Load and Space Utilization
abstract
This paper proposes a new data-placement method named adaptive overlapped declustering, which can be applied to a parallel storage system using a value range partitioning-based distributed directory and primary-backup data replication, to improve the space utilization by balancing their access loads. The proposed method reduces data skews generated by data migration for balancing access load. While some data-placement methods capable of balancing access load or reducing data skew have been proposed, both requirements satisfied simultaneously. The proposed method also improves the reliability and availability of the system because it reduces recovery time for damaged backups after a disk failure. The method achieves this acceleration by reducing a large amount of network communications and disk I/O. Mathematical analysis shows the efficiency of space utilization under skewed access workloads. Queuing simulations demonstrated that the proposed method halves backup restoration time, compared with the traditional chained declustering method.
Akitsugu Watanabe, Haruo Yokota
ICDE2
2005 A Failure-Aware Model for Estimating and Analyzing the Efficiency of Web Services Compositions
abstract
More and more within the last couple of years, there is a recognition that the Web services composition concept constitute a major breakthrough and revolutionize the way we deal with integrating disparate and distributed computing environments. Yet, to rise to such a position, a chief concern is to guarantee a high-dependability level of the Web services compositions, which is significantly critical, specially in view of the particularities of the Web services environment unpredictability, heterogeneity, autonomy), if confronted with other computing environments. Our present work falls within this context since we tackle the problem of QoS (quality of service) in the Web services context by verifying to what extent fault-tolerant and dynamically-executed Web services compositions are efficiently serving their purposes. And in pursuing this goal, we introduce a novel model that characterizes, estimates and analyzes several QoS properties of dynamically-executed fault-tolerant Web services compositions - namely the reliability and the execution time. Our model allows acquiring more accurate estimations since it confers a paramount importance to the repercussions of failures. In addition, contrary to other QoS estimations models in the Web services context, which use QoS estimations published in UDDIs by the Web services owners/providers, our model computes QoS estimations on the base of the compositions execution observations, where the observation results are collected in a history. Finally, since Web services are stateless, tracking the failures and determining their locations is almost impossible. To overcome this limitation, we propose to attach to each of the composition's component a state. In doing so, obtained estimations can contribute in acquiring more accurate information about the failures locations and can be used later to improve the composition QoS in the future.
Neila Ben Lakhal, Takashi Kobayashi 0001, Haruo Yokota
PRDC3
2004 Availabilities and Costs of Reliable Fat-Btrees
abstract
The Fat-Btree is expected to be used as a parallel directory structure for high performance and highly reliable data intensive systems, such as databases and file systems. Though the Fat-Btree is essentially high performance, it needs to be highly reliable as well. In order to achieve high reliability, we introduce five configuration methods. We show not only their steady-state and computational availabilities but also cost related comparisons, so as to evaluate the properties of these configurations.
Jun Miyazaki, Yohei Abe, Haruo Yokota
PRDC3
2001 Automatic Reconfiguration of an Autonomous Disk Cluster
abstract
Recently, storage-centric configurations, such as NAS and SAN architectures, have attracted attention in advanced data processing. For these configurations, scalability, flexibility, and availability are key features, and central control is unsuitable. We propose autonomous disks to enable distributed control in the storage-centric configurations. Autonomous disks configure a cluster in a network, and implement the above key features using ECA rules with a distributed directory. The combination of the rules and directory also manages load distribution, and has a capability of reconfiguring the cluster automatically. We focus on the cluster reconfiguration, utilizing the skew handling mechanism. Autonomous disks adopt a multi-phase synchronization protocol with a dynamic coordinator for reconfiguration, apart from data migration. They allow the usual operations to be accepted during reconfiguration. We also report preliminary results of our experimental system. They indicate that synchronization cost is adequately acceptable.
Daisuke Ito, Haruo Yokota
PRDC2
2000 A performance comparison between the DR-net and a hierarchical RAID system
abstract
When the number of disks rises in disk array systems that contain multiple disk drives, system performance is limited by a bottleneck at a centralized controller and/or at a communication path that uses a bus. We evaluate, through simulation, a scalable architecture called DR-net, in which the controller functions are distributed to all disk drives and each disk has autonomy in processing its tasks. DR-net provided high throughput in proportion to the number of disks. In a conventional system, the influence of the bus setup time and concentration of the parity calculation load caused the throughput to saturate. We also show that DR-net can take advantage of disk autonomy when reconstructing data stored in failed disks. Our results indicate that the distribution of functions and the autonomy of disk drives enable better scalability and more effective utilization of system resources than with a hierarchical system.
Yasuyuki Mimatsu, Haruo Yokota
PRDC2
1999 Fat-Btree: An Update-Conscious Parallel Directory Structure
abstract
We propose a parallel directory structure, Fat-Btree, to improve high speed access for parallel database systems in shared nothing environments. The Fat-Btree has a threefold aim: to provide an indexing mechanism for fast retrieval in each processor; to balance the amount of data among distributed disks, and to reduce synchronization costs between processors during update operations. We use a probability based model to compare the throughput and response time of the Fat-Btree with two ordinary parallel Btree structures, with copies of a whole Btree in each processor and storing index nodes in a processor. The comparison results indicate that the Fat-Btree is suitable for actual parallel database systems that accept update operations.
Haruo Yokota, Yasuhiko Kanemasa, Jun Miyazaki
ICDE1
1999 FBD: A Fault-tolerant Buffering Disk System for Improving Write Performance of RAID5 Systems
abstract
The parity calculation technique of the RAID5 provides high reliability, efficient disk space usage, and good read performance for parallel-disk-array configurations. However, it requires four disk accesses for each write request. The write performance of a RAID5 is therefore poor compared with its read performance. We propose a buffering system to improve write performance while maintaining the reliability of the total system. The buffering system uses two to four disks clustered into two groups, primary and backup. Write performance is improved by sequential accesses of the primary disks without interruption and by reduction of irrelevant disk accesses for the RAID5 by packing. The backup disks are used to tolerate a disk failure and to accept read requests for data stored in the buffering system so as not to disturb sequential accesses in the primary disks. We developed an experimental system using an off-the-shelf personal computer and disks, and a commercial RAID5 system. The experiments indicate that the buffering system considerably improves both system throughput and average response time.
Haruo Yokota, Masanori Goto
PRDC1
1989 Term Indexing for Retrieval by Unification
abstract
A method is presented for indexing terms in a knowledge-base retrieval-by-unification (RBU) system. The term is a well-defined structure which represents knowledge using variables. RBU operations are an extension of relational database operations using unification and backtracking to retrieve terms from term relations. The term indexing proposed uses hashing and trie structures to reduce the number of comparisons between elements of a search condition and of an object term relation. Unification on a trie structure is suited to backtracking bindings of variables. The search and updating speed of an RBU prototype is measured to evaluate the indexing method. This method is effective for fast term retrieval for a large number of similar and varied form terms. The overhead for maintaining indexes in updating is low.>
Haruo Yokota, Hajime Kitakami, Akira Hattori
ICDE1
1986 Deductive Database System based on Unit Resolution
abstract
This paper presents a methodology for constructing a deductive database system consisting of an intensional processor and a relational database management system. A setting evaluation approach is introduced. The intensional processor derives a setting from the in-tensional database and a given goal and sends the setting and the relationship between setting elements to the management system. The management system performs a unit resolution with setting using relational operations for the extensional databases. An extended least fixed point operation is introduced to terminate all types of recursive queries.
Haruo Yokota, Sko Sakai, Hidenori Itoh
ICDE1
1986 A Model and an Architecture for a Relational Knowledge Base
abstract
A relational knowledge base model and an architecture which manipulates the model are presented. An item stored in the relational knowledge base is called a term. A unification operation on terms in the relational knowledge base is used as the retrieval mechanism. The relational knowledge base architecture we propose consists of a number of unification engines, several disk systems, a control processor, and a multiport page-memory. The system has a knowledge compiler to support a variety of knowledge representations.
Haruo Yokota, Hidenori Itoh
ISCA1
1986 Retrieval-By-Unification Operation on a Relational Knowledge Base
Yukihiro Morita, Haruo Yokota, Kenji Nishida, Hidenori Itoh
VLDB2
1984 An Enhanced Inference Mechanism for Generating Relational Algebra Queries
abstract
A system for interfacing Prolog programs with relational algebra is presented. The system produces relational algebra queries using a deferred evaluation approach. Least fixed point (LFP) queries are automatically managed. An optimization method for removing redundant relations is also presented.
Haruo Yokota, Susumu Kunifuji, Takeo Kakuta, Nobuyoshi Miyazaki, Shigeki Shibayama, Kunio Murakami
PODS1
1983 A Relational Data Base Machine: First Step to Knowledge Base Machine
abstract
The Japan's Fifth Generation Computer System project is divided into three stages. In the first three-year stage, a working relational data base machine is developed for a software development support system to be used in the second stage and also for an experimental system which provides a research tool for the knowledge base machine.
Kunio Murakami, Takeo Kakuta, Nobuyoshi Miyazaki, Shigeki Shibayama, Haruo Yokota
ISCA5