VLDB 2026 Research / reviewers in the wild / expert
Jun'ichi Tatemura
dblp:84/284
· DBLP profile ↗
48ranked-venue papers
9as first author
2since 2021 · last 2023
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 35 · 5 first-author · 2 since 2021Artificial intelligence and machine learning · 3Systems, architecture and hardware · 3Human-computer interaction and ubiquitous computing · 3 · 3 first-authorComputer networks · 2Software engineering, systems software and programming languages · 2 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2Applied, interdisciplinary, general and emerging computing · 2
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
26 papers |
Query processing and optimization · 32% Database system architecture and tuning · 20% Distributed and cloud data management · 14% | |
| Computer architecture, parallel and distributed computing, and storage systems
4 papers |
Storage systems · 61% Distributed systems · 26% Cloud and datacenter computing · 14% |
Topics — the 30 heaviest of 59, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Query processing and optimization
parallel query processing |
0.7 | 1 | 2023 | Progressive Partitioning for Parallelized Query Execution in Google's Napa · Proc. VLDB Endow. 2023 |
Query processing and optimization
view maintenance |
0.6 | 2 | 2021 | Napa: Powering Scalable Data Warehousing with Robust Query Performance at Google · Proc. VLDB Endow. 2021 Incremental Maintenance of Path Expression Views · SIGMOD Conference 2005 |
Data integration and cleaning
data warehouse |
0.5 | 1 | 2021 | Napa: Powering Scalable Data Warehousing with Robust Query Performance at Google · Proc. VLDB Endow. 2021 |
Distributed and cloud data management
geo-distributed data management |
0.5 | 1 | 2021 | Napa: Powering Scalable Data Warehousing with Robust Query Performance at Google · Proc. VLDB Endow. 2021 |
Indexing and storage engines
key-value store |
0.2 | 2 | 2014 | Partiqle: an elastic SQL engine over key-value stores · SIGMOD Conference 2012 SAGE: A logical and physical design tool for entity-group based new SQL systems · ICDE 2014 |
Query processing and optimization › materialized view
materialized view reuse |
0.2 | 1 | 2014 | Opportunistic physical design for big data analytics · SIGMOD Conference 2014 |
Database system architecture and tuning › database design
physical database design |
0.2 | 1 | 2014 | MISO: souping up big data query processing with a multistore system · SIGMOD Conference 2014 |
Query processing and optimization
query rewriting |
0.2 | 1 | 2014 | Opportunistic physical design for big data analytics · SIGMOD Conference 2014 |
Query processing and optimization
cardinality estimation |
0.2 | 1 | 2013 | Predicting query execution time: Are optimizer cost models really unusable? · ICDE 2013 |
Query processing and optimization › cost estimation
query execution time prediction |
0.2 | 1 | 2013 | Predicting query execution time: Are optimizer cost models really unusable? · ICDE 2013 |
Database theory › conjunctive query
tree pattern query |
0.1 | 2 | 2008 | Scalable Filtering of Multiple Generalized-Tree-Pattern Queries over XML Streams · IEEE Trans. Knowl. Data Eng. 2008 Twig2Stack: Bottom-up Processing of Generalized-Tree-Pattern Queries over XML Documents · VLDB 2006 |
Transaction processing and concurrency control › OLTP
OLTP engine |
0.1 | 1 | 2012 | Partiqle: an elastic SQL engine over key-value stores · SIGMOD Conference 2012 |
Query processing and optimization
query optimization |
0.1 | 1 | 2011 | Towards Cost-Effective Storage Provisioning for DBMSs · Proc. VLDB Endow. 2011 |
Storage systems › storage management › storage resource management
storage provisioning |
0.1 | 1 | 2011 | Towards Cost-Effective Storage Provisioning for DBMSs · Proc. VLDB Endow. 2011 |
Web and social media mining › social media analysis
blog analysis |
0.1 | 2 | 2009 | Structural and temporal analysis of the blogosphere through community factorization · KDD 2007 Efficient overlap and content reuse detection in blogs and online news articles · WWW 2009 |
Natural language and speech › Information extraction and text analysis
web information extraction |
0.1 | 1 | 2009 | Extracting data records from the web using tag path clustering · WWW 2009 |
Information retrieval › similarity search
near-duplicate detection |
0.1 | 1 | 2009 | Efficient overlap and content reuse detection in blogs and online news articles · WWW 2009 |
Information retrieval › document processing › document analysis
text reuse detection |
0.1 | 1 | 2009 | Efficient overlap and content reuse detection in blogs and online news articles · WWW 2009 |
Data stream processing
complex event processing |
0.1 | 1 | 2008 | Runtime Semantic Query Optimization for Event Stream Processing · ICDE 2008 |
Query processing and optimization › adaptive query processing
dynamic query optimization |
0.1 | 1 | 2008 | Runtime Semantic Query Optimization for Event Stream Processing · ICDE 2008 |
Query processing and optimization
OLAP |
0.1 | 1 | 2008 | Supporting OLAP operations over imperfectly integrated taxonomies · SIGMOD Conference 2008 |
Data models and query languages › query interface
query by example |
0.1 | 1 | 2008 | UQBE: uncertain query by example for web service mashup · SIGMOD Conference 2008 |
Data integration and cleaning
schema matching |
0.1 | 1 | 2008 | UQBE: uncertain query by example for web service mashup · SIGMOD Conference 2008 |
Data integration and cleaning › semantic integration
taxonomy integration |
0.1 | 1 | 2008 | Supporting OLAP operations over imperfectly integrated taxonomies · SIGMOD Conference 2008 |
Data integration and cleaning › schema matching
uncertain schema matching |
0.1 | 1 | 2008 | UQBE: uncertain query by example for web service mashup · SIGMOD Conference 2008 |
Data stream processing
XML stream processing |
0.1 | 1 | 2008 | Scalable Filtering of Multiple Generalized-Tree-Pattern Queries over XML Streams · IEEE Trans. Knowl. Data Eng. 2008 |
Services computing and microservices › service composition
service mashup |
0.1 | 1 | 2008 | UQBE: uncertain query by example for web service mashup · SIGMOD Conference 2008 |
Data mining
clustering |
0.1 | 1 | 2007 | Structural and temporal analysis of the blogosphere through community factorization · KDD 2007 |
Data mining › structured data mining › graph mining
community detection |
0.1 | 1 | 2007 | Structural and temporal analysis of the blogosphere through community factorization · KDD 2007 |
Data stream processing
continuous query processing |
0.1 | 1 | 2007 | Mashup Feeds: : continuous queries over web services · SIGMOD Conference 2007 |
Methods — techniques the papers use, named apart from their topics
load balancing · 0.7b-tree statistics · 0.7multi-datacenter replication · 0.5materialized view maintenance · 0.5workload-driven design · 0.2workload characterization · 0.2view rewriting · 0.2semantic modeling of UDFs · 0.2performance comparison · 0.2online tuning · 0.2iterative user feedback · 0.2formal methods · 0.2cost modeling · 0.2service level agreement modeling · 0.1heuristic optimization · 0.1visual signal similarity · 0.1tag path clustering · 0.1optimization · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Progressive Partitioning for Parallelized Query Execution in Google's NapaabstractNapa holds Google's critical data warehouses in log-structured merge trees for real-time data ingestion and sub-second response for billions of queries per day. These queries are often multi-key look-ups in highly skewed tables and indexes. In our production experience, only progressive query-specific partitioning can achieve Napa's strict query latency SLOs. Here we advocate good-enough partitioning that keeps the per-query partitioning time low without risking uneven work distribution. Our design combines pragmatic system choices and algorithmic innovations. For instance, B-trees are augmented with statistics of key distributions, thus serving the dual purpose of aiding lookups and partitioning. Furthermore, progressive partitioning is designed to be "good enough" thereby balancing partitioning time with performance. The resulting system is robust and successfully serves day-in-day-out billions of queries with very high quality of service forming a core infrastructure at Google. Jun'ichi Tatemura, Tao Zou 0002, Jagan Sankaranarayanan, Yanlai Huang, Jim Chen, Hao Zhang 0029, Gokul Nath Babu Manoharan, Goetz Graefe, Divyakant Agrawal, Brad Adelberg, Shilpa Kolhar, Indrajit Roy 0001 |
Proc. VLDB Endow. | 1 |
| 2021 | Napa: Powering Scalable Data Warehousing with Robust Query Performance at GoogleabstractGoogle services continuously generate vast amounts of application data. This data provides valuable insights to business users. We need to store and serve these planet-scale data sets under the extremely demanding requirements of scalability, sub-second query response times, availability, and strong consistency; all this while ingesting a massive stream of updates from applications used around the globe. We have developed and deployed in production an analytical data management system, Napa, to meet these requirements. Napa is the backend for numerous clients in Google. These clients have a strong expectation of variance-free, robust query performance. At its core, Napa's principal technologies for robust query performance include the aggressive use of materialized views, which are maintained consistently as new data is ingested across multiple data centers. Our clients also demand flexibility in being able to adjust their query performance, data freshness, and costs to suit their unique needs. Robust query processing and flexible configuration of client databases are the hallmark of Napa design. Most of the related work in this area takes advantage of full flexibility to design the whole system without the need to support a diverse set of preexisting use cases. In comparison, a particular challenge we faced is that Napa needs to deal with hard constraints from existing applications and infrastructure, so we could not do a "green field" system, but rather had to satisfy existing constraints. These constraints led us to make particular design decisions and also devise new techniques to meet the challenges. In this paper, we share our experiences in designing, implementing, deploying, and running Napa in production with some of Google's most demanding applications. Ankur Agiwal, Gokul Nath Babu Manoharan, Indrajit Roy 0001, Jagan Sankaranarayanan, Hao Zhang 0029, Tao Zou 0002, Jim Chen, Thanh Do, Haoyan Geng, Raman Grover, Yanlai Huang, Adam Li, Jianyi Liang, Xi Mao, Maya Meng, Prashant Mishra, Rajesh Sr, Vijayshankar Raman, Sourashis Roy, Mayank Singh Shishodia, Tianhang Sun, Justin Tang, Jun'ichi Tatemura, Sagar Trehan, Ramkumar Vadali, Prasanna Venkatasubramanian, Joey Zhang, Zeleng Zhuang, Goetz Graefe, Divyakant Agrawal, Jeffrey F. Naughton, Sujata Kosalge, Hakan Hacigümüs |
Proc. VLDB Endow. | 32 |
| 2016 | Strudel: A Framework for Transaction Performance Analyses on SQL/NoSQL SystemsabstractThe paper introduces Strudel, a development and execution framework for transactional workloads both on SQL and NoSQL systems. Whereas a rich set of benchmarks and performance analysis platforms have been developed for SQLbased systems (RDBMSs), it is challenging for application developers to evaluate both SQL and NoSQL systems for their specific needs. The Strudel framework, which we have released as open-source software, helps such developers (as well as providers of NoSQL stores) to build, customize, and share benchmarks that can run on various SQL/NoSQL systems. We describe Strudel’s architecture and APIs, its components for supporting various NoSQL stores (e.g., HBase, MongoDB), example benchmarks included in the release, and performance experiments to demonstrate usefulness of the framework. Jun'ichi Tatemura, Oliver Po, Hakan Hacigümüs |
EDBT | 1 |
| 2015 | Transactional Replication in Hybrid Data Store ArchitecturesabstractWe present a transactional and concurrent replication scheme that is designed for hybrid data store architectures. The system design and the requirements are motivated by the real business cases we encountered during the development of our commercial database product. We consider two databases where the original database handles read/write transactional application workloads while the second database handles read-only workloads from the same applications over the data periodically replicated from the original database. The main requirement is ensuring the application of the updates on the replica database in the exact same order they were executed in the original database, which is called executiondefined order. Although this requirement could easily be satisfied by the serial execution of the updates in the commit order, doing so in an ecient manner by exploiting concurrency is a challenging problem. We present a novel concurrency control algorithm to addresses that problem by also allowing the read-only workloads on the replica database to interleave with the concurrent replication. The extensive experiments show the ecacy of the proposed solution. Hojjat Jafarpour, Jun'ichi Tatemura, Hakan Hacigümüs |
EDBT | 2 |
| 2014 | SAGE: A logical and physical design tool for entity-group based new SQL systemsabstractEntity-group based new SQL systems achieve scalability and consistency at the same time by using a key-value store as the storage layer and limiting each transaction's boundary to a collection of data (called an entity-group). Examples of such systems are Google's Megastore, NEC's Partiqle, and LinkedIn's Espresso. Application developers of such systems face tremendous challenges, both in designing entity-groups (and hence transaction boundaries) and physical layout of data in key-value stores. Entity-group designs directly impact consistency semantics of the workload, and physical layout impacts the application throughput. Both problems are challenging for users to solve manually. In this demonstration, we show a system that solves both problems with a user-friendly GUI built on our principled formal methods. We demonstrate the system using a simple Auction benchmark, and a more complex TPC-W benchmark. Wang-Pin Hsiung, Jun'ichi Tatemura, Hakan Hacigümüs |
ICDE | 3 |
| 2014 | Automatic entity-grouping for OLTP workloadsabstractSupporting an online transaction processing (OLTP) workload in a scalable and elastic fashion is a challenging task. Recently, a new breed of scalable systems have shown significant throughput gains by limiting consistency to small units of data called “entity-groups” (e.g., a user's account information stored together with all her emails in an online email service.) Transactions that access the data from only one entity-group are guaranteed of full ACID, but those that access multiple entity-groups are not. Defining entity-groups has direct impact on workload consistency and performance, and doing so for data with a complex schema is very challenging. It is prone to go to extremes - groups that are too fine-grained cause excessive number of expensive distributed transactions while those that are too coarse lead to excessive serialization and performance degradation. It is also difficult to balance conflicting requirements from different transactions. In commercially available entity-group systems, creating entity-groups is usually a manual process, which severely limits the usability of those systems. This paper is the first systematic effort on automating the entity-group design process. Our goal is to build a user-friendly design tool for automatically creating entity-groups based on a given workload and to help users trade consistency for performance in a principled manner. For advanced users, we allow them to provide feedback to the entity-group design and iteratively improve the final output. We demonstrate the effectiveness of our approach with widely used benchmarks. We also present the user experience of a prototype we built. Jun'ichi Tatemura, Oliver Po, Wang-Pin Hsiung, Hakan Hacigümüs |
ICDE | 2 |
| 2014 | Opportunistic physical design for big data analyticsabstractBig data analytical systems, such as MapReduce, perform aggressive materialization of intermediate job results in order to support fault tolerance. When jobs correspond to exploratory queries submitted by data analysts, these materializations yield a large set of materialized views that we propose to treat as an opportunistic physical design. We present a semantic model for UDFs that enables effective reuse of views containing UDFs along with a rewrite algorithm that provably finds the minimum-cost rewrite under certain assumptions. An experimental study on real-world datasets using our prototype based on Hive shows that our approach can result in dramatic performance improvements. Jeff LeFevre, Jagan Sankaranarayanan, Hakan Hacigümüs, Jun'ichi Tatemura, Neoklis Polyzotis, Michael J. Carey 0001 |
SIGMOD Conference | 4 |
| 2014 | MISO: souping up big data query processing with a multistore systemabstractMultistore systems utilize multiple distinct data stores such as Hadoop's HDFS and an RDBMS for query processing by allowing a query to access data and computation in both stores. Current approaches to multistore query processing fail to achieve the full potential benefits of utilizing both systems due to the high cost of data movement and loading between the stores. Tuning the physical design of a multistore, i.e., deciding what data resides in which store, can reduce the amount of data movement during query processing, which is crucial for good multistore performance. In this work, we provide what we believe to be the first method to tune the physical design of a multistore system, by focusing on which store to place data. Our method, called MISO for MultISstore Online tuning, is adaptive, lightweight, and works in an online fashion utilizing only the by-products of query processing, which we term as opportunistic views. We show that MISO significantly improves the performance of ad-hoc big data query processing by leveraging the specific characteristics of the individual stores while incurring little additional overhead on the stores. Jeff LeFevre, Jagan Sankaranarayanan, Hakan Hacigümüs, Jun'ichi Tatemura, Neoklis Polyzotis, Michael J. Carey 0001 |
SIGMOD Conference | 4 |
| 2014 | Re-evaluating designs for multi-tenant OLTP workloads on SSD-basedI/O subsystemsabstractMulti-tenancy is a common practice that is employed to maximize server resources and reduce the total cloud operation costs. The focus of this work is on multi-tenancy for OLTP workloads. Several designs for OLTP multi-tenancy have been proposed that vary the trade-offs made between performance and isolation. However, existing studies have not considered the impact of OLTP multi-tenancy designs when using an SSD-based I/O subsystem. In this paper, we compare three designs using both open-source and proprietary DBMSs on SSD-based I/O subsystems. Our study reveals that in contrast to the case of an HDD-based I/O subsystem, VM-based designs have fairly competitive performance compared to the non-virtualized designs (generally within 1.3-2X of the best performing case) on SSD-based I/O subsystems. Whereas previous studies were based on traditional hard disk-based environments, our results indicate that switching to a pure SSD-based I/O subsystem requires rethinking the trade-offs for multi-tenant OLTP workloads. Ning Zhang 0013, Jun'ichi Tatemura, Jignesh M. Patel, Hakan Hacigümüs |
SIGMOD Conference | 2 |
| 2014 | Toward cost-effective storage provisioning for DBMSs
Ning Zhang 0013, Jun'ichi Tatemura, Jignesh M. Patel, Hakan Hacigümüs |
VLDB J. | 2 |
| 2013 | Predicting query execution time: Are optimizer cost models really unusable?abstractPredicting query execution time is useful in many database management issues including admission control, query scheduling, progress monitoring, and system sizing. Recently the research community has been exploring the use of statistical machine learning approaches to build predictive models for this task. An implicit assumption behind this work is that the cost models used by query optimizers are insufficient for query execution time prediction. In this paper we challenge this assumption and show while the simple approach of scaling the optimizer's estimated cost indeed fails, a properly calibrated optimizer cost model is surprisingly effective. However, even a well-tuned optimizer cost model will fail in the presence of errors in cardinality estimates. Accordingly we investigate the novel idea of spending extra resources to refine estimates for the query plan after it has been chosen by the optimizer but before execution. In our experiments we find that a well calibrated query optimizer model along with cardinality estimation refinement provides a low overhead way to provide estimates that are always competitive and often much better than the best reported numbers from the machine learning approaches. Wentao Wu 0001, Yun Chi, Shenghuo Zhu, Jun'ichi Tatemura, Hakan Hacigümüs, Jeffrey F. Naughton |
ICDE | 4 |
| 2013 | Odyssey: A Multi-Store System for Evolutionary AnalyticsabstractNo abstract available. Hakan Hacigümüs, Jagan Sankaranarayanan, Jun'ichi Tatemura, Jeff LeFevre, Neoklis Polyzotis |
Proc. VLDB Endow. | 3 |
| 2012 | Towards principled design support for scalable OLTP workloadsabstractSupporting online transaction processing (OLTP) workload in a scalable and elastic fashion is a challenging task. With the advent of cloud-based systems, supporting entity group based consistency is a viable, scalable, and cost-effective option. This approach remains attractive in the presence of systems supporting the highest level of consistency, due to the relative high cost and performance degradation of the latter. In this paper, we briefly introduce our on-going work for assisting application developers to design OLTP workload for entity group based systems. The goal is providing a suite of user-friendly design tools for new-breed databases to achieve scalability and elasticity. Jun'ichi Tatemura, Hakan Hacigümüs |
EDBT | 2 |
| 2012 | Partiqle: an elastic SQL engine over key-value storesabstractThe demo features Partiqle, a SQL engine over key-value stores as a relational alternative for the recent procedural approaches to support OLTP workloads elastically. Based on our microsharding framework [12], it employs a declarative specification, called transaction classes, of constraints applied on the transactions in a workload. We demonstrate use of a transaction class in design and analysis of OLTP workloads. We then demonstrate live-scaling of our fully functioning system on a server cluster. Jun'ichi Tatemura, Oliver Po, Wang-Pin Hsiung, Hakan Hacigümüs |
SIGMOD Conference | 1 |
| 2012 | Performance Evaluation of Range Queries in Key Value Stores
Pouria Pirzadeh, Jun'ichi Tatemura, Oliver Po, Hakan Hacigümüs |
J. Grid Comput. | 2 |
| 2011 | ActiveSLA: a profit-oriented admission control framework for database-as-a-service providersabstractThe system overload is a common problem in a Database-as-a-Serice (DaaS) environment because of unpredictable and bursty workloads from various clients. Due to the service delivery nature of DaaS, such system overload usually has direct economic impact on the service provider, who has to pay penalties if the system performance does not meet clients' service level agreements (SLAs). In this paper, we investigate techniques that prevent system overload by using admission control. We propose a profit-oriented admission control framework, called ActiveSLA, for DaaS providers. ActiveSLA is an end-to-end framework that consists of two components. First, a prediction module estimates the probability for a new query to finish the execution before its deadline. Second, based on the predicted probability, a decision module determines whether or not to admit the given query into the database system. The decision is made with the profit optimization objective, where the expected profit is derived from the service level agreements between a service provider and its clients. We present extensive real system experiments with standard database benchmarks, under different traffic patterns, DBMS settings, and SLAs. The results demonstrate that ActiveSLA is able to make admission control decisions that are both more accurate and more profit-effective than several state-of-the-art methods. PengCheng Xiong, Yun Chi, Shenghuo Zhu, Jun'ichi Tatemura, Calton Pu, Hakan Hacigümüs |
SoCC | 4 |
| 2011 | SLA-tree: a framework for efficiently supporting SLA-based decisions in cloud computingabstractAs cloud computing becomes increasingly important in database systems, many new challenges and opportunities have arisen. One challenge is that in cloud computing, business profit plays a central role. Hence, it is very important for a cloud service provider to quickly make profit-oriented decisions. In this paper, we propose a novel data structure, called SLA-tree, to efficiently support profit-oriented decision making. SLA-tree is built on two pieces of information: (1) a set of buffered queries waiting to be executed, which represents the scheduled events that will happen in the near future, and (2) a service level agreement (SLA) for each query, which indicates the different profits for the query for varying query response times. By constructing the SLA-tree, we efficiently support the answering of certain profit-oriented "what if" questions. Answers to these questions in turn can be applied to different profit-oriented decisions in cloud computing such as profit-aware scheduling, dispatching, and capacity planning. Extensive experimental results based on both synthetic and real-world data demonstrate the effectiveness and efficiency of our SLA-tree framework. Yun Chi, Hyun Jin Moon, Hakan Hacigümüs, Jun'ichi Tatemura |
EDBT | 4 |
| 2011 | COSMOS: A Platform for Seamless Mobile Services in the CloudabstractThe mobility of today is defined by the multitude of apps, which while working in isolation, can achieve a variety of tasks for the mobile user. The mobility of tomorrow is envisioned as one where mobile apps work together by sharing information to create a seamless mobile experience, where the focus is the mobility of the user but not the device. A Platform as a Service (PaaS) system called COSMOS (stands for Clouddb for Seamless Mobile Services) is proposed to provide the necessary support for seamless mobility. COSMOS is a multitenant, SLA-aware, cloud based PaaS system, which is currently under active development. The core component of COSMOS is the Sharing Middleware (SMILE), which provides the infrastructure for mobile apps residing on COSMOS to share data actively with one another. SMILE allows for management of SLAs on the shared data, which means that some serious technical challenges will have to be overcome in order to guarantee the desired level of access on the shared data to all those who access it. The key challenge is in ensuring that SLA guarantees are provided in the face of multiple users with diverse workloads and SLA requirements, while providing performance guarantees for the data owners in sharing data with others using performance isolation policies. Additional services for mobile apps, such as mobile context, recommendation and analytics are proposed by leveraging on SMILE. The challenges in designing COSMOS and SMILE as well as solution strategies are discussed. Jagan Sankaranarayanan, Hakan Hacigümüs, Jun'ichi Tatemura |
Mobile Data Management (1) | 3 |
| 2011 | Towards Cost-Effective Storage Provisioning for DBMSsabstractData center operators face a bewildering set of choices when considering how to provision resources on machines with complex I/O subsystems. Modern I/O subsystems often have a rich mix of fast, high performing, but expensive SSDs sitting alongside with cheaper but relatively slower (for random accesses) traditional hard disk drives. The data center operators need to determine how to provision the I/O resources for specific workloads so as to abide by existing Service Level Agreements (SLAs), while minimizing the total operating cost (TOC) of running the workload, where the TOC includes the amortized hardware costs and the run time energy costs. The focus of this paper is on introducing this new problem of TOC-based storage allocation, cast in a framework that is compatible with traditional DBMS query optimization and query processing architecture. We also present a heuristic-based solution to this problem, called DOT. We have implemented DOT in PostgreSQL, and experiments using TPC-H and TPC-C demonstrate significant TOC reduction by DOT in various settings. Ning Zhang 0013, Jun'ichi Tatemura, Jignesh M. Patel, Hakan Hacigümüs |
Proc. VLDB Endow. | 2 |
| 2010 | CloudDB: One Size Fits All RevivedabstractWe present a data management platform in the cloud, CloudDB. The guiding principle of CloudDB’s design is establishing data independence for the applications that need to use diverse underlying data stores that are optimized for varying workload needs and characteristics. The applications should not have to be aware of the physical organization of the data and how the data is accessed. Ideally, an application only needs a logical specification of the data access layer and the data access requests are handled in a declarative way. CloudDB hosts variety of specialized databases that deliver high performance, scalability, and cost efficiency for varying application needs. CloudDB’s API layer is designed in such a way to give data independence to the higher level applications. The goal is to let the clients use just a simple, standard, and uniform language API to access data management functions as a service. Hakan Hacigümüs, Jun'ichi Tatemura, Wang-Pin Hsiung, Hyun Jin Moon, Oliver Po, Arsany Sawires, Yun Chi, Hojjat Jafarpour |
SERVICES | 2 |
| 2009 | Efficient overlap and content reuse detection in blogs and online news articlesabstractThe use of blogs to track and comment on real world (political, news, entertainment) events is growing. Similarly, as more individuals start relying on the Web as their primary information source and as more traditional media outlets try reaching consumers through alternative venues, the number of news sites on the Web is also continuously increasing. Content-reuse, whether in the form of extensive quotations or content borrowing across media outlets, is very common in blogs and news entries outlets tracking the same real-world event. Knowledge about which web entries re-use content from which others can be an effective asset when organizing these entries for presentation. On the other hand, this knowledge is not cheap to acquire: considering the size of the related space web entries, it is essential that the techniques developed for identifying re-use are fast and scalable. Furthermore, the dynamic nature of blog and news entries necessitates incremental processing for reuse detection. In this paper, we develop a novel qSign algorithm that efficiently and effectively analyze the blogosphere for quotation and reuse identification. Experiment results show that with qSign processing time gains from 10X to 100X are possible while maintaining reuse detection rates of upto 90%. Furthermore, processing time gains can be pushed multiple orders of magnitude (from 100X to 1000X) for 70% recall. Jong Wook Kim, K. Selçuk Candan, Jun'ichi Tatemura |
WWW | 3 |
| 2009 | Extracting data records from the web using tag path clusteringabstractFully automatic methods that extract lists of objects from the Web have been studied extensively. Record extraction, the first step of this object extraction process, identifies a set of Web page segments, each of which represents an individual object (e.g., a product). State-of-the-art methods suffice for simple search, but they often fail to handle more complicated or noisy Web page structures due to a key limitation -- their greedy manner of identifying a list of records through pairwise comparison (i.e., similarity match) of consecutive segments. This paper introduces a new method for record extraction that captures a list of objects in a more robust way based on a holistic analysis of a Web page. The method focuses on how a distinct tag path appears repeatedly in the DOM tree of the Web document. Instead of comparing a pair of individual segments, it compares a pair of tag path occurrence patterns (called visual signals) to estimate how likely these two tag paths represent the same list of objects. The paper introduces a similarity measure that captures how closely the visual signals appear and interleave. Clustering of tag paths is then performed based on this similarity measure, and sets of tag paths that form the structure of data records are extracted. Experiments show that this method achieves higher accuracy than previous methods. Gengxin Miao, Jun'ichi Tatemura, Wang-Pin Hsiung, Arsany Sawires, Louise E. Moser |
WWW | 2 |
| 2008 | Runtime Semantic Query Optimization for Event Stream ProcessingabstractDetecting complex patterns in event streams, i.e., complex event processing (CEP), has become increasingly important for modern enterprises to react quickly to critical situations. In many practical cases business events are generated based on pre-defined business logics. Hence constraints, such as occurrence and order constraints, often hold among events. Reasoning using these known constraints enables us to predict the non-occurrences of certain future events, thereby helping us to identify and then terminate the long running query processes that are guaranteed to not lead to successful matches. In this work, we focus on exploiting event constraints to optimize CEP over large volumes of business transaction streams. Since the optimization opportunities arise at runtime, we develop a runtime query unsatisfiability (RunSAT) checking technique that detects optimal points for terminating query evaluation. To assure efficiency of RunSAT checking, we propose mechanisms to precompute the query failure conditions to be checked at runtime. This guarantees a constant-time RunSAT reasoning cost, making our technique highly scalable. We realize our optimal query termination strategies by augmenting the query with Event-Condition-Action rules encoding the pre-computed failure conditions. This results in an event processing solution compatible with state-of-the-art CEP architectures. Extensive experimental results demonstrate that significant performance gains are achieved, while the optimization overhead is small. Luping Ding, Songting Chen, Elke A. Rundensteiner, Jun'ichi Tatemura, Wang-Pin Hsiung, K. Selçuk Candan |
ICDE | 4 |
| 2008 | Monitoring Moving Objects Using Low Frequency Snapshots in Sensor NetworksabstractMonitoring moving objects is one of the key application domains for sensor networks. In the absence of cooperative objects and devices attached to these objects, target tracking algorithms have to be used for monitoring. In this paper, we present that many of the applications of moving object monitoring systems could be addressed with low-frequency snapshot-based queries. With the realization of this query type, we show that existing target tracking algorithms may not be the least expensive solutions. We introduce an approach that uses two alternating strategies. We maintain a cheap low-quality knowledge of moving objects' location between snapshots and trigger expensive sensor readings only when a snapshot period has elapsed. With extensive experiments we show that our approach is significantly more energy efficient than established methods. It is also more effective than existing data-and-query centric in-network query processing schemes as it can maintain object identities between snapshots. Egemen Tanin, Songting Chen, Jun'ichi Tatemura, Wang-Pin Hsiung |
MDM | 3 |
| 2008 | Supporting OLAP operations over imperfectly integrated taxonomiesabstractOLAP is an important tool in decision support. With the help of domain knowledge, such as hierarchies of attribute values, OLAP helps the user observe the effects of various decisions. One assumption of most OLAP operations is that the available domain knowledge is precise. In particular, they assume that the hierarchy of values over which the user can navigate forms a taxonomy. In this paper, we first note that when multiple heterogeneous data sources are involved in the gathering of the data and the associated domain knowledge, the integrated knowledge-base, constructed by combining locally available taxonomies based on the concept matchings, may not be a taxonomy itself. Specifically, existence of intersections among concepts from different sources compromises the tree-structure of the integrated taxonomy and prevents effective use of hierarchical navigation techniques, such as drill-down and roll-up. To cope with this, we introduce concept un-classification, where a select few of the concepts are eliminated to ensure that the remaining structure is a navigable taxonomy, without concept intersections. Since un-classifying an originally classified data is not desirable, we consider ways to minimize un-classification in the process. We introduce a cost model which captures the imprecision caused by the un-classification process and we formulate the problem of finding an un-classification strategy which eliminates intersections and which adds minimal imprecision to the resulting structure. We show that, when performed naively, this task can be very costly and thus we propose a bottom-up preprocessing strategy which supports basic navigational analytics operations, such as drill-down and roll-up. Experiments over synthetic and real-life data verified the effectiveness and efficiency of our approach. Yan Qi 0002, K. Selçuk Candan, Jun'ichi Tatemura, Songting Chen, Fenglin Liao |
SIGMOD Conference | 3 |
| 2008 | UQBE: uncertain query by example for web service mashupabstractThe UQBE is a mashup tool for non-programmers that supports query-by-example (QBE) over a schema made up by the user without knowing the schema of the original sources. Based on automated schema matching with uncertainty, the UQBE system returns the best confident results. The system lets the user refine them interactively. A tuple in the query result is associated with lineage that is a boolean formula over schema matching decisions representing underlying conditions on which the corresponding tuple is included in the result. Given binary feedbacks on tuples by the user, which are possibly imprecise, the system solves it as an optimization problem to refine confidence values of matching decisions. The demo features graphical user interaction on the UQBE system, including querying and refinement. Jun'ichi Tatemura, Songting Chen, Fenglin Liao, Oliver Po, K. Selçuk Candan, Divyakant Agrawal |
SIGMOD Conference | 1 |
| 2008 | Scalable Filtering of Multiple Generalized-Tree-Pattern Queries over XML StreamsabstractAn XML publish/subscribe system needs to filter a large number of queries over XML streams. Most existing systems only consider filtering the simple XPath statements. In this paper, we focus on filtering of the more complex Generalized-Tree-Pattern (GTP) queries. Our filtering mechanism is based on a novel Tree-of-Path (TOP) encoding scheme, which compactly represents the path matches for the entire document. First, we show that the TOP encodings can be efficiently produced via a shared bottom-up path matching. Second, with the aid of this TOP encoding, we can 1) achieve polynomial time and space complexity for post processing, 2) avoid redundant predicate evaluations, 3) allow an efficient duplicate-free and merge join-based algorithm for merging multiple encoded path matches and 4) simplify the processing of GTP queries. Overall our approach maximizes the sharing opportunity across queries by exploiting the suffix as well as prefix sharing. At the same time, our TOP encodings allow efficient post processing for GTP queries. Extensive performance studies show that our GFilter solution not only achieves significantly better filtering performance than state-of-the-art algorithms, but also is capable of efficiently filtering the more complex GTP queries. Songting Chen, Hua-Gang Li, Jun'ichi Tatemura, Wang-Pin Hsiung, Divyakant Agrawal, K. Selçuk Candan |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2008 | Detecting splogs via temporal dynamics using self-similarity analysisabstractThis article addresses the problem of spam blog (splog) detection using temporal and structural regularity of content, post time and links. Splogs are undesirable blogs meant to attract search engine traffic, used solely for promoting affiliate sites. Blogs represent popular online media, and splogs not only degrade the quality of search engine results, but also waste network resources. The splog detection problem is made difficult due to the lack of stable content descriptors. We have developed a new technique for detecting splogs, based on the observation that a blog is a dynamic, growing sequence of entries (or posts) rather than a collection of individual pages. In our approach, splogs are recognized by their temporal characteristics and content. There are three key ideas in our splog detection framework. (a) We represent the blog temporal dynamics using self-similarity matrices defined on the histogram intersection similarity measure of the time, content, and link attributes of posts, to investigate the temporal changes of the post sequence. (b) We study the blog temporal characteristics using a visual representation derived from the self-similarity measures. The visual signature reveals correlation between attributes and posts, depending on the type of blogs (normal blogs and splogs). (c) We propose two types of novel temporal features to capture the splog temporal characteristics. In our splog detector, these novel features are combined with content based features. We extract a content based feature vector from blog home pages as well as from different parts of the blog. The dimensionality of the feature vector is reduced by Fisher linear discriminant analysis. We have tested an SVM-based splog detector using proposed features on real world datasets, with appreciable results (90% accuracy). Yu-Ru Lin, Hari Sundaram, Yun Chi, Jun'ichi Tatemura, Belle L. Tseng |
ACM Trans. Web | 4 |
| 2007 | Splog Detection using Content, Time and Link StructuresabstractThis paper focuses on spam blog (splog) detection. Blogs are highly popular, new media social communication mechanisms and splogs corrupt blog search results as well as waste network resources. In our approach we exploit unique blog temporal dynamics to detect splogs. The key idea is that splogs exhibit high temporal regularity in content and post time, as well as consistent linking patterns. Temporal content regularity is detected using a novel autocorrelation of post content. Temporal structural regularity is determined using the entropy of the post time difference distribution, while the link regularity is computed using a HITS based hub score measure. Experiments based on the annotated ground truth on real world dataset show excellent results on splog detection tasks with 90% accuracy. Yu-Ru Lin, Hari Sundaram, Yun Chi, Jun'ichi Tatemura, Belle L. Tseng |
ICME | 4 |
| 2007 | Structural and temporal analysis of the blogosphere through community factorizationabstractThe blogosphere has unique structural and temporal properties since blogs are typically used as communication media among human individuals. In this paper, we propose a novel technique that captures the structure and temporal dynamics of blog communities. In our framework, a community is a set of blogs that communicate with each other triggered by some events (such as a news article). The community is represented by its structure and temporal dynamics: a community graph indicates how often one blog communicates with another, and a community intensity indicates the activity level of the community that varies over time. Our method, community factorization, extracts such communities from the blogosphere, where the communication among blogs is observed as a set of subgraphs (i.e., threads of discussion). This community extraction is formulated as a factorization problem in the framework of constrained optimization, in which the objective is to best explain the observed interactions in the blogosphere over time. We further provide a scalable algorithm for computing solutions to the constrained optimization problems. Extensive experimental studies on both synthetic and real blog data demonstrate that our technique is able to discover meaningful communities that are not detectable by traditional methods. Yun Chi, Shenghuo Zhu, Xiaodan Song, Jun'ichi Tatemura, Belle L. Tseng |
KDD | 4 |
| 2007 | Application semantics in query optimization for WSNsabstractEfficient data acquisition in WSNs has attracted significant interest. For example, TinyDB [2] introduced query dissemination and data aggregation trees. Later, a probabilistic model of the physical world is used in [1]. Recently, [3] argues that probabilistic models of the physical world used in acquisition may miss outliers and introduces spatio-temporal suppression-based methods. We classify these established approaches as query-and-data centric approaches for optimizing the data acquisition process. Egemen Tanin, Songting Chen, Jun'ichi Tatemura, Wang-Pin Hsiung |
SenSys | 3 |
| 2007 | Mashup Feeds: : continuous queries over web servicesabstractMashup Feeds is a system that supports integrated web service feeds as continuous queries. We introduce collection-based stream processing semantics to enable information extraction by monitoring source evolution over time. Jun'ichi Tatemura, Arsany Sawires, Oliver Po, Songting Chen, K. Selçuk Candan, Divyakant Agrawal, Maria Goveas |
SIGMOD Conference | 1 |
| 2007 | Blog Community Discovery and Evolution Based on Mutual Awareness ExpansionabstractThere are information needs involving costly decisions that cannot be efficiently satisfied through conventional Web search engines. Alternately, community centric search can provide multiple viewpoints to facilitate decision making. We propose to discover and model the temporal dynamics of thematic communities based on mutual awareness, where the awareness arises due to observable blogger actions and the expansion of mutual awareness leads to community formation. Given a query, we construct a directed action graph that is time-dependent, and weighted with respect to the query. We model the process of mutual awareness expansion using a random walk process and extract communities based on the model. We propose an interaction space based representation to quantify community dynamics. Each community is represented as a vector in the interaction space and its evolution is determined by a novel interaction correlation method. We have conducted experiments with a real-world blog dataset and have promising results for detection as well as insightful results for community evolution. Yu-Ru Lin, Hari Sundaram, Yun Chi, Jun'ichi Tatemura, Belle L. Tseng |
Web Intelligence | 4 |
| 2006 | Identifying Agitators as Important Blogger Based on Analyzing Blog Threads
Shinsuke Nakajima, Jun'ichi Tatemura, Yoshinori Hara, Katsumi Tanaka, Shunsuke Uemura |
APWeb | 2 |
| 2006 | Eigen-trend: trend analysis in the blogosphere based on singular value decompositionsabstractThe blogosphere - the totality of blog-related Web sites - has become a great source of trend analysis in areas such as product survey, customer relationship, and marketing. Existing approaches are based on simple counts, such as the number of entries or the number of links. In this paper, we introduce a novel concept, coined eigen-trend, to represent the temporal trend in a group of blogs with common interests and propose two new techniques for extracting eigen-trends in blogs. First, we propose a trend analysis technique based on the singular value decomposition. Extracted eigen-trends provide new insights into multiple trends on the same keyword. Second, we propose another trend analysis technique based on a higher-order singular value decomposition. This analyzes the blogosphere as a dynamic graph structure and extracts eigen-trends that reflect the structural changes of the blogosphere over time. Experimental studies based on synthetic data sets and a real blog data set show that our new techniques can reveal a lot of interesting trend information and insights in the blogosphere that are not obtainable from traditional count-based methods. Yun Chi, Belle L. Tseng, Jun'ichi Tatemura |
CIKM | 3 |
| 2006 | AFilter: Adaptable XML Filtering with Prefix-Caching and Suffix-Clustering
K. Selçuk Candan, Wang-Pin Hsiung, Songting Chen, Jun'ichi Tatemura, Divyakant Agrawal |
VLDB | 4 |
| 2006 | Twig2Stack: Bottom-up Processing of Generalized-Tree-Pattern Queries over XML Documents
Songting Chen, Hua-Gang Li, Jun'ichi Tatemura, Wang-Pin Hsiung, Divyakant Agrawal, K. Selçuk Candan |
VLDB | 3 |
| 2006 | Safety Guarantee of Continuous Join Queries over Punctuated Data Streams
Hua-Gang Li, Songting Chen, Jun'ichi Tatemura, Divyakant Agrawal, K. Selçuk Candan, Wang-Pin Hsiung |
VLDB | 3 |
| 2006 | Maintaining XPath Views In Loosely Coupled Systems
Arsany Sawires, Jun'ichi Tatemura, Oliver Po, Divyakant Agrawal, Amr El Abbadi, K. Selçuk Candan |
VLDB | 2 |
| 2005 | DHT overlay schemes for scalable p-range resource discoveryabstractThe information service is a critical component of a grid infrastructure for resource discovery. Although P2P computing paradigm could address some of the scalability issues that plagued grid resource discovery, most existing distributed hash table (DHT) based P2P overlays have difficulty in treating attribute range queries that are common in resource discovery lookups due to the inherent randomness of hash functions. Recently, there have been various attempts to solve the range search problem over DHT networks [Aberer, et al. (2003), Gao and Steenkiste (2004), Ratnasamy, et al. (2003), and Tanin, et al. (2005)]. Central to all of these is a mapping scheme which maps the tree-structured logical index space to some DHT-based physical node space. In this paper, we propose a general framework to put all these under the same umbrella based on how mapping of the tree-structured index that identifies a physical node responsible of a particular range is done through replication. We identify three schemes which should cover the spectrum of all meaningful replication schemes: the tree replication scheme (TRS) replicates the entire tree; the path caching scheme (PCS) replicates paths from root to leaves; and the node replication scheme (NRS) replicates individual logical nodes. K. Selçuk Candan, Jun'ichi Tatemura, Divyakant Agrawal, Dirceu Cavendish |
HPDC | 3 |
| 2005 | Dynamic Stochastic Models for Workflow Response OptimizationabstractIn this paper we propose a solution for optimizing (Web service) business workflow response times through dynamic resource allocation. On-the-fly monitoring is combined with a novel workflow modeling algorithm that discovers critical execution paths and builds "dynamic" stochastic models in the associated "critical graph". One novel contribution of this work is the ability to naturally handle parallel workflow execution paths. This is essential in applications where workflows include multiple concurrent service calls/paths that need to be "joined" at a later point in time. We discuss the automatic deployment of on-the-fly monitoring mechanisms within the resource management mechanisms. We implement, deploy and experiment with a proof of concept within a generalized Web services business process (BPEL4WS/SOAP) framework. In the experimental setup we explore and show the natural adaptation to changing workflow conditions and appropriate automatic re-allocation of resources to reduce execution times. Radu Sion, Jun'ichi Tatemura |
ICWS | 2 |
| 2005 | WreC: A Scalable Middleware Architecture to Enable XML Caching for Web-Services
Jun'ichi Tatemura, Oliver Po, Arsany Sawires, Divyakant Agrawal, K. Selçuk Candan |
Middleware | 1 |
| 2005 | On Overlay Schemes to Support Point-in-Range Queries for Scalable Grid Resource DiscoveryabstractA resource directory is a critical component of a grid architecture. P2P computing paradigm could address some of the scalability issues that make distributed resource discovery services challenging. Unfortunately, most existing distributed hash table (DHT) based P2P overlays have difficulty in treating attribute range queries that are common in resource discovery lookups. This paper proposes a general framework for range-based resource discovery. In particular, the proposed framework maps tree-structured logical data (i.e., range indexing) onto a DHT-based physical node space (i.e., resource brokers). In this paper, we consider three mapping schemes from the logical space onto the physical space. Each mapping scheme uses a different replication mechanism to reduce range search time and to achieve load balance. We analytically and experimentally compare the performance characteristics (query/update costs and workload distributions) of these schemes and discuss their applicability under different resource discovery service scenarios. K. Selçuk Candan, Jun'ichi Tatemura, Divyakant Agrawal, Dirceu Cavendish |
Peer-to-Peer Computing | 3 |
| 2005 | Incremental Maintenance of Path Expression ViewsabstractCaching data by maintaining materialized views typically requires updating the cache appropriately to reflect dynamic source updates. Extensive research has addressed the problem of incremental view maintenance for relational data but only few works have addressed it for semi-structured data. In this paper we address the problem of incremental maintenance of views defined over XML documents using path-expressions. The approach described in this paper has the following main features that distinguish it from the previous works: (1) The view specification language is powerful and standardized enough to be used in realistic applications. (2) The size of the auxiliary data maintained with the views depends on the expression size and the answer size regardless of the source data size.(3) No source schema is assumed to exist; the source data can be any general well-formed XML document. Experimental evaluation is conducted to assess the performance benefits of the proposed approach. Arsany Sawires, Jun'ichi Tatemura, Oliver Po, Divyakant Agrawal, K. Selçuk Candan |
SIGMOD Conference | 2 |
| 2000 | Dynamic Label Sampling on Fisheye Maps for Information ExplorationabstractFor data with large dimensionality, placing labels is critical for users' comprehension of a scatterplot or a map of items. We propose a dynamic label sampling technique that, combined with graphical fisheye views, selects appropriate labels out of a large set of items on a map. Labels are sampled to give focus and contextual information according to users' panning/zooming and filtering operation. The paper also demonstrates an example of visual exploration with the image browser based on our technique. Jun'ichi Tatemura |
Advanced Visual Interfaces | 1 |
| 2000 | Virtual reviewers for collaborative exploration of movie reviewsabstractWe propose a collaborative exploration system that helps users to explore recommendations from various viewpoints. Given ratings and reviews on movies from reviewers, the system provides “virtual reviewers” that represent particular viewpoints. Each virtual reviewer navigates the user by recommending and characterizing both movies and reviewers according to its viewpoint. We have developed a browsing method with virtual reviewers and visual interfaces. Jun'ichi Tatemura |
IUI | 1 |
| 1999 | Visual Querying and Explanation of Recommendations from Collaborative Filtering SystemsabstractNo abstract available. Jun'ichi Tatemura |
IUI | 1 |
| 1997 | Four Promising Multimedia Databases and Their Embodiments
Yoshitomo Yaginuma, Tomoyuki Yatabe, Takashi Satou, Jun'ichi Tatemura, Masao Sakauchi |
Multim. Tools Appl. | 4 |