EDBT 2026 Demo / reviewers in the wild / expert
Wang-Pin Hsiung
dblp:37/236
· DBLP profile ↗
30ranked-venue papers
0as first author
4since 2021 · last 2023
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 24Artificial intelligence and machine learning · 5 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
15 papers |
Transaction processing and concurrency control · 20% Distributed and cloud data management · 17% Data stream processing · 15% | |
| Computer architecture, parallel and distributed computing, and storage systems
6 papers |
Cloud and datacenter computing · 67% Distributed systems · 22% Memory systems · 10% | |
| Computer networks
5 papers |
Content delivery and video streaming · 58% Internet of things and sensor networks · 36% Edge and fog computing · 6% | |
| Artificial intelligence
1 paper |
Information extraction and text analysis · 100% |
Topics — the 29 heaviest of 35, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Indexing and storage engines
key-value store |
0.2 | 2 | 2014 | Partiqle: an elastic SQL engine over key-value stores · SIGMOD Conference 2012 SAGE: A logical and physical design tool for entity-group based new SQL systems · ICDE 2014 |
Query processing and optimization
query scheduling |
0.2 | 1 | 2013 | Distribution-Based Query Scheduling · Proc. VLDB Endow. 2013 |
Database theory › conjunctive query
tree pattern query |
0.1 | 2 | 2008 | Scalable Filtering of Multiple Generalized-Tree-Pattern Queries over XML Streams · IEEE Trans. Knowl. Data Eng. 2008 Twig2Stack: Bottom-up Processing of Generalized-Tree-Pattern Queries over XML Documents · VLDB 2006 |
Transaction processing and concurrency control › OLTP
OLTP engine |
0.1 | 1 | 2012 | Partiqle: an elastic SQL engine over key-value stores · SIGMOD Conference 2012 |
Natural language and speech › Information extraction and text analysis
web information extraction |
0.1 | 1 | 2009 | Extracting data records from the web using tag path clustering · WWW 2009 |
Data stream processing
complex event processing |
0.1 | 1 | 2008 | Runtime Semantic Query Optimization for Event Stream Processing · ICDE 2008 |
Query processing and optimization › adaptive query processing
dynamic query optimization |
0.1 | 1 | 2008 | Runtime Semantic Query Optimization for Event Stream Processing · ICDE 2008 |
Data stream processing
XML stream processing |
0.1 | 1 | 2008 | Scalable Filtering of Multiple Generalized-Tree-Pattern Queries over XML Streams · IEEE Trans. Knowl. Data Eng. 2008 |
Cloud and datacenter computing › datacenter services › online service systems › internet services
dynamic content caching |
0.1 | 2 | 2003 | Engineering and hosting adaptive freshness-sensitive web applications on data centers · WWW 2003 View Invalidation for Dynamic Content Caching in Multitiered Architectures · VLDB 2002 |
Internet of things and sensor networks
wireless sensor network |
0.1 | 1 | 2007 | Application semantics in query optimization for WSNs · SenSys 2007 |
Database theory › datalog evaluation
bottom-up evaluation |
0.1 | 1 | 2006 | Twig2Stack: Bottom-up Processing of Generalized-Tree-Pattern Queries over XML Documents · VLDB 2006 |
Data stream processing › XML stream processing
XML filtering |
0.1 | 1 | 2006 | AFilter: Adaptable XML Filtering with Prefix-Caching and Suffix-Clustering · VLDB 2006 |
Query processing and optimization
XML query processing |
0.1 | 1 | 2006 | Twig2Stack: Bottom-up Processing of Generalized-Tree-Pattern Queries over XML Documents · VLDB 2006 |
Cloud and datacenter computing
database-as-a-service |
0.0 | 1 | 2013 | Distribution-Based Query Scheduling · Proc. VLDB Endow. 2013 |
Cloud and datacenter computing › job scheduling
query scheduling |
0.0 | 1 | 2013 | Distribution-Based Query Scheduling · Proc. VLDB Endow. 2013 |
Content delivery and video streaming › caching
dynamic content caching |
0.0 | 1 | 2004 | Challenges and practices in deploying web acceleration solutions for distributed enterprise systems · WWW 2004 |
Content delivery and video streaming › web content delivery
web acceleration |
0.0 | 1 | 2004 | Challenges and practices in deploying web acceleration solutions for distributed enterprise systems · WWW 2004 |
Distributed systems › replication
database replication |
0.0 | 1 | 2003 | Engineering and hosting adaptive freshness-sensitive web applications on data centers · WWW 2003 |
Indexing and storage engines › caching
database caching |
0.0 | 1 | 2002 | Issues and Evaluations of Caching Solutions for Web Application Acceleration · VLDB 2002 |
Memory systems
cache |
0.0 | 1 | 2002 | Issues and Evaluations of Caching Solutions for Web Application Acceleration · VLDB 2002 |
Distributed systems
query result caching |
0.0 | 1 | 2002 | Issues and Evaluations of Caching Solutions for Web Application Acceleration · VLDB 2002 |
Database system architecture and tuning
cache invalidation |
0.0 | 1 | 2001 | Enabling Dynamic Content Caching for Database-Driven Web Sites · SIGMOD Conference 2001 |
Distributed and cloud data management › web caching
dynamic content caching |
0.0 | 1 | 2001 | Enabling Dynamic Content Caching for Database-Driven Web Sites · SIGMOD Conference 2001 |
Web and social media mining › web mining
web page structure analysis |
0.0 | 1 | 2009 | Extracting data records from the web using tag path clustering · WWW 2009 |
Data stream processing
publish/subscribe |
0.0 | 1 | 2008 | Scalable Filtering of Multiple Generalized-Tree-Pattern Queries over XML Streams · IEEE Trans. Knowl. Data Eng. 2008 |
Content delivery and video streaming › caching
web caching |
0.0 | 2 | 2002 | View Invalidation for Dynamic Content Caching in Multitiered Architectures · VLDB 2002 Enabling Dynamic Content Caching for Database-Driven Web Sites · SIGMOD Conference 2001 |
Data stream processing
punctuated data streams |
0.0 | 1 | 2006 | Safety Guarantee of Continuous Join Queries over Punctuated Data Streams · VLDB 2006 |
Edge and fog computing
edge caching |
0.0 | 1 | 2003 | Engineering and hosting adaptive freshness-sensitive web applications on data centers · WWW 2003 |
Indexing and storage engines
caching |
0.0 | 1 | 2001 | Cache Portal: Technology for Accelerating Database-driven e-commerce Web Sites · VLDB 2001 |
Methods — techniques the papers use, named apart from their topics
histogram-based estimation · 0.3visual signal similarity · 0.2tag path clustering · 0.2workload-driven design · 0.2iterative user feedback · 0.2formal methods · 0.2microsharding · 0.1case study · 0.1satisfiability reasoning · 0.1merge join · 0.1event-condition-action rules · 0.1adaptive caching policy · 0.1caching · 0.1spatio-temporal suppression · 0.1probabilistic model · 0.1query result caching · 0.0cache invalidation · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | AnB: Application-in-a-Box to Rapidly Deploy and Self-optimize 5G AppsabstractWe present "Application in a Box" (AnB) product concept aimed at simplifying the deployment and operation of remote 5G applications. AnB comes pre-configured with all necessary hardware and software components, including sensors like cameras, hardware and software components for a local 5G wireless network, and 5G-ready apps. Enterprises can easily download additional apps from an App Store. Setting up a 5G infrastructure and running applications on it is a significant challenge, but AnB is designed to make it fast, convenient, and easy, even for those without extensive knowledge of software, computers, wireless networks, or AI-based analytics. With AnB, customers only need to open the box, set up the sensors, turn on the 5G networking and edge computing devices, and start running their applications. Our system software automatically deploys and optimizes the pipeline of microservices in the application on a tiered computing infrastructure that includes device, edge, and cloud computing. Application scalability, dynamic resource management, placement of critical tasks for low-latency response, and dynamic network bandwidth allocation for efficient 5G network usage are all automatically orchestrated.AnB offers cost savings, simplified setup and management, and increased reliability and security. We’ve implemented several real-world applications, such as collision prediction at busy traffic light intersections and remote construction site monitoring using video analytics. With AnB, deployment and optimization effort can be reduced from several months to just a few minutes. This is the first-of-its-kind approach to easing deployment effort and automating self-optimization of the application during system operation. Kunal Rao, Murugan Sankaradass, Giuseppe Coviello, Ciro Giuseppe De Vita, Gennaro Mellone, Wang-Pin Hsiung, Srimat T. Chakradhar |
SMARTCOMP | 6 |
| 2022 | ROMA: Resource Orchestration for Microservices-based 5G ApplicationsabstractWith the growth of 5G, Internet of Things (IoT), edge computing and cloud computing technologies, the infrastructure (compute and network) available to emerging applications (AR/VR, autonomous driving, industry 4.0, etc.) has become quite complex. There are multiple tiers of computing (IoT devices, near edge, far edge, cloud, etc.) that are connected with different types of networking technologies (LAN, LTE, 5G, MAN, WAN, etc.). Deployment and management of applications in such an environment is quite challenging. In this paper, we propose ROMA, which performs resource orchestration for microservices-based 5G applications in a dynamic, heterogeneous, multi-tiered compute and network fabric. We assume that only application-level requirements are known, and the detailed requirements of the individual microservices in the application are not specified. As part of our solution, ROMA identifies and leverages the coupling relationship between compute and network usage for various microservices and solves an optimization problem in order to appropriately identify how each microservice should be deployed in the complex, multi-tiered compute and network fabric, so that the end-to-end application requirements are optimally met. We implemented two real-world 5G applications in video surveillance and intelligent transportation system (ITS) domains. Through extensive experiments, we show that ROMA is able to save up to 90%, 55% and 44% compute and up to 80%, 95% and 75% network bandwidth for the surveillance (watchlist) and transportation application (person and car detection), respectively. This improvement is achieved while honoring the application performance requirements, and it is over an alternative scheme that employs a static and overprovisioned resource allocation strategy by ignoring the resource coupling relationships. Anousheh Gholami, Kunal Rao, Wang-Pin Hsiung, Oliver Po, Murugan Sankaradass, Srimat T. Chakradhar |
NOMS | 3 |
| 2021 | ECO: Edge-Cloud Optimization of 5G applicationsabstractCentralized cloud computing with 100+ milliseconds network latencies cannot meet the tens of milliseconds to sub-millisecond response times required for emerging 5G applications like autonomous driving, smart manufacturing, tactile internet, and augmented or virtual reality. We describe a new, dynamic runtime that enables such applications to make effective use of a 5G network, computing at the edge of this network, and resources in the centralized cloud, at all times. Our runtime continuously monitors the interaction among the microservices, estimates the data produced and exchanged among the microservices, and uses a novel graph min-cut algorithm to dynamically map the microservices to the edge or the cloud to satisfy application-specific response times. Our runtime also handles temporary network partitions, and maintains data consistency across the distributed fabric by using microservice proxies to reduce WAN bandwidth by an order of magnitude, all in an application-specific manner by leveraging knowledge about the application's functions, latency-critical pipelines and intermediate data. We illustrate the use of our runtime by successfully mapping two complex, representative real-world video analytics applications to the AWS/Verizon Wavelength edge-cloud architecture, and improving application response times by 2x when compared with a static edge-cloud implementation. Kunal Rao, Giuseppe Coviello, Wang-Pin Hsiung, Srimat T. Chakradhar |
CCGRID | 3 |
| 2021 | F3S: Free Flow Fever ScreeningabstractIdentification of people with elevated body temperature can reduce or dramatically slow down the spread of infectious diseases like COVID-19. We present a novel fever-screening system, F3S, that uses edge machine learning techniques to accurately measure core body temperatures of multiple individuals in a free-flow setting. F3S performs real-time sensor fusion of visual camera with thermal camera data streams to detect elevated body temperature, and it has several unique features: (a) visual and thermal streams represent very different modalities, and we dynamically associate semantically-equivalent regions across visual and thermal frames by using a new, dynamic alignment technique that analyzes content and context in real-time, (b) we track people through occlusions, identify the eye (inner canthus), forehead, face and head regions where possible, and provide an accurate temperature reading by using a prioritized refinement algorithm, and (c) we robustly detect elevated body temperature even in the presence of personal protective equipment like masks, or sunglasses or hats, all of which can be affected by hot weather and lead to spurious temperature readings. F3S has been deployed at over a dozen large commercial establishments, providing contact-less, free-flow, real-time fever screening for thousands of employees and customers in indoors and outdoor settings. Kunal Rao, Giuseppe Coviello, Min Feng 0001, Biplob Debnath, Wang-Pin Hsiung, Murugan Sankaradass, Yi Yang 0018, Oliver Po, Utsav Drolia, Srimat T. Chakradhar |
SMARTCOMP | 5 |
| 2014 | SAGE: A logical and physical design tool for entity-group based new SQL systemsabstractEntity-group based new SQL systems achieve scalability and consistency at the same time by using a key-value store as the storage layer and limiting each transaction's boundary to a collection of data (called an entity-group). Examples of such systems are Google's Megastore, NEC's Partiqle, and LinkedIn's Espresso. Application developers of such systems face tremendous challenges, both in designing entity-groups (and hence transaction boundaries) and physical layout of data in key-value stores. Entity-group designs directly impact consistency semantics of the workload, and physical layout impacts the application throughput. Both problems are challenging for users to solve manually. In this demonstration, we show a system that solves both problems with a user-friendly GUI built on our principled formal methods. We demonstrate the system using a simple Auction benchmark, and a more complex TPC-W benchmark. Wang-Pin Hsiung, Jun'ichi Tatemura, Hakan Hacigümüs |
ICDE | 2 |
| 2014 | Automatic entity-grouping for OLTP workloadsabstractSupporting an online transaction processing (OLTP) workload in a scalable and elastic fashion is a challenging task. Recently, a new breed of scalable systems have shown significant throughput gains by limiting consistency to small units of data called “entity-groups” (e.g., a user's account information stored together with all her emails in an online email service.) Transactions that access the data from only one entity-group are guaranteed of full ACID, but those that access multiple entity-groups are not. Defining entity-groups has direct impact on workload consistency and performance, and doing so for data with a complex schema is very challenging. It is prone to go to extremes - groups that are too fine-grained cause excessive number of expensive distributed transactions while those that are too coarse lead to excessive serialization and performance degradation. It is also difficult to balance conflicting requirements from different transactions. In commercially available entity-group systems, creating entity-groups is usually a manual process, which severely limits the usability of those systems. This paper is the first systematic effort on automating the entity-group design process. Our goal is to build a user-friendly design tool for automatically creating entity-groups based on a given workload and to help users trade consistency for performance in a principled manner. For advanced users, we allow them to provide feedback to the entity-group design and iteratively improve the final output. We demonstrate the effectiveness of our approach with widely used benchmarks. We also present the user experience of a prototype we built. Jun'ichi Tatemura, Oliver Po, Wang-Pin Hsiung, Hakan Hacigümüs |
ICDE | 4 |
| 2013 | PMAX: tenant placement in multitenant databases for profit maximizationabstractThere has been a great interest in exploiting the cloud as a platform for database as a service. As with other cloud-based services, database services may enjoy cost efficiency through consolidation: hosting multiple databases within a single physical server. Aggressive consolidation, however, may hurt the service quality, leading to SLA violation penalty, which in turn reduces the total business profit, called SLA profit. In this paper, we consider the problem of tenant placement in the cloud for SLA profit maximization, which, as will be shown in the paper, is strongly NP-hard. We propose SLA profit-aware solutions for database tenant placement based on our model for expected penalty computation for multitenant servers. Specifically, we present two approximation algorithms, which have constant approximation ratios, and we further discuss improving the quality of tenant placement using a dynamic programming algorithm. Extensive experiments based on TPC-W workload verified the performance of the proposed approaches. Ziyang Liu 0001, Hakan Hacigümüs, Hyun Jin Moon, Yun Chi, Wang-Pin Hsiung |
EDBT | 5 |
| 2013 | SWAT: a lightweight load balancing method for multitenant databasesabstractMultitenant databases achieve cost efficiency through the consolidation of multiple small tenants. However, performance isolation is an inherent problem in multitenant databases due to resource sharing among the tenants. That is, a bursty workload from a co-located tenant, i.e., a noisy neighbor, may affect the performance of the other tenants sharing the same system resources. We address this issue by using a load balancing method that is based on database replica swap. Unlike the traditional data migration-based load balancing, replica swap-based load balancing does not incur data movement, which makes it highly resource- and time-efficient. We propose a novel method of choosing which tenants should be subject to swaps. Our experimental results show that swap-based load balancing effectively reduces the number of SLA violations, which is the main performance metric we choose. Hyun Jin Moon, Hakan Hacigümüs, Yun Chi, Wang-Pin Hsiung |
EDBT | 4 |
| 2013 | Distribution-Based Query SchedulingabstractQuery scheduling, a fundamental problem in database management systems, has recently received a renewed attention, perhaps in part due to the rise of the "database as a service" (DaaS) model for database deployment. While there has been a great deal of work investigating different scheduling algorithms, there has been comparatively little work investigating what the scheduling algorithms can or should know about the queries to be scheduled. In this work, we investigate the efficacy of using histograms describing the distribution of likely query execution times as input to the query scheduler. We propose a novel distribution-based scheduling algorithm, Shepherd, and show that Shepherd substantially outperforms state-of-the-art point-based methods through extensive experimentation with both synthetic and TPC workloads. Yun Chi, Hakan Hacigümüs, Wang-Pin Hsiung, Jeffrey F. Naughton |
Proc. VLDB Endow. | 3 |
| 2012 | Partiqle: an elastic SQL engine over key-value storesabstractThe demo features Partiqle, a SQL engine over key-value stores as a relational alternative for the recent procedural approaches to support OLTP workloads elastically. Based on our microsharding framework [12], it employs a declarative specification, called transaction classes, of constraints applied on the transactions in a workload. We demonstrate use of a transaction class in design and analysis of OLTP workloads. We then demonstrate live-scaling of our fully functioning system on a server cluster. Jun'ichi Tatemura, Oliver Po, Wang-Pin Hsiung, Hakan Hacigümüs |
SIGMOD Conference | 3 |
| 2010 | CloudDB: One Size Fits All RevivedabstractWe present a data management platform in the cloud, CloudDB. The guiding principle of CloudDB’s design is establishing data independence for the applications that need to use diverse underlying data stores that are optimized for varying workload needs and characteristics. The applications should not have to be aware of the physical organization of the data and how the data is accessed. Ideally, an application only needs a logical specification of the data access layer and the data access requests are handled in a declarative way. CloudDB hosts variety of specialized databases that deliver high performance, scalability, and cost efficiency for varying application needs. CloudDB’s API layer is designed in such a way to give data independence to the higher level applications. The goal is to let the clients use just a simple, standard, and uniform language API to access data management functions as a service. Hakan Hacigümüs, Jun'ichi Tatemura, Wang-Pin Hsiung, Hyun Jin Moon, Oliver Po, Arsany Sawires, Yun Chi, Hojjat Jafarpour |
SERVICES | 3 |
| 2009 | Extracting data records from the web using tag path clusteringabstractFully automatic methods that extract lists of objects from the Web have been studied extensively. Record extraction, the first step of this object extraction process, identifies a set of Web page segments, each of which represents an individual object (e.g., a product). State-of-the-art methods suffice for simple search, but they often fail to handle more complicated or noisy Web page structures due to a key limitation -- their greedy manner of identifying a list of records through pairwise comparison (i.e., similarity match) of consecutive segments. This paper introduces a new method for record extraction that captures a list of objects in a more robust way based on a holistic analysis of a Web page. The method focuses on how a distinct tag path appears repeatedly in the DOM tree of the Web document. Instead of comparing a pair of individual segments, it compares a pair of tag path occurrence patterns (called visual signals) to estimate how likely these two tag paths represent the same list of objects. The paper introduces a similarity measure that captures how closely the visual signals appear and interleave. Clustering of tag paths is then performed based on this similarity measure, and sets of tag paths that form the structure of data records are extracted. Experiments show that this method achieves higher accuracy than previous methods. Gengxin Miao, Jun'ichi Tatemura, Wang-Pin Hsiung, Arsany Sawires, Louise E. Moser |
WWW | 3 |
| 2008 | Runtime Semantic Query Optimization for Event Stream ProcessingabstractDetecting complex patterns in event streams, i.e., complex event processing (CEP), has become increasingly important for modern enterprises to react quickly to critical situations. In many practical cases business events are generated based on pre-defined business logics. Hence constraints, such as occurrence and order constraints, often hold among events. Reasoning using these known constraints enables us to predict the non-occurrences of certain future events, thereby helping us to identify and then terminate the long running query processes that are guaranteed to not lead to successful matches. In this work, we focus on exploiting event constraints to optimize CEP over large volumes of business transaction streams. Since the optimization opportunities arise at runtime, we develop a runtime query unsatisfiability (RunSAT) checking technique that detects optimal points for terminating query evaluation. To assure efficiency of RunSAT checking, we propose mechanisms to precompute the query failure conditions to be checked at runtime. This guarantees a constant-time RunSAT reasoning cost, making our technique highly scalable. We realize our optimal query termination strategies by augmenting the query with Event-Condition-Action rules encoding the pre-computed failure conditions. This results in an event processing solution compatible with state-of-the-art CEP architectures. Extensive experimental results demonstrate that significant performance gains are achieved, while the optimization overhead is small. Luping Ding, Songting Chen, Elke A. Rundensteiner, Jun'ichi Tatemura, Wang-Pin Hsiung, K. Selçuk Candan |
ICDE | 5 |
| 2008 | Monitoring Moving Objects Using Low Frequency Snapshots in Sensor NetworksabstractMonitoring moving objects is one of the key application domains for sensor networks. In the absence of cooperative objects and devices attached to these objects, target tracking algorithms have to be used for monitoring. In this paper, we present that many of the applications of moving object monitoring systems could be addressed with low-frequency snapshot-based queries. With the realization of this query type, we show that existing target tracking algorithms may not be the least expensive solutions. We introduce an approach that uses two alternating strategies. We maintain a cheap low-quality knowledge of moving objects' location between snapshots and trigger expensive sensor readings only when a snapshot period has elapsed. With extensive experiments we show that our approach is significantly more energy efficient than established methods. It is also more effective than existing data-and-query centric in-network query processing schemes as it can maintain object identities between snapshots. Egemen Tanin, Songting Chen, Jun'ichi Tatemura, Wang-Pin Hsiung |
MDM | 4 |
| 2008 | Scalable Filtering of Multiple Generalized-Tree-Pattern Queries over XML StreamsabstractAn XML publish/subscribe system needs to filter a large number of queries over XML streams. Most existing systems only consider filtering the simple XPath statements. In this paper, we focus on filtering of the more complex Generalized-Tree-Pattern (GTP) queries. Our filtering mechanism is based on a novel Tree-of-Path (TOP) encoding scheme, which compactly represents the path matches for the entire document. First, we show that the TOP encodings can be efficiently produced via a shared bottom-up path matching. Second, with the aid of this TOP encoding, we can 1) achieve polynomial time and space complexity for post processing, 2) avoid redundant predicate evaluations, 3) allow an efficient duplicate-free and merge join-based algorithm for merging multiple encoded path matches and 4) simplify the processing of GTP queries. Overall our approach maximizes the sharing opportunity across queries by exploiting the suffix as well as prefix sharing. At the same time, our TOP encodings allow efficient post processing for GTP queries. Extensive performance studies show that our GFilter solution not only achieves significantly better filtering performance than state-of-the-art algorithms, but also is capable of efficiently filtering the more complex GTP queries. Songting Chen, Hua-Gang Li, Jun'ichi Tatemura, Wang-Pin Hsiung, Divyakant Agrawal, K. Selçuk Candan |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2007 | Application semantics in query optimization for WSNsabstractEfficient data acquisition in WSNs has attracted significant interest. For example, TinyDB [2] introduced query dissemination and data aggregation trees. Later, a probabilistic model of the physical world is used in [1]. Recently, [3] argues that probabilistic models of the physical world used in acquisition may miss outliers and introduces spatio-temporal suppression-based methods. We classify these established approaches as query-and-data centric approaches for optimizing the data acquisition process. Egemen Tanin, Songting Chen, Jun'ichi Tatemura, Wang-Pin Hsiung |
SenSys | 4 |
| 2006 | AFilter: Adaptable XML Filtering with Prefix-Caching and Suffix-Clustering
K. Selçuk Candan, Wang-Pin Hsiung, Songting Chen, Jun'ichi Tatemura, Divyakant Agrawal |
VLDB | 2 |
| 2006 | Twig2Stack: Bottom-up Processing of Generalized-Tree-Pattern Queries over XML Documents
Songting Chen, Hua-Gang Li, Jun'ichi Tatemura, Wang-Pin Hsiung, Divyakant Agrawal, K. Selçuk Candan |
VLDB | 4 |
| 2006 | Safety Guarantee of Continuous Join Queries over Punctuated Data Streams
Hua-Gang Li, Songting Chen, Jun'ichi Tatemura, Divyakant Agrawal, K. Selçuk Candan, Wang-Pin Hsiung |
VLDB | 6 |
| 2004 | Challenges and practices in deploying web acceleration solutions for distributed enterprise systemsabstractFor most Web-based applications, contents are created dynamically based on the current state of a business, such as product prices and inventory, stored in database systems. These applications demand personalized content and track user behavior while maintaining application integrity. Many of such practices are not compatible with Web acceleration solutions. Consequently, although many web acceleration solutions have shown promising performance improvement and scalability, architecting and engineering distributed enterprise Web applications to utilize available content delivery networks remains a challenge. In this paper, we examine the challenge to accelerate J2EE-based enterprise web applications. We list obstacles and recommend some practices to transform typical database-driven J2EE applications to cache friendly Web applications where Web acceleration solutions can be applied. Furthermore, such transformation should be done without modification to the underlying application business logic and without sacrificing functions that are essential to e-commerce. We take the J2EE reference software, the Java PetStore, as a case study. By using the proposed guideline, we are able to cache more than 90% of the content in the PetStore and scale up the Web site more than 20 times. Wen-Syan Li, Wang-Pin Hsiung, Oliver Po, Koji Hino, K. Selçuk Candan, Divyakant Agrawal |
WWW | 2 |
| 2003 | Freshness-driven Adaptive Caching for Dynamic ContentabstractWith the wide availability of content delivery networks, many e-commerce Web applications utilize edge cache servers to cache and deliver dynamic contents at locations much closer to users, avoiding network latency. By caching a large number of dynamic content pages in the edge cache servers, response time can be reduced, benefiting from higher cache hit rates. However this is achieved at the expense of higher invalidation cost. On the other hand, a higher invalidation cost leads to a longer invalidation cycle (time to perform the invalidation check on the pages in caches) at the expense of freshness of cached dynamic content. In this paper we propose a freshness-driven adaptive dynamic content caching technique, which monitors response time and invalidation cycle length and dynamically adjusts caching policies. We have implemented the proposed technique within NECs CachePortal Web acceleration solution. The experimental results show that the proposed technique consistently maintains the best content freshness to users. The experimental results also show that even a Web site with dynamic content caching enabled can further benefit from deployment of our solution with improvement of its content freshness up to 10 times especially during heavy traffic. Wen-Syan Li, Oliver Po, Wang-Pin Hsiung, K. Selçuk Candan, Divyakant Agrawal |
DASFAA | 3 |
| 2003 | CachePortal II: Acceleration of Very Large Scale Data Center-Hosted Database-driven Web Applications
Wen-Syan Li, Oliver Po, Wang-Pin Hsiung, K. Selçuk Candan, Divyakant Agrawal, Yusuf Akca, Kunihiro Taniguchi |
VLDB | 3 |
| 2003 | Engineering and hosting adaptive freshness-sensitive web applications on data centersabstractWide-area database replication technologies and the availability of content delivery networks allow Web applications to be hosted and served from powerful data centers. This form of application support requires a complete Web application suite to be distributed along with the database replicas. A major advantage of this approach is that dynamic content is served from locations closer to users, leading into reduced network latency and fast response times. However, this is achieved at the expense of overheads due to (a) invalidation of cached dynamic content in the edge caches and (b) synchronization of database replicas in the data center. These have adverse effects on the freshness of delivered content. In this paper, we propose a freshness-driven adaptive dynamic content caching, which monitors the system status and adjusts caching policies to provide content freshness guarantees. The proposed technique has been intensively evaluated to validate its effectiveness. The experimental results show that the freshness-driven adaptive dynamic content caching technique consistently provides good content freshness. Furthermore, even a Web site that enables dynamic content caching can further benefit from our solution, which improves content freshness up to 7 times, especially under heavy user request traffic and long network latency conditions. Our approach also provides better scalability and significantly reduced response times up to 70% in the experiments. Wen-Syan Li, Oliver Po, Wang-Pin Hsiung, K. Selçuk Candan, Divyakant Agrawal |
WWW | 3 |
| 2003 | Freshness-driven adaptive caching for dynamic content Web sites
Wen-Syan Li, Oliver Po, Wang-Pin Hsiung, K. Selçuk Candan, Divyakant Agrawal |
Data Knowl. Eng. | 3 |
| 2003 | Corrigendum to: "Freshness-driven adaptive caching for dynamic content web sites" [Data & Knowledge Engineering 47 (2) (2003) 269-296]
Wen-Syan Li, Oliver Po, Wang-Pin Hsiung, K. Selçuk Candan, Divyakant Agrawal |
Data Knowl. Eng. | 3 |
| 2002 | View Invalidation for Dynamic Content Caching in Multitiered Architectures
K. Selçuk Candan, Divyakant Agrawal, Wen-Syan Li, Oliver Po, Wang-Pin Hsiung |
VLDB | 5 |
| 2002 | Issues and Evaluations of Caching Solutions for Web Application Acceleration
Wen-Syan Li, Wang-Pin Hsiung, Dmitri V. Kalashnikov, Radu Sion, Oliver Po, Divyakant Agrawal, K. Selçuk Candan |
VLDB | 2 |
| 2002 | Evaluations of architectural designs and implementation for database-driven web sites
Wen-Syan Li, Wang-Pin Hsiung, Oliver Po, K. Selçuk Candan, Divyakant Agrawal |
Data Knowl. Eng. | 2 |
| 2001 | Enabling Dynamic Content Caching for Database-Driven Web SitesabstractWeb performance is a key differentiation among content providers. Snafus and slowdowns at major web sites demonstrate the difficulty that companies face trying to scale to a large amount of web traffic. One solution to this problem is to store web content at server-side and edge-caches for fast delivery to the end users. However, for many e-commerce sites, web pages are created dynamically based on the current state of business processes, represented in application servers and databases. Since application servers, databases, web servers, and caches are independent components, there is no efficient mechanism to make changes in the database content reflected to the cached web pages. As a result, most application servers have to mark dynamically generated web pages as non-cacheable. In this paper, we describe the architectural framework of the CachePortal system for enabling dynamic content caching for database-driven e-commerce sites. We describe techniques for intelligently invalidating dynamically generated web pages in the caches, thereby enabling caching of web pages generated based on database contents. We use some of the most popular components in the industry to illustrate the deployment and applicability of the proposed architecture. K. Selçuk Candan, Wen-Syan Li, Qiong Luo 0001, Wang-Pin Hsiung, Divyakant Agrawal |
SIGMOD Conference | 4 |
| 2001 | Cache Portal: Technology for Accelerating Database-driven e-commerce Web Sites
Wen-Syan Li, K. Selçuk Candan, Wang-Pin Hsiung, Oliver Po, Divyakant Agrawal, Qiong Luo 0001, Wei-Kuang Waine Huang, Yusuf Akca |
VLDB | 3 |