EDBT 2026 Demo / reviewers in the wild / expert
Gang Luo 0001
dblp:22/793
· DBLP profile ↗
36ranked-venue papers
31as first author
3since 2021 · last 2026
0000-0001-7217-4008ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 29 · 25 first-author · 3 since 2021Artificial intelligence and machine learning · 10 · 8 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 5 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-authorSoftware engineering, systems software and programming languages · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Using Brand Knowledge Bases and LLM Agents to Enhance E-commerce Retailers' Catalog QualityabstractFor e-commerce retailers, high-quality product catalogs are vital to customer experience. Yet, despite lots of data cleaning efforts, catalog quality, especially in large catalogs, remains suboptimal. This paper shows how to use unstructured brand knowledge base data as a reference and a large language model agent to automatically enhance an e-commerce retailer's catalog quality. Unlike prior methods that usually repair and match product entries separately, our method does both concurrently. Our evaluation results show its effectiveness. Hayreddin Çeker, Gang Luo 0001, Kee Kiat Koo, Prashant Mathur, Wencong You, Atharva Amdekar, Rob Barton, Navaneet K. L., Vidit Bansal, Karim Bouyarmane |
WSDM | 2 |
| 2025 | Using Large Language Models to Improve Product Information in E-commerce CatalogsabstractTo give customers good experience, an e-commerce retailer needs high-quality product information in its catalog. Yet, the raw product information often lacks sufficient quality. For a large catalog that can contain billions of products, manually fixing this information is highly labor-intensive. To address this issue, we propose using the tool use functionality of large language models to automatically improve product information. In this talk, we show why existing data cleaning methods are not well suited for this task and how we designed our automated system to improve product information. When evaluated on a random sample of products from an e-commerce catalog, our system improved product information completeness by 78% with no major drop in information accuracy. Gang Luo 0001, Julien Han, Hayreddin Çeker, Karim Bouyarmane |
CIKM | 1 |
| 2023 | Proactive and Automatic Detection of Product Misclassifications at Massive ScaleabstractIn e-commerce, product classification is widely used for various purposes. Misclassifying products can cause compliance issues and hurt the company's reputation. To address this problem, we propose an automated system to proactively detect product misclassifications by overcoming several challenges. A large e-commerce retailer can sell billions of distinct products, on which many thousands of classification tasks are performed. At this massive scale, we need to quickly detect misclassifications under a limited budget. In this talk, we point out these challenges and show how we design our system to handle them. When evaluated on a set of Amazon's product classification data, at an overhead of <10% of the classification cost, our system automatically identified and corrected many misclassifications, which would take a human many thousand years to manually find and 14.6 years to manually review and correct if our system were not used. Ling Jiang 0003, Xiaoyu Chu, Saaransh Gulati, Pulkit Garg, Andrew Borthwick, Gang Luo 0001 |
CIKM | 6 |
| 2019 | Guest editorial: special issue on data management and analytics for healthcare
Fusheng Wang 0001, Gang Luo 0001 |
Distributed Parallel Databases | 2 |
| 2017 | Automatic identification of high impact articles in PubMed to support clinical decision makingabstractOBJECTIVES: The practice of evidence-based medicine involves integrating the latest best available evidence into patient care decisions. Yet, critical barriers exist for clinicians' retrieval of evidence that is relevant for a particular patient from primary sources such as randomized controlled trials and meta-analyses. To help address those barriers, we investigated machine learning algorithms that find clinical studies with high clinical impact from PubMed®. METHODS: Our machine learning algorithms use a variety of features including bibliometric features (e.g., citation count), social media attention, journal impact factors, and citation metadata. The algorithms were developed and evaluated with a gold standard composed of 502 high impact clinical studies that are referenced in 11 clinical evidence-based guidelines on the treatment of various diseases. We tested the following hypotheses: (1) our high impact classifier outperforms a state-of-the-art classifier based on citation metadata and citation terms, and PubMed's® relevance sort algorithm; and (2) the performance of our high impact classifier does not decrease significantly after removing proprietary features such as citation count. RESULTS: The mean top 20 precision of our high impact classifier was 34% versus 11% for the state-of-the-art classifier and 4% for PubMed's® relevance sort (p=0.009); and the performance of our high impact classifier did not decrease significantly after removing proprietary features (mean top 20 precision=34% vs. 36%; p=0.085). CONCLUSION: The high impact classifier, using features such as bibliometrics, social media attention and MEDLINE® metadata, outperformed previous approaches and is a promising alternative to identifying high impact studies for clinical decision support. Jiantao Bian, Mohammad Amin Morid, Siddhartha Jonnalagadda, Gang Luo 0001, Guilherme Del Fiol |
J. Biomed. Informatics | 4 |
| 2016 | Efficient Execution Methods of Pivoting for Bulk Extraction of Entity-Attribute-Value-Modeled DataabstractEntity-attribute-value (EAV) tables are widely used to store data in electronic medical records and clinical study data management systems. Before they can be used by various analytical (e.g., data mining and machine learning) programs, EAV-modeled data usually must be transformed into conventional relational table format through pivot operations. This time-consuming and resource-intensive process is often performed repeatedly on a regular basis, e.g., to provide a daily refresh of the content in a clinical data warehouse. Thus, it would be beneficial to make pivot operations as efficient as possible. In this paper, we present three techniques for improving the efficiency of pivot operations: 1) filtering out EAV tuples related to unneeded clinical parameters early on; 2) supporting pivoting across multiple EAV tables; and 3) conducting multi-query optimization. We demonstrate the effectiveness of our techniques through implementation. We show that our optimized execution method of pivoting using these techniques significantly outperforms the current basic execution method of pivoting. Our techniques can be used to build a data extraction tool to simplify the specification of and improve the efficiency of extracting data from the EAV tables in electronic medical records and clinical study data management systems. Gang Luo 0001, Lewis J. Frey |
IEEE J. Biomed. Health Informatics | 1 |
| 2010 | V Locking Protocol for Materialized Aggregate Join Views on B-Tree Indices
Gang Luo 0001 |
WAIM | 1 |
| 2010 | Transaction reordering
Gang Luo 0001, Jeffrey F. Naughton, Curt J. Ellmann, Michael Watzke |
Data Knowl. Eng. | 1 |
| 2009 | Automatic Home Nursing Activity Recommendation
Gang Luo 0001, Chunqiang Tang |
AMIA | 1 |
| 2009 | Design and Evaluation of the iMed Intelligent Medical Search EngineabstractSearching for medical information on the Web is popular and important. However, medical search has its own unique requirements that are poorly handled by existing medical Web search engines. This paper presents iMed, the first intelligent medical Web search engine that extensively uses medical knowledge and questionnaire to facilitate ordinary Internet users to search for medical information. iMed introduces and extends expert system technology into the search engine domain. It uses several key techniques to improve its usability and search result quality. First, since ordinary users often cannot clearly describe their situations due to lack of medical background, iMed uses a questionnaire-based query interface to guide searchers to provide the most important information about their situations. Second, iMed uses medical knowledge to automatically form multiple queries from a searcher' answers to the questions. Using these queries to perform search can significantly improve the quality of search results. Third, iMed structures all the search results into a multi-level hierarchy with explicitly marked medical meanings to facilitate searchers' viewing. Lastly, iMed suggests diversified, related medical phrases at each level of the search result hierarchy. These medical phrases are extracted from the MeSH ontology and can help searchers quickly digest search results and refine their inputs. We evaluated iMed under a wide range of medical scenarios. The results show that iMed is effective and efficient for medical search. Gang Luo 0001 |
ICDE | 1 |
| 2009 | Answering linear optimization queries with an approximate stream index
Gang Luo 0001, Kun-Lung Wu, Philip S. Yu |
Knowl. Inf. Syst. | 1 |
| 2008 | Intelligent Output Interface for Intelligent Medical Search Engine
Gang Luo 0001 |
AAAI | 1 |
| 2008 | Transaction reordering with application to synchronized scansabstractTraditional workload management methods mainly focus on the current system status while information about the interaction between queued and running transactions is largely ignored. An exception to this is the transaction reordering method, which reorders the transaction sequence submitted to the RDBMS and improves the transaction throughput by considering both the current system status and information about the interaction between queued and running transactions. The existing transaction reordering method only considers the reordering opportunities provided by analyzing the lock conflict information among multiple transactions. This significantly limits the applicability of the transaction reordering method. In this paper, we extend the existing transaction reordering method into a general transaction reordering framework that can incorporate various factors as the reordering criteria. We show that by analyzing the resource utilization information of transactions, the transaction reordering method can also improve the system throughput by increasing the resource sharing opportunities among multiple transactions. We provide a concrete example on synchronized scans and demonstrate the advantages of our method through experiments with a commercial parallel RDBMS. Gang Luo 0001, Jeffrey F. Naughton, Curt J. Ellmann, Michael Watzke |
CIKM | 1 |
| 2008 | MedSearch: a specialized search engine for medical information retrievalabstractPeople are thirsty for medical information. Existing Web search engines often cannot handle medical search well because they do not consider its special requirements. Often a medical information searcher is uncertain about his exact questions and unfamiliar with medical terminology. Therefore, he sometimes prefers to pose long queries, describing his symptoms and situation in plain English, and receive comprehensive, relevant information from search results. This paper presents MedSearch, a specialized medical Web search engine, to address these challenges. MedSearch uses several key techniques to improve its usability and the quality of search results. First, it accepts queries of extended length and reforms long queries into shorter queries by extracting a subset of important and representative words. This not only significantly increases the query processing speed but also improves the quality of search results. Second, it provides diversified search results. Lastly, it suggests related medical phrases to help the user quickly digest search results and refine the query. We evaluated MedSearch using medical questions posted on medical discussion forums. The results show that MedSearch can handle various medical queries effectively and efficiently. Gang Luo 0001, Chunqiang Tang |
CIKM | 1 |
| 2008 | Content-based filtering for efficient online materialized view maintenanceabstractReal-time materialized view maintenance has become increasingly popular, especially in real-time data warehousing and data streaming environments. Upon updates to base relations, maintaining the corresponding materialized views can bring a heavy burden to the RDBMS. A traditional method to mitigate this problem is to use the where clause condition in the materialized view definition to detect whether an update to a base relation is relevant and can affect the materialized view. However, this detection method does not consider the content in the base relations and hence misses a large number of filtering opportunities. In this paper, we propose a content-based method for detecting irrelevant updates to base relations of a materialized view. At the cost of using more space, this method increases the probability of catching irrelevant updates by judiciously designing filtering relations to capture the content in the base relations. Based on the content-based method, a prototype real-time data warehouse has been implemented on top of IBM's System S using IBM DB2. Using an analytical model and our prototype, we show that the content-based method can catch most (or all) irrelevant updates to base relations that are missed by the traditional method. Thus, when the fraction of irrelevant updates is non-negligible, the load on the RDBMS due to materialized view maintenance can be significantly reduced. Gang Luo 0001, Philip S. Yu |
CIKM | 1 |
| 2008 | Real-time new event detection for video streamsabstractOnline detection of video clips that present previously unseen events in a video stream is still an open challenge to date. For this online new event detection (ONED) task, existing studies mainly focus on optimizing the detection accuracy instead of the detection efficiency. As a result, it is difficult for existing systems to detect new events in real time, especially for large-scale video collections such as the video content available on the Web. In this paper, we propose several scalable techniques to improve the video processing speed of a baseline ONED system by orders of magnitude without sacrificing much detection accuracy. First, we use text features alone to filter out most of the non-new-event clips and to skip those expensive but unnecessary steps including image feature extraction and image similarity computation. Second, we use a combination of indexing and compression methods to speed up text processing. We implemented a prototype of our optimized ONED system on top of IBM's System S. The effectiveness of our techniques is evaluated on the standard TRECVID 2005 benchmark, which demonstrates that our techniques can achieve a 480-fold speedup with detection accuracy degraded less than 5%. Gang Luo 0001, Philip S. Yu |
CIKM | 1 |
| 2008 | Transaction reordering with application to synchronized scansabstractTraditional workload management methods mainly focus on the current system status while information about the interaction between queued and running transactions is largely ignored. An exception to this is the transaction reordering method, which reorders the transaction sequence submitted to the RDBMS and improves the transaction throughput by considering both the current system status and information about the interaction between queued and running transactions. The existing transaction reordering method only considers the reordering opportunities provided by analyzing the lock conflict information among multiple transactions. This significantly limits the applicability of the transaction reordering method. In this paper, we extend the existing transaction reordering method into a general transaction reordering framework that can incorporate various factors as the reordering criteria. We show that by analyzing the resource utilization information of transactions, the transaction reordering method can also improve the system throughput by increasing the resource sharing opportunities among multiple transactions. We provide a concrete example on synchronized scans and demonstrate the advantages of our method through experiments with a commercial parallel RDBMS. Gang Luo 0001, Jeffrey F. Naughton, Curt J. Ellmann, Michael Watzke |
DOLAP | 1 |
| 2008 | Challenging issues in iterative intelligent medical searchabstractSearching for medical information on the Web is highly popular these days. To facilitate ordinary people to perform medical search and preliminary disease self-diagnosis, we have built an intelligent medical Web search engine called iMed. iMed introduces and extends pattern recognition and expert system technology into the search engine domain. It uses medical knowledge and an interactive questionnaire to help searchers form queries. Due to searcherspsila limited medical knowledge and the taskpsilas inherent difficulty, searchers often cannot find desired search results in a single pass and have to search iteratively for multiple passes. For this purpose, iMed provides an iterative search advisor that guides searchers to refine their inputs. Based on our experience in building and using iMed, this paper summarizes the common difficulties faced by ordinary medical information searchers and the research issues that deserve attention from people working in the pattern recognition and medical search areas. Gang Luo 0001, Chunqiang Tang |
ICPR | 1 |
| 2008 | On iterative intelligent medical searchabstractSearching for medical information on the Web has become highly popular, but it remains a challenging task because searchers are often uncertain about their exact medical situations and unfamiliar with medical terminology. To address this challenge, we have built an intelligent medical Web search engine called iMed, which uses medical knowledge and an interactive questionnaire to help searchers form queries. This paper focuses on iMed's iterative search advisor, which integrates medical and linguistic knowledge to help searchers improve search results iteratively. Such an iterative process is common for general Web search, and especially crucial for medical Web search, because searchers often miss desired search results due to their limited medical knowledge and the task's inherent difficulty. iMed's iterative search advisor helps the searcher in several ways. First, relevant symptoms and signs are automatically suggested based on the searcher's description of his situation. Second, instead of taking for granted the searcher's answers to the questions, iMed ranks and recommends alternative answers according to their likelihoods of being the correct answers. Third, related MeSH medical phrases are suggested to help the searcher refine his situation description. We demonstrate the effectiveness of iMed's iterative search advisor by evaluating it using real medical case records and USMLE medical exam questions. Gang Luo 0001, Chunqiang Tang |
SIGIR | 1 |
| 2007 | Subject-Adaptive Real-Time Sleep Stage Classification Based on Conditional Random Field
Gang Luo 0001, Wanli Min |
AMIA | 1 |
| 2007 | Partial Materialized ViewsabstractEarly access to partial query results is highly desirable during exploration of massive data sets. However, it is challenging to provide transactionally consistent, immediate partial results without significantly increasing queries' execution time. To address this problem, this paper proposes a partial materialized view (PMV) method to cache some of the most frequently accessed results rather than all the possible results. Compared to traditional materialized views, the proposed PMVs do not require maintenance during insertion into base relations, and have much smaller storage and maintenance overhead. Upon the arrival of a query, the RDBMS first searches the PMV and returns to the user the cached partial results. Since a large portion of the PMV is cached in memory, this usually finishes within a millisecond. Then the RDBMS continues to execute the query to find the remaining results. The efficiency of our PMV method is evaluated through a simulation study, a theoretical analysis, and an initial implementation in PostgreSQL. Gang Luo 0001 |
ICDE | 1 |
| 2007 | SAO: A Stream Index for Answering Linear Optimization QueriesabstractLinear optimization queries retrieve the top-K tuples in a sliding window of a data stream that maximize/minimize the linearly weighted sums of certain attribute values. To efficiently answer such queries against a large relation, an onion index was previously proposed to properly organize all the tuples in the relation. However, such an onion index does not work in a streaming environment due to fast tuple arrival rate and limited memory. In this paper, we propose a SAO index to approximately answer arbitrary linear optimization queries against a data stream. It uses a small amount of memory to efficiently keep track of the most "important" tuples in a sliding window of a data stream. The index maintenance cost is small because the great majority of the incoming tuples do not cause any changes to the index and are quickly discarded. At any time, for any linear optimization query, we can retrieve from the SAO index the approximate top-K tuples in the sliding window almost instantly. The larger the amount of available memory, the better the quality of the answers is. More importantly, for a given amount of memory, the quality of the answers can be further improved by dynamically allocating a larger portion of the memory to the outer layers of the SAO index. We evaluate the effectiveness of this SAO index through a prototype implementation. Gang Luo 0001, Kun-Lung Wu, Philip S. Yu |
ICDE | 1 |
| 2007 | Resource-adaptive real-time new event detectionabstractIn a document streaming environment, online detection of the first documents that mention previously unseen events is an open challenge. For this online new event detection (ONED) task, existing studies usually assume that enough resources are always available and focus entirely on detection accuracy without considering efficiency. Moreover, none of the existing work addresses the issue of providing an effective and friendly user interface. As a result, there is a significant gap between the existing systems and a system that can be used in practice. In this paper, we propose an ONED framework with the following prominent features. First, a combination of indexing and compression methods is used to improve the document processing rate by orders of magnitude without sacrificing much detection accuracy. Second, when resources are tight, a resource-adaptive computation method is used to maximize the benefit that can be gained from the limited resources. Third, when the new event arrival rate is beyond the processing capability of the consumer of the ONED system, new events are further filtered and prioritized before they are presented to the consumer. Fourth, implicit citation relationships are created among all the documents and used to compute the importance of document sources. This importance information can guide the selection of document sources. We implemented a prototype of our framework on top of IBM's Stream Processing Core middleware. We also evaluated the effectiveness of our techniques on the standard TDT5 benchmark. To the best of our knowledge, this is the first implementation of a real application in a large-scale stream processing system. Gang Luo 0001, Chunqiang Tang, Philip S. Yu |
SIGMOD Conference | 1 |
| 2007 | Challenges and Experience in Prototyping a Multi-Modal Stream Analytic and Monitoring Application on System S
Kun-Lung Wu, Philip S. Yu, Bugra Gedik, Kirsten Hildrum, Charu C. Aggarwal, Eric Bouillet, Wei Fan 0001, Xiaohui Gu, Gang Luo 0001, Haixun Wang |
VLDB | 10 |
| 2007 | Answering relationship queries on the webabstractFinding relationships between entities on the Web, e.g., the connections between different places or the commonalities of people, is a novel and challenging problem. Existing Web search engines excel in keyword matching and document ranking, but they cannot well handle many relationship queries. This paper proposes a new method for answering relationship queries on two entities. Our method first respectively retrieves the top Web pages for either entity from a Web search engine. It then matches these Web pages and generates an ordered list of Web page pairs. Each Web page pair consists of one Web page for either entity. The top ranked Web page pairs are likely to contain the relationships between the two entities. One main challenge in the ranking process is to effectively filter out the large amount of noise in the Web pages without losing much useful information. To achieve this, our method assigns appropriate weights to terms in Web pages and intelligently identifies the potential connecting terms that capture the relationships between the two entities. Only those top potential connecting terms with large weights are used to rank Web page pairs. Finally, the top ranked Web page pairs are presented to the searcher. For each such pair, the query terms and the top potential connecting terms are properly highlighted so that the relationships between the two entities can be easily identified. We implemented a prototype on top of the Google search engine and evaluated it under a wide variety of query scenarios. The experimental results show that our method is effective at finding important relationships with low overhead. Gang Luo 0001, Chunqiang Tang, Yingli Tian |
WWW | 1 |
| 2007 | MedSearch: a specialized search engine for medical informationabstractPeople are thirsty for medical information. Existing Web search engines cannot handle medical search well because they do not consider its special requirements. Often a medical information searcher is uncertain about his exact questions and unfamiliar with medical terminology. Therefore, he prefers to pose long queries, describing his symptoms and situation in plain English, and receive comprehensive, relevant information from search results. This paper presents MedSearch, a specialized medical Web search engine, to address these challenges. MedSearch can assist ordinary Internet users to search for medical information, by accepting queries of extended length, providing diversified search results, and suggesting related medical phrases. Gang Luo 0001, Chunqiang Tang |
WWW | 1 |
| 2007 | Toward a progress indicator for program compilationabstractAbstract For user‐friendliness purposes, many modern software systems provide progress indicators for long‐running tasks. These progress indicators continuously estimate the percentage of the task that has been completed and when the task will finish. However, none of the existing program compilation tools provide a non‐trivial progress indicator, although it often takes minutes or hours to build a large program. In this paper, we investigate the problem of supporting such progress indicators. We first discuss the goals and challenges inherent in this problem. Then we present a set of techniques that are sufficient for implementing a simple yet useful progress indicator for program compilation. Finally, we report on an initial implementation of these techniques in GNU Make. Copyright © 2006 John Wiley & Sons, Ltd. Gang Luo 0001 |
Softw. Pract. Exp. | 1 |
| 2006 | Multi-query SQL Progress Indicators
Gang Luo 0001, Jeffrey F. Naughton, Philip S. Yu |
EDBT | 1 |
| 2006 | Efficient Detection of Empty-Result Queries
Gang Luo 0001 |
VLDB | 1 |
| 2005 | Increasing the Accuracy and Coverage of SQL Progress IndicatorsabstractRecently, progress indicators have been proposed for long-running SQL queries in RDBMSs. Although the proposed techniques work well for a subset of SQL queries, they are preliminary in the sense that (1) they cannot provide non-trivial estimates for some SQL queries, and (2) the provided estimates can be rather imprecise in certain cases. In this paper, we consider the problem of supporting non-trivial progress indicators for a wider class of SQL queries with more precise estimates. We present a set of techniques in achieving this goal. We report an initial implementation of these techniques in PostgreSQL. Gang Luo 0001, Jeffrey F. Naughton, Curt J. Ellmann, Michael Watzke |
ICDE | 1 |
| 2005 | Locking Protocols for Materialized Aggregate Join ViewsabstractThe maintenance of materialized aggregate join views is a well-studied problem. However, to date the published literature has largely ignored the issue of concurrency control. Clearly, immediate materialized view maintenance with transactional consistency, if enforced by generic concurrency control mechanisms, can result in low levels of concurrency and high rates of deadlock. While this problem is superficially amenable to well-known techniques, such as fine-granularity locking and special lock modes for updates that are associative and commutative, we show that these previous high concurrency locking techniques do not fully solve the problem, but a combination of a "value-based" latch pool and these previous high concurrency locking techniques can solve the problem. Gang Luo 0001, Jeffrey F. Naughton, Curt J. Ellmann, Michael Watzke |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2004 | Toward a Progress Indicator for Database QueriesabstractMany modern software systems provide progress indicators for long-running tasks. These progress indicators make systems more user-friendly by helping the user quickly estimate how much of the task has been completed and when the task will finish. However, none of the existing commercial RDBMSs provides a non-trivial progress indicator for long-running queries. In this paper, we consider the problem of supporting such progress indicators. After discussing the goals and challenges inherent in this problem, we present a set of techniques sufficient for implementing a simple yet useful progress indicator for a large subset of RDBMS queries. We report an initial implementation of these techniques in PostgreSQL. 1. Gang Luo 0001, Jeffrey F. Naughton, Curt J. Ellmann, Michael Watzke |
SIGMOD Conference | 1 |
| 2003 | A Comparison of Three Methods for Join View Maintenance in Parallel RDBMSabstractIn a typical data warehouse, materialized views are used to speed up query execution. Upon updates to the base relations in the warehouse, these materialized views must also be maintained. The need to maintain these materialized views can have a negative impact on performance that is exacerbated in parallel RDBMSs, since simple single-node updates to base relations can give rise to expensive all-node operations for materialized view maintenance. We present a comparison of three materialized join view maintenance methods in a parallel RDBMS, which we refer to as the naive, auxiliary relation, and global index methods. The last two methods improve performance at the cost of using more space. The results of this study show that the method of choice depends on the environment, in particular, the update activity on base relations and the amount of available storage space. Gang Luo 0001, Jeffrey F. Naughton, Curt J. Ellmann, Michael Watzke |
ICDE | 1 |
| 2003 | Locking Protocols for Materialized Aggregate Join Views
Gang Luo 0001, Jeffrey F. Naughton, Curt J. Ellmann, Michael Watzke |
VLDB | 1 |
| 2002 | A Non-Blocking Parallel Spatial Join AlgorithmabstractInterest in incremental and adaptive query processing has led to the investigation of equijoin evaluation algorithms that are non-blocking. This investigation has yielded a number of algorithms, including the symmetric hash join, the XJoin, the Ripple Join, and their variants. However, to our knowledge no one has proposed a nonblocking spatial join algorithm. In this paper, we propose a parallel non-blocking spatial join algorithm that uses duplicate avoidance rather than duplicate elimination. Results from a prototype implementation in a commercial parallel object-relational DBMS show that it generates answer tuples steadily even in the presence of memory overflow, and that its rate of producing answer tuples scales with the number of processors. Also, when allowed to run to completion, its performance is comparable with the state-of-the-art blocking parallel spatial join algorithm. Gang Luo 0001, Jeffrey F. Naughton, Curt J. Ellmann |
ICDE | 1 |
| 2002 | A scalable hash ripple join algorithmabstractRecently, Haas and Hellerstein proposed the hash ripple join algorithm in the context of online aggregation. Although the algorithm rapidly gives a good estimate for many join-aggregate problem instances, the convergence can be slow if the number of tuples that satisfy the join predicate is small or if there are many groups in the output. Furthermore, if memory overflows (for example, because the user allows the algorithm to run to completion for an exact answer), the algorithm degenerates to block ripple join and performance suffers. In this paper, we build on the work of Haas and Hellerstein and propose a new algorithm that (a) combines parallelism with sampling to speed convergence, and (b) maintains good performance in the presence of memory overflow. Results from a prototype implementation in a parallel DBMS show that its rate of convergence scales with the number of processors, and that when allowed to run to completion, even in the presence of memory overflow, it is competitive with the traditional parallel hybrid hash join algorithm. Gang Luo 0001, Curt J. Ellmann, Peter J. Haas, Jeffrey F. Naughton |
SIGMOD Conference | 1 |