EDBT 2026 Demo / reviewers in the wild / expert
Hongrae Lee
dblp:73/5604
· DBLP profile ↗
23ranked-venue papers
4as first author
2since 2021 · last 2023
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 20 · 4 first-author · 1 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Systems, architecture and hardware · 1Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
15 papers |
Information retrieval · 42% Data integration and cleaning · 21% Query processing and optimization · 12% | |
| Artificial intelligence
4 papers |
Language models and text generation · 90% Information extraction and text analysis · 10% | |
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
Storage systems · 49% Distributed systems · 28% Cloud and datacenter computing · 23% | |
| Theoretical computer science
1 paper |
Mathematical optimization · 50% Computational geometry · 50% |
Topics — the 29 heaviest of 40, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation
text generation |
1.0 | 2 | 2023 | RARR: Researching and Revising What Language Models Say, Using Language Models · ACL (1) 2023 Generating Titles for Web Tables · WWW 2019 |
Natural language and speech › Language models and text generation › trustworthy language model › large language model reliability › factuality
factual consistency |
0.7 | 1 | 2023 | RARR: Researching and Revising What Language Models Say, Using Language Models · ACL (1) 2023 |
Information retrieval
evidence retrieval |
0.7 | 1 | 2023 | RARR: Researching and Revising What Language Models Say, Using Language Models · ACL (1) 2023 |
Information retrieval
similarity search |
0.6 | 2 | 2021 | Substring Similarity Search with Synonyms · ICDE 2021 Efficient Exact Similarity Searches Using Multiple Token Orderings · ICDE 2012 |
Information retrieval
string matching |
0.6 | 2 | 2021 | Substring Similarity Search with Synonyms · ICDE 2021 Extending Q-Grams to Estimate Selectivity of String Matching with Low Edit Distance · VLDB 2007 |
Data integration and cleaning › approximate matching
synonym-aware string similarity |
0.5 | 1 | 2021 | Substring Similarity Search with Synonyms · ICDE 2021 |
Data models and query languages › natural language interface › natural language interface to database
text-to-SQL |
0.4 | 1 | 2020 | Natural language to SQL: Where are we today? · Proc. VLDB Endow. 2020 |
Data integration and cleaning
table understanding |
0.4 | 1 | 2019 | Generating Titles for Web Tables · WWW 2019 |
Data integration and cleaning
web data integration |
0.3 | 1 | 2018 | Ten Years of WebTables · Proc. VLDB Endow. 2018 |
Storage systems
flash and SSD |
0.2 | 1 | 2016 | Using SSDs to scale up Google Fusion Tables, a database-in-the-cloud · ICDE 2016 |
Query processing and optimization
similarity join |
0.2 | 2 | 2011 | Similarity Join Size Estimation using Locality Sensitive Hashing · Proc. VLDB Endow. 2011 Power-Law Based Estimation of Set Similarity Join Size · Proc. VLDB Endow. 2009 |
Natural language and speech › Information extraction and text analysis
sentiment analysis |
0.2 | 1 | 2015 | Mining Subjective Properties on the Web · SIGMOD Conference 2015 |
Mathematical optimization › integer programming
integer linear programming formulation |
0.2 | 1 | 2013 | Consistent thinning of large geographical data for map visualization · ACM Trans. Database Syst. 2013 |
Information retrieval › similarity search
exact similarity search |
0.1 | 1 | 2012 | Efficient Exact Similarity Searches Using Multiple Token Orderings · ICDE 2012 |
Spatial and temporal data management
spatial sampling |
0.1 | 1 | 2012 | Efficient spatial sampling of large geographical tables · SIGMOD Conference 2012 |
Information retrieval › search engines › structured data search
table retrieval |
0.1 | 1 | 2012 | Finding related tables · SIGMOD Conference 2012 |
Visualization and visual analytics › geospatial visualization
cartographic visualization |
0.1 | 1 | 2012 | Efficient spatial sampling of large geographical tables · SIGMOD Conference 2012 |
Distributed systems › distributed algorithms
distributed sorting |
0.1 | 1 | 2012 | CloudRAMSort: fast and efficient large-scale distributed RAM sort on shared-nothing cluster · SIGMOD Conference 2012 |
Query processing and optimization
cardinality estimation |
0.1 | 1 | 2011 | Similarity Join Size Estimation using Locality Sensitive Hashing · Proc. VLDB Endow. 2011 |
Query processing and optimization › cardinality estimation
sampling-based estimation |
0.1 | 1 | 2011 | Similarity Join Size Estimation using Locality Sensitive Hashing · Proc. VLDB Endow. 2011 |
Query processing and optimization
parameterized queries |
0.1 | 1 | 2010 | Variance aware optimization of parameterized queries · SIGMOD Conference 2010 |
Query processing and optimization › query planning
query plan selection |
0.1 | 1 | 2010 | Variance aware optimization of parameterized queries · SIGMOD Conference 2010 |
Data mining
pattern mining |
0.1 | 1 | 2009 | Power-Law Based Estimation of Set Similarity Join Size · Proc. VLDB Endow. 2009 |
Cloud and datacenter computing › cloud data management
cloud data service |
0.1 | 1 | 2016 | Using SSDs to scale up Google Fusion Tables, a database-in-the-cloud · ICDE 2016 |
Query processing and optimization
selectivity estimation |
0.1 | 1 | 2007 | Extending Q-Grams to Estimate Selectivity of String Matching with Low Edit Distance · VLDB 2007 |
Information retrieval › similarity search
near-duplicate detection |
0.0 | 1 | 2012 | Efficient Exact Similarity Searches Using Multiple Token Orderings · ICDE 2012 |
Information retrieval › ranking › content ranking
table ranking |
0.0 | 1 | 2012 | Finding related tables · SIGMOD Conference 2012 |
Cloud and datacenter computing › datacenter architecture
shared-nothing architecture |
0.0 | 1 | 2012 | CloudRAMSort: fast and efficient large-scale distributed RAM sort on shared-nothing cluster · SIGMOD Conference 2012 |
Information retrieval › similarity search
approximate string matching |
0.0 | 1 | 2007 | Extending Q-Grams to Estimate Selectivity of String Matching with Low Edit Distance · VLDB 2007 |
Methods — techniques the papers use, named apart from their topics
retrieval-augmented generation · 1.3language model revision · 1.3sequence-to-sequence · 0.8copy mechanism · 0.8filtering methods · 0.5probabilistic model · 0.4query rewrite · 0.4deep learning · 0.4database testing · 0.4DFS traversal · 0.3recurrent neural network · 0.3policy gradient · 0.3LSTM · 0.3randomized algorithm · 0.2integer programming · 0.2spatial sampling · 0.1single-node sorting · 0.1inter-node communication optimization · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | RARR: Researching and Revising What Language Models Say, Using Language ModelsabstractLuyu Gao, Zhuyun Dai, Panupong Pasupat, Anthony Chen, Arun Tejasvi Chaganty, Yicheng Fan, Vincent Zhao, Ni Lao, Hongrae Lee, Da-Cheng Juan, Kelvin Guu. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Luyu Gao, Zhuyun Dai, Panupong Pasupat, Anthony Chen, Arun Tejasvi Chaganty, Yicheng Fan, Vincent Y. Zhao, Ni Lao, Hongrae Lee, Da-Cheng Juan, Kelvin Guu |
ACL (1) | 9 |
| 2021 | Substring Similarity Search with SynonymsabstractTo allow deeper semantic understanding of strings, string matching with synonyms has recently received increasing attention. However, all the works focus on string semantics, which requires matching of entire strings, and cannot handle partial matching with substring semantics, resulting in its limited applications. To remedy this issue, we first propose a novel similarity measure between two strings allowing substring matching with synonyms, and develop an efficient algorithm to find the strings that have a substring semantically similar to the query string. Since considering synonyms on substrings enlarges the search space significantly, we devise efficient filtering methods to reduce the number of expensive similarity computations. Experiments with real-life datasets show the efficiency of our algorithm. Gwangho Song, Kyuseok Shim, Hongrae Lee |
ICDE | 3 |
| 2020 | String Joins with Synonyms
Gwangho Song, Hongrae Lee, Kyuseok Shim, Yoonjae Park, Wooyeol Kim |
DASFAA (3) | 2 |
| 2020 | Natural language to SQL: Where are we today?abstractTranslating natural language to SQL (NL2SQL) has received extensive attention lately, especially with the recent success of deep learning technologies. However, despite the large number of studies, we do not have a thorough understanding of how good existing techniques really are and how much is applicable to real-world situations. A key difficulty is that different studies are based on different datasets, which often have their own limitations and assumptions that are implicitly hidden in the context or datasets. Moreover, a couple of evaluation metrics are commonly employed but they are rather simplistic and do not properly depict the accuracy of results, as will be shown in our experiments. To provide a holistic view of NL2SQL technologies and access current advancements, we perform extensive experiments under our unified framework using eleven of recent techniques over 10+ benchmarks including a new benchmark (WTQ) and TPC-H. We provide a comprehensive survey of recent NL2SQL methods, introducing a taxonomy of them. We reveal major assumptions of the methods and classify translation errors through extensive experiments. We also provide a practical tool for validation by using existing, mature database technologies such as query rewrite and database testing. We then suggest future research directions so that the translation can be used in practice. Hyeonji Kim, Byeong-Hoon So, Wook-Shin Han, Hongrae Lee |
Proc. VLDB Endow. | 4 |
| 2019 | Generating Titles for Web TablesabstractDescriptive titles provide crucial context for interpreting tables that are extracted from web pages and are a key component of search features such as tabular featured snippets from Google and Bing. Prior approaches have attempted to produce titles by selecting existing text snippets associated with the table. These approaches, however, are limited by their dependence on suitable titles existing a priori. In our user study, we observe that the relevant information for the title tends to be scattered across the page, and often-more than 80% of the time-does not appear verbatim anywhere in the page. We propose instead the application of a sequence-to-sequence neural network model as a more generalizable approach for generating high-quality table titles. This is accomplished by extracting many text snippets that have potentially relevant information to the table, encoding them into an input sequence, and using both copy and generation mechanisms in the decoder to balance relevance and readability of the generated title. We validate this approach with human evaluation on sample web tables and report that while sequence models with only a copy mechanism or only a generation mechanism are easily outperformed by simple selection-based baselines, the model with both capabilities performs the best, approaching the quality of crowdsourced titles while training on fewer than ten thousand examples. To the best of our knowledge, the proposed technique is the first to consider text-generation methods for table titles, and establishes a new state of the art. Braden Hancock, Hongrae Lee, Cong Yu 0001 |
WWW | 2 |
| 2018 | Ten Years of WebTablesabstractIn 2008, we wrote about WebTables, an effort to exploit the large and diverse set of structured databases casually published online in the form of HTML tables. The past decade has seen a flurry of research and commercial activities around the WebTables project itself, as well as the broad topic of informal online structured data. In this paper, we 1 will review the WebTables project, and try to place it in the broader context of the decade of work that followed. We will also show how the progress over the past ten years sets up an exciting agenda for the future, and will draw upon many corners of the data management community. Michael J. Cafarella, Alon Y. Halevy, Hongrae Lee, Jayant Madhavan, Cong Yu 0001, Daisy Zhe Wang, Eugene Wu 0002 |
Proc. VLDB Endow. | 3 |
| 2017 | Learning to Skim TextabstractRecurrent Neural Networks are showing much promise in many sub-areas of natural language processing, ranging from document classification to machine translation to automatic question answering.Despite their promise, many recurrent models have to read the whole text word by word, making it slow to handle long documents.For example, it is difficult to use a recurrent network to read a book and answer questions about it.In this paper, we present an approach of reading text while skipping irrelevant information if needed.The underlying model is a recurrent network that learns how far to jump after reading a few words of the input text.We employ a standard policy gradient method to train the model to make discrete jumping decisions.In our benchmarks on four different tasks, including number prediction, sentiment analysis, news article classification and automatic Q&A, our proposed model, a modified LSTM with jumping, is up to 6 times faster than the standard sequential LSTM, while maintaining the same or even better accuracy. Adams Wei Yu, Hongrae Lee, Quoc V. Le |
ACL (1) | 2 |
| 2016 | Using SSDs to scale up Google Fusion Tables, a database-in-the-cloudabstractFlash memory solid state drives (SSDs) have increasingly been advocated and adopted as a means of speeding up and scaling up data-driven applications. SSDs are becoming more widely available as an option in the cloud. However, when an application considers SSDs in the cloud, the best option for the application may not be immediate, among a number of choices for placing SSDs in the layers of the cloud. Although there have been many studies on SSDs, they often concern a specific setting, and how different SSD options in the cloud compare with each other is less well understood. In this paper, we describe how Google Fusion Tables (GFT) used SSDs and what optimizations were implemented to scale up its in-memory processing, clearly showing opportunities and limitations of SSDs in the cloud with quantitative analyses. We first discuss various SSD placement strategies and compare them with low-level measurements, and propose SSD-placement guidelines for a variety of cloud data services. We then present internals of our column engine and optimizations to better use the performance characteristics of SSDs. We empirically demonstrate that the optimizations enable us to scale our application to much larger datasets while retaining the low-latency and simple query processing architecture. Yingyi Bu, Felix Halim, Changkyu Kim, Hongrae Lee, Jayant Madhavan |
ICDE | 4 |
| 2016 | EIC EditorialabstractPresents the introductory editorial for this issue of the publication. Jian Pei 0001, Leman Akoglu, Hongrae Lee, Justin J. Levandoski, Xuelong Li 0001, Rosa Meo, Carlos Ordonez 0001, Jeff M. Phillips, Barbara Poblete, K. Selçuk Candan, Meng Wang 0001, Ji-Rong Wen, Li Xiong 0001, Wenjie Zhang 0001 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2015 | Applying WebTables in Practice
Sreeram Balakrishnan, Alon Y. Halevy, Boulos Harb, Hongrae Lee, Jayant Madhavan, Afshin Rostamizadeh, Warren Shen, Kenneth Wilder, Fei Wu 0003, Cong Yu 0001 |
CIDR | 4 |
| 2015 | Mining Subjective Properties on the WebabstractEven with the recent developments in Web search of answering queries from structured data, search engines are still limited to queries with an objective answer, such as EUROPEAN CAPITALS or WOODY ALLEN MOVIES. However, many queries are subjective, such as SAFE CITIES, or CUTE ANIMALS. The underlying knowledge bases of search engines do not contain answers to these queries because they do not have a ground truth. We describe the Surveyor system that mines the dominant opinion held by authors of Web content about whether a subjective property applies to a given entity. The evidence on which SURVEYOR relies is statements extracted from Web text that either support the property or claim its negation. The key challenge that SURVEYOR faces is that simply counting the number of positive and negative statements does not suffice, because there are multiple hidden biases with which content tends to be authored on the Web. SURVEYOR employs a probabilistic model of how content is authored on the Web. As one example, this model accounts for correlations between the subjective property and the frequency with which it is mentioned on the Web. The parameters of the model are specialized to each property and entity type. Immanuel Trummer, Alon Y. Halevy, Hongrae Lee, Sunita Sarawagi |
SIGMOD Conference | 3 |
| 2013 | Comparing SSD-placement strategies to scale a database-in-the-cloudabstractFlash memory solid state drives (SSDs) have increasingly been advocated and adopted as a means of speeding up and scaling up data-driven applications. However, given the layered software architecture of cloud-based services, there are a number of options available for placing SSDs. In this work, we studied the trade-offs involved in different SSD placement strategies, their impact of response time and throughput, and ultimately the potential in achieving scalability in Google Fusion Tables (GFT), a cloud-based service for data management and visualization [1]. Yingyi Bu, Hongrae Lee, Jayant Madhavan |
SoCC | 2 |
| 2013 | Recent progress towards an ecosystem of structured data on the WebabstractGoogle Fusion Tables aims to support an ecosystem of structured data on the Web by providing a tool for managing and visualizing data on the one hand, and for searching and exploring for data on the other. This paper describes a few recent developments in our efforts to further the ecosystem. Nitin Gupta 0003, Alon Y. Halevy, Boulos Harb, Heidi Lam, Hongrae Lee, Jayant Madhavan, Fei Wu 0003, Cong Yu 0001 |
ICDE | 5 |
| 2013 | Consistent thinning of large geographical data for map visualizationabstractLarge-scale map visualization systems play an increasingly important role in presenting geographic datasets to end-users. Since these datasets can be extremely large, a map rendering system often needs to select a small fraction of the data to visualize them in a limited space. This article addresses the fundamental challenge of thinning : determining appropriate samples of data to be shown on specific geographical regions and zoom levels. Other than the sheer scale of the data, the thinning problem is challenging because of a number of other reasons: (1) data can consist of complex geographical shapes, (2) rendering of data needs to satisfy certain constraints, such as data being preserved across zoom levels and adjacent regions, and (3) after satisfying the constraints, an optimal solution needs to be chosen based on objectives such as maximality , fairness , and importance of data. This article formally defines and presents a complete solution to the thinning problem. First, we express the problem as an integer programming formulation that efficiently solves thinning for desired objectives. Second, we present more efficient solutions for maximality, based on DFS traversal of a spatial tree. Third, we consider the common special case of point datasets, and present an even more efficient randomized algorithm. Fourth, we show that contiguous regions are tractable for a general version of maximality for which arbitrary regions are intractable. Fifth, we examine the structure of our integer programming formulation and show that for point datasets, our program is integral. Finally, we have implemented all techniques from this article in Google Maps [Google 2005] visualizations of fusion tables [Gonzalez et al. 2010], and we describe a set of experiments that demonstrate the trade-offs among the algorithms. Anish Das Sarma, Hongrae Lee, Hector Gonzalez, Jayant Madhavan, Alon Y. Halevy |
ACM Trans. Database Syst. | 2 |
| 2012 | Efficient Exact Similarity Searches Using Multiple Token OrderingsabstractSimilarity searches are essential in many applications including data cleaning and near duplicate detection. Many similarity search algorithms first generate candidate records, and then identify true matches among them. A major focus of those algorithms has been on how to reduce the number of candidate records in the early stage of similarity query processing. One of the most commonly used techniques to reduce the candidate size is the prefix filtering principle, which exploits the document frequency ordering of tokens. In this paper, we propose a novel partitioning technique that considers multiple token orderings based on token co-occurrence statistics. Experimental results show that the proposed technique is effective in reducing the number of candidate records and as a result improves the performance of existing algorithms significantly. Jongik Kim, Hongrae Lee |
ICDE | 2 |
| 2012 | CloudRAMSort: fast and efficient large-scale distributed RAM sort on shared-nothing clusterabstractSorting is a fundamental kernel used in many database operations. The total memory available across cloud computers is now sufficient to store even hundreds of terabytes of data in-memory. Applications requiring high-speed data analysis typically use in-memory sorting. The two most important factors in designing a high-speed in-memory sorting system are the single-node sorting performance and inter-node communication. Changkyu Kim, Jongsoo Park, Nadathur Satish, Hongrae Lee, Pradeep Dubey, Jatin Chhugani |
SIGMOD Conference | 4 |
| 2012 | Finding related tablesabstractWe consider the problem of finding related tables in a large corpus of heterogenous tables. Detecting related tables provides users a powerful tool for enhancing their tables with additional data and enables effective reuse of available public data. Our first contribution is a framework that captures several types of relatedness, including tables that are candidates for joins and tables that are candidates for union. Our second contribution is a set of algorithms for detecting related tables that can be either unioned or joined. We describe a set of experiments that demonstrate that our algorithms produce highly related tables. We also show that we can often improve the results of table search by pulling up tables that are ranked much lower based on their relatedness to top-ranked tables. Finally, we describe how to scale up our algorithms and show the results of running it on a corpus of over a million tables extracted from Wikipedia. Anish Das Sarma, Lujun Fang, Nitin Gupta 0003, Alon Y. Halevy, Hongrae Lee, Fei Wu 0003, Reynold Xin, Cong Yu 0001 |
SIGMOD Conference | 5 |
| 2012 | Efficient spatial sampling of large geographical tablesabstractLarge-scale map visualization systems play an increasingly important role in presenting geographic datasets to end users. Since these datasets can be extremely large, a map rendering system often needs to select a small fraction of the data to visualize them in a limited space. This paper addresses the fundamental challenge of thinning: determining appropriate samples of data to be shown on specific geographical regions and zoom levels. Other than the sheer scale of the data, the thinning problem is challenging because of a number of other reasons: (1) data can consist of complex geographical shapes, (2) rendering of data needs to satisfy certain constraints, such as data being preserved across zoom levels and adjacent regions, and (3) after satisfying the constraints, an optimal solution needs to be chosen based on objectives such as maximality, fairness, and importance of data. Anish Das Sarma, Hongrae Lee, Hector Gonzalez, Jayant Madhavan, Alon Y. Halevy |
SIGMOD Conference | 2 |
| 2011 | Similarity Join Size Estimation using Locality Sensitive HashingabstractSimilarity joins are important operations with a broad range of applications. In this paper, we study the problem of vector similarity join size estimation (VSJ). It is a generalization of the previously studied set similarity join size estimation (SSJ) problem and can handle more interesting cases such as TF-IDF vectors. One of the key challenges in similarity join size estimation is that the join size can change dramatically depending on the input similarity threshold. We propose a sampling based algorithm that uses Locality-Sensitive-Hashing (LSH). The proposed algorithm LSH-SS uses an LSH index to enable effective sampling even at high thresholds. We compare the proposed technique with random sampling and the state-of-the-art technique for SSJ (adapted to VSJ) and demonstrate LSH-SS offers more accurate estimates throughout the similarity threshold range and small variance using real-world data sets. Hongrae Lee, Raymond T. Ng, Kyuseok Shim |
Proc. VLDB Endow. | 1 |
| 2010 | Variance aware optimization of parameterized queriesabstractParameterized queries are commonly used in database applications. In a parameterized query, the same SQL statement is potentially executed multiple times with different parameter values. In today's DBMSs the query optimizer typically chooses a single execution plan that is reused for multiple instances of the same query. A key problem is that even if a plan with low average cost across instances is chosen, its variance can be high, which is undesirable in many production settings. In this paper, we describe techniques for selecting a plan that can better address the trade-off between the average and variance of cost across instances of a parameterized query. We show how to efficiently compute the skyline in the average-variance cost space. We have implemented our techniques on top of a commercial DBMS. We present experimental results on benchmark and real-world decision support queries. Surajit Chaudhuri, Hongrae Lee, Vivek R. Narasayya |
SIGMOD Conference | 2 |
| 2009 | Approximate substring selectivity estimationabstractWe study the problem of estimating selectivity of approximate substring queries. Its importance in databases is ever increasing as more and more data are input by users and are integrated with many typographical errors and different spelling conventions. To begin with, we consider edit distance for the similarity between a pair of strings. Based on information stored in an extended N-gram table, we propose two estimation algorithms, MOF and LBS for the task. The latter extends the former with ideas from set hashing signatures. The experimental results show that MOF is a light-weight algorithm that gives fairly accurate estimations. However, if more space is available, LBS can give better accuracy than MOF and other baseline methods. Next, we extend the proposed solution to other similarity predicates, SQL LIKE operator and Jaccard similarity. Hongrae Lee, Raymond T. Ng, Kyuseok Shim |
EDBT | 1 |
| 2009 | Power-Law Based Estimation of Set Similarity Join SizeabstractWe propose a novel technique for estimating the size of set similarity join. The proposed technique relies on a succinct representation of sets using Min-Hash signatures. We exploit frequent patterns in the signatures for the Set Similarity Join (SSJoin) size estimation by counting their support. However, there are overlaps among the counts of signature patterns and we need to use the set Inclusion-Exclusion (IE) principle. We develop a novel lattice-based counting method for efficiently evaluating the IE principle. The proposed counting technique is linear in the lattice size. To make the mining process very light-weight, we exploit a recently discovered Power-law relationship of pattern count and frequency. Extensive experimental evaluations show the proposed technique is capable of accurate and efficient estimation. Hongrae Lee, Raymond T. Ng, Kyuseok Shim |
Proc. VLDB Endow. | 1 |
| 2007 | Extending Q-Grams to Estimate Selectivity of String Matching with Low Edit Distance
Hongrae Lee, Raymond T. Ng, Kyuseok Shim |
VLDB | 1 |