EDBT 2026 Demo / reviewers in the wild / expert
Wang Lian
dblp:56/3112
· DBLP profile ↗
5ranked-venue papers
4as first author
0since 2021 · last 2005
0000-0003-0895-061XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 5 · 4 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
3 papers |
Data mining · 49% Indexing and storage engines · 19% Information retrieval · 19% |
Topics — the 9 heaviest of 10, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Data mining
pattern mining |
0.1 | 1 | 2005 | Indexing Useful Structural Patterns for XML Query Processing · IEEE Trans. Knowl. Data Eng. 2005 |
Query processing and optimization
XML query processing |
0.1 | 1 | 2005 | Indexing Useful Structural Patterns for XML Query Processing · IEEE Trans. Knowl. Data Eng. 2005 |
Data mining
clustering |
0.0 | 1 | 2004 | An Efficient and Scalable Algorithm for Clustering XML Documents by Structure · IEEE Trans. Knowl. Data Eng. 2004 |
Data mining › clustering
hierarchical clustering |
0.0 | 1 | 2004 | An Efficient and Scalable Algorithm for Clustering XML Documents by Structure · IEEE Trans. Knowl. Data Eng. 2004 |
Data mining › clustering › document clustering
XML document clustering |
0.0 | 1 | 2004 | An Efficient and Scalable Algorithm for Clustering XML Documents by Structure · IEEE Trans. Knowl. Data Eng. 2004 |
Indexing and storage engines
multidimensional indexing |
0.0 | 1 | 2003 | Similarity Search in Sets and Categorical Data Using the Signature Tree · ICDE 2003 |
Information retrieval › similarity search
set similarity search |
0.0 | 1 | 2003 | Similarity Search in Sets and Categorical Data Using the Signature Tree · ICDE 2003 |
Information retrieval
similarity search |
0.0 | 1 | 2003 | Similarity Search in Sets and Categorical Data Using the Signature Tree · ICDE 2003 |
Indexing and storage engines
vector index |
0.0 | 1 | 2003 | Similarity Search in Sets and Categorical Data Using the Signature Tree · ICDE 2003 |
Methods — techniques the papers use, named apart from their topics
filtering · 0.1data mining algorithms · 0.1structure graph · 0.0distance metric · 0.0signature tree · 0.0bitmap signatures · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2005 | Indexing Useful Structural Patterns for XML Query ProcessingabstractQueries on semistructured data are hard to process due to the complex nature of the data and call for specialized techniques. Existing path-based indexes and query processing algorithms are not efficient for searching complex structures beyond simple paths, even when the queries are high-selective. We introduce the definition of minimal infrequent structures (MIS), which are structures that 1) exist in the data, 2) are not frequent with respect to a support threshold, and 3) all substructures of them are frequent. By indexing the occurrences of MIS, we can efficiently locate the high-selective substructures of a query, improving search performance significantly. An efficient data mining algorithm is proposed, which finds the minimal infrequent structures. Their occurrences in the XML data are then indexed by a lightweight data structure and used as a fast filter step in query evaluation. We validate the efficiency and applicability of our methods through experimentation on both synthetic and real data. Wang Lian, Nikos Mamoulis, David Wai-Lok Cheung, Siu-Ming Yiu |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2004 | Discovering Minimal Infrequent Structures from XML Documents
Wang Lian, Nikos Mamoulis, David Wai-Lok Cheung, Siu-Ming Yiu |
WISE | 1 |
| 2004 | An Efficient and Scalable Algorithm for Clustering XML Documents by StructureabstractWith the standardization of XML as an information exchange language over the Internet, a huge amount of information is formatted in XML documents. In order to analyze this information efficiently, decomposing the XML documents and storing them in relational tables is a popular practice. However, query processing becomes expensive since, in many cases, an excessive number of joins is required to recover information from the fragmented data. If a collection consists of documents with different structures (for example, they come from different DTDs), mining clusters in the documents could alleviate the fragmentation problem. We propose a hierarchical algorithm (S-GRACE) for clustering XML documents based on structural information in the data. The notion of structure graph (s-graph) is proposed, supporting a computationally efficient distance metric defined between documents and sets of documents. This simple metric yields our new clustering algorithm which is efficient and effective, compared to other approaches based on tree-edit distance. Experiments on real data show that our algorithm can discover clusters not easily identified by manual inspection. Wang Lian, David Wai-Lok Cheung, Nikos Mamoulis, Siu-Ming Yiu |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2003 | Similarity Search in Sets and Categorical Data Using the Signature TreeabstractData mining applications analyze large collections of set data and high dimensional categorical data. Search on these data types is not restricted to the classic problems of mining association rules and classification, but similarity search is also a frequently applied operation. Access methods/or multidimensional numerical data are inappropriate for this problem and specialized indexes are needed. We propose a method that represents set data as bitmaps (signatures) and organizes them into a hierarchical index, suitable for similarity search and other related query types. In contrast to a previous technique, the signature tree is dynamic and does not rely on hardwired constants. Experiments with synthetic and real datasets show that it is robust to different data characteristics, scalable to the database size and efficient for various queries. Nikos Mamoulis, David Wai-Lok Cheung, Wang Lian |
ICDE | 3 |
| 2003 | A Filter Index for Complex Queries on Semi-structured Data
Wang Lian, Nikos Mamoulis, David Wai-Lok Cheung |
WAIM | 1 |