VLDB 2026 Research / reviewers in the wild / expert
Nan Liu 0009
dblp:86/4643-9
· DBLP profile ↗
6ranked-venue papers
3as first author
0since 2021 · last 2009
0000-0003-3610-4883ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 4 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 2 first-authorArtificial intelligence and machine learning · 1 · 1 first-authorSecurity and privacy · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
2 papers |
Information retrieval · 100% | |
| Theoretical computer science
1 paper |
Graph algorithms and graph theory · 100% |
Topics — the 5 heaviest of 5, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Information retrieval
web search |
0.1 | 1 | 2007 | A link classification based approach to website topic hierarchy generation · WWW 2007 |
Information retrieval › similarity measure
document similarity |
0.1 | 1 | 2006 | Measuring similarity of semi-structured documents with context weights · SIGIR 2006 |
Information retrieval
retrieval models |
0.1 | 1 | 2006 | Measuring similarity of semi-structured documents with context weights · SIGIR 2006 |
Information retrieval › retrieval models
vector space model |
0.1 | 1 | 2006 | Measuring similarity of semi-structured documents with context weights · SIGIR 2006 |
Graph algorithms and graph theory › spanning tree
minimum spanning tree |
0.0 | 1 | 2007 | A link classification based approach to website topic hierarchy generation · WWW 2007 |
Methods — techniques the papers use, named apart from their topics
link classification · 0.1directed minimum spanning tree · 0.1shortest-path tree · 0.1shortest path tree · 0.1extended vector space model · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2009 | Keyphrase extraction for labeling a website topic hierarchyabstractLooking for web pages to identify useful information from a website is tedious and time consuming. Search engines are not always helpful due to the vocabulary difference between queries and web pages. Users may also have difficulty to accurately represent their information needs as queries at the beginning of exploration stage. A site map of website provides an outline of the overall structure of website. Without navigating through the website from the root page, users can easily identify the exact webpage to extract useful information to satisfy their information needs. However, site maps are not always available. In our previous work, we develop techniques to generate a website topic hierarchy. In this paper, we extend our work to extract keyphrases to label the web site topic hierarchy. The keyphrases serve in the purpose of summarizing the content so that users can efficiently browse through the site map to pin point the web page that provides the useful information they need. In the proposed keyphrase extraction, there are three major components. The first component is the candidate phrases identification. The second component computes the feature scores for summarization. The features include thematic and presentation features. The third component extracts the keyphrases by combining the feature scores. We have conducted an experiment and obtained promising result. Nan Liu 0009, Christopher C. Yang |
ICEC | 1 |
| 2009 | Web site topic-hierarchy generation based on link structureabstractAbstract Navigating through hyperlinks within a Web site to look for information from one of its Web pages without the support of a site map can be inefficient and ineffective. Although the content of a Web site is usually organized with an inherent structure like a topic hierarchy, which is a directed tree rooted at a Web site's homepage whose vertices and edges correspond to Web pages and hyperlinks, such a topic hierarchy is not always available to the user. In this work, we studied the problem of automatic generation of Web sites' topic hierarchies. We modeled a Web site's link structure as a weighted directed graph and proposed methods for estimating edge weights based on eight types of features and three learning algorithms, namely decision trees, naïve Bayes classifiers, and logistic regression. Three graph algorithms, namely breadth‐first search, shortest‐path search, and directed minimum‐spanning tree, were adapted to generate the topic hierarchy based on the graph model. We have tested the model and algorithms on real Web sites. It is found that the directed minimum‐spanning tree algorithm with the decision tree as the weight learning algorithm achieves the highest performance with an average accuracy of 91.9%. Christopher C. Yang, Nan Liu 0009 |
J. Assoc. Inf. Sci. Technol. | 2 |
| 2007 | A link classification based approach to website topic hierarchy generationabstractHierarchical models are commonly used to organize a Website's content. A Website's content structure can be represented by a topic hierarchy, a directed tree rooted at a Website's homepage in which the vertices and edges correspond to Web pages and hyperlinks. In this work, we propose a new method for constructing the topic hierarchy of a Website. We model the Website's link structure using weighted directed graph, in which the edge weights are computed using a classifier that predicts if an edge connects a pair of nodes representing a topic and a sub-topic. We then pose the problem of building the topic hierarchy as finding the shortest-path tree and directed minimum spanning tree in the weighted graph. We've done extensive experiments using real Websites and obtained very promising results. Nan Liu 0009, Christopher C. Yang |
WWW | 1 |
| 2006 | Analyzing the Terrorist Social Networks with Visualization Tools
Christopher C. Yang, Nan Liu 0009, Marc Sageman |
ISI | 2 |
| 2006 | Measuring similarity of semi-structured documents with context weightsabstractIn this work, we study similarity measures for text-centric XML documents based on an extended vector space model, which considers both document content and structure. Experimental results based on a benchmark showed superior performance of the proposed measure over the baseline which ignores structural knowledge of XML documents. Christopher C. Yang, Nan Liu 0009 |
SIGIR | 2 |
| 2005 | Extracting a website's content structure from its link structureabstractHierarchical models are commonly used to organize a Website's content. A Website's content structure can be represented by a topic hierarchy, a directed tree rooted at a Website's homepage in which the vertices and edges correspond to Web pages and hyperlinks. In this work, we propose an algorithm for extracting a Website's topic hierarchy from its link structure. The proposed algorithm consists of a construction stage and a refining stage, in which we analyze the semantic relationships between web pages based on link structure, web page content and directory structure. We've done extensive experiments using different Websites and obtained very promising results. Nan Liu 0009, Christopher C. Yang |
CIKM | 1 |