Nan Liu 0009

dblp:86/4643-9 · DBLP profile ↗
← Back
6ranked-venue papers
3as first author
0since 2021 · last 2009
0000-0003-3610-4883ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 4 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 2 first-authorArtificial intelligence and machine learning · 1 · 1 first-authorSecurity and privacy · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
2 papers
Information retrieval · 100%
Theoretical computer science
1 paper
Graph algorithms and graph theory · 100%

Topics — the 5 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Information retrieval
web search
0.112007
A link classification based approach to website topic hierarchy generation · WWW 2007
Information retrieval › similarity measure
document similarity
0.112006
Measuring similarity of semi-structured documents with context weights · SIGIR 2006
Information retrieval
retrieval models
0.112006
Measuring similarity of semi-structured documents with context weights · SIGIR 2006
Information retrieval › retrieval models
vector space model
0.112006
Measuring similarity of semi-structured documents with context weights · SIGIR 2006
Graph algorithms and graph theory › spanning tree
minimum spanning tree
0.012007
A link classification based approach to website topic hierarchy generation · WWW 2007

Methods — techniques the papers use, named apart from their topics

link classification · 0.1directed minimum spanning tree · 0.1shortest-path tree · 0.1shortest path tree · 0.1extended vector space model · 0.1
YearPublicationVenuePosition
2009 Keyphrase extraction for labeling a website topic hierarchy
abstract
Looking for web pages to identify useful information from a website is tedious and time consuming. Search engines are not always helpful due to the vocabulary difference between queries and web pages. Users may also have difficulty to accurately represent their information needs as queries at the beginning of exploration stage. A site map of website provides an outline of the overall structure of website. Without navigating through the website from the root page, users can easily identify the exact webpage to extract useful information to satisfy their information needs. However, site maps are not always available. In our previous work, we develop techniques to generate a website topic hierarchy. In this paper, we extend our work to extract keyphrases to label the web site topic hierarchy. The keyphrases serve in the purpose of summarizing the content so that users can efficiently browse through the site map to pin point the web page that provides the useful information they need. In the proposed keyphrase extraction, there are three major components. The first component is the candidate phrases identification. The second component computes the feature scores for summarization. The features include thematic and presentation features. The third component extracts the keyphrases by combining the feature scores. We have conducted an experiment and obtained promising result.
Nan Liu 0009, Christopher C. Yang
ICEC1
2009 Web site topic-hierarchy generation based on link structure
abstract
Abstract Navigating through hyperlinks within a Web site to look for information from one of its Web pages without the support of a site map can be inefficient and ineffective. Although the content of a Web site is usually organized with an inherent structure like a topic hierarchy, which is a directed tree rooted at a Web site's homepage whose vertices and edges correspond to Web pages and hyperlinks, such a topic hierarchy is not always available to the user. In this work, we studied the problem of automatic generation of Web sites' topic hierarchies. We modeled a Web site's link structure as a weighted directed graph and proposed methods for estimating edge weights based on eight types of features and three learning algorithms, namely decision trees, naïve Bayes classifiers, and logistic regression. Three graph algorithms, namely breadth‐first search, shortest‐path search, and directed minimum‐spanning tree, were adapted to generate the topic hierarchy based on the graph model. We have tested the model and algorithms on real Web sites. It is found that the directed minimum‐spanning tree algorithm with the decision tree as the weight learning algorithm achieves the highest performance with an average accuracy of 91.9%.
Christopher C. Yang, Nan Liu 0009
J. Assoc. Inf. Sci. Technol.2
2007 A link classification based approach to website topic hierarchy generation
abstract
Hierarchical models are commonly used to organize a Website's content. A Website's content structure can be represented by a topic hierarchy, a directed tree rooted at a Website's homepage in which the vertices and edges correspond to Web pages and hyperlinks. In this work, we propose a new method for constructing the topic hierarchy of a Website. We model the Website's link structure using weighted directed graph, in which the edge weights are computed using a classifier that predicts if an edge connects a pair of nodes representing a topic and a sub-topic. We then pose the problem of building the topic hierarchy as finding the shortest-path tree and directed minimum spanning tree in the weighted graph. We've done extensive experiments using real Websites and obtained very promising results.
Nan Liu 0009, Christopher C. Yang
WWW1
2006 Analyzing the Terrorist Social Networks with Visualization Tools
Christopher C. Yang, Nan Liu 0009, Marc Sageman
ISI2
2006 Measuring similarity of semi-structured documents with context weights
abstract
In this work, we study similarity measures for text-centric XML documents based on an extended vector space model, which considers both document content and structure. Experimental results based on a benchmark showed superior performance of the proposed measure over the baseline which ignores structural knowledge of XML documents.
Christopher C. Yang, Nan Liu 0009
SIGIR2
2005 Extracting a website's content structure from its link structure
abstract
Hierarchical models are commonly used to organize a Website's content. A Website's content structure can be represented by a topic hierarchy, a directed tree rooted at a Website's homepage in which the vertices and edges correspond to Web pages and hyperlinks. In this work, we propose an algorithm for extracting a Website's topic hierarchy from its link structure. The proposed algorithm consists of a construction stage and a refining stage, in which we analyze the semantic relationships between web pages based on link structure, web page content and directory structure. We've done extensive experiments using different Websites and obtained very promising results.
Nan Liu 0009, Christopher C. Yang
CIKM1