Lizhu Zhou

dblp:z/LizhuZhou · also Li-Zhu Zhou · DBLP profile ↗
← Back
99ranked-venue papers
2as first author
1since 2021 · last 2023
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 79 · 2 first-authorArtificial intelligence and machine learning · 22 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 15Human-computer interaction and ubiquitous computing · 5Graphics, computer vision, multimedia, augmented reality and games · 2Systems, architecture and hardware · 1Software engineering, systems software and programming languages · 1Theory of computation · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
25 papers
Information retrieval · 25% Spatial and temporal data management · 21% Data mining · 18%
Human-computer interaction and pervasive computing
3 papers
Ubiquitous computing and smart environments · 37% Learning and educational technologies · 28% User interface design and tools · 20%
Computer graphics and multimedia
1 paper
Multimedia analysis and retrieval · 100%
Theoretical computer science
3 papers
Graph algorithms and graph theory · 63% Mathematical optimization · 37%

Topics — the 30 heaviest of 68, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Spatial and temporal data management › spatial query processing › nearest neighbor query
k-nearest neighbor query
0.212015
G-Tree: An Efficient and Scalable Index for Spatial Search on Road Networks · IEEE Trans. Knowl. Data Eng. 2015
Spatial and temporal data management › spatial indexing
road network indexing
0.212015
G-Tree: An Efficient and Scalable Index for Spatial Search on Road Networks · IEEE Trans. Knowl. Data Eng. 2015
Spatial and temporal data management
spatial indexing
0.212015
G-Tree: An Efficient and Scalable Index for Spatial Search on Road Networks · IEEE Trans. Knowl. Data Eng. 2015
Spatial and temporal data management
spatial query processing
0.212015
G-Tree: An Efficient and Scalable Index for Spatial Search on Road Networks · IEEE Trans. Knowl. Data Eng. 2015
Data mining
pattern mining
0.232008
Efficient mining of frequent sequence generators · WWW 2008
Mining Complex Time-Series Data by Learning Markovian Models · ICDM 2006
Incremental Mining of Frequent Query Patterns from XML Queries for Caching · ICDM 2006
Spatial and temporal data management › spatial query processing › spatial network query
direction-aware search
0.212013
TsingNUS: a location-based service system towards live city · SIGMOD Conference 2013
Information retrieval
keyword search
0.222008
An effective and versatile keyword search engine on heterogenous data sources · Proc. VLDB Endow. 2008
EASE: an effective 3-in-1 keyword search method for unstructured, semi-structured and structured data · SIGMOD Conference 2008
Information retrieval › web search
location-based search
0.212013
TsingNUS: a location-based service system towards live city · SIGMOD Conference 2013
Spatial and temporal data management
location-based services
0.212013
TsingNUS: a location-based service system towards live city · SIGMOD Conference 2013
Information retrieval › web search
real-time search
0.212013
TsingNUS: a location-based service system towards live city · SIGMOD Conference 2013
Recommender systems
collaborative filtering
0.222010
Personalizing Web Page Recommendation via Collaborative Filtering and Topic-Aware Markov Model · ICDM 2010
Similarity measure and instance selection for collaborative filtering · WWW 2003
Query processing and optimization
XML query processing
0.122008
Efficient vectorial operators for processing xml twig queries · WWW 2008
Incremental Mining of Frequent Query Patterns from XML Queries for Caching · ICDM 2006
Information retrieval
similarity search
0.112012
SEAL: Spatio-Textual Similarity Search · Proc. VLDB Endow. 2012
Multimedia analysis and retrieval › video classification
web video categorization
0.112012
Keyword-propagation-based information enriching and noise removal for web news videos · KDD 2012
Multimedia analysis and retrieval › video retrieval
web video retrieval
0.112012
Keyword-propagation-based information enriching and noise removal for web news videos · KDD 2012
Data mining › structured data mining
graph mining
0.122007
Out-of-core coherent closed quasi-clique mining from large dense graph databases · ACM Trans. Database Syst. 2007
Coherent closed quasi-clique discovery from large dense graph databases · KDD 2006
Data mining › structured data mining › graph mining › dense subgraph mining
quasi-clique mining
0.122007
Out-of-core coherent closed quasi-clique mining from large dense graph databases · ACM Trans. Database Syst. 2007
Coherent closed quasi-clique discovery from large dense graph databases · KDD 2006
Information retrieval
search engines
0.122013
Sailer: an effective search engine for unified retrieval of heterogeneous xml and web documents · WWW 2008
TsingNUS: a location-based service system towards live city · SIGMOD Conference 2013
Query processing and optimization › query execution
concurrent query execution
0.112011
KEMB: A Keyword-Based XML Message Broker · IEEE Trans. Knowl. Data Eng. 2011
Query processing and optimization
keyword query processing
0.112011
KEMB: A Keyword-Based XML Message Broker · IEEE Trans. Knowl. Data Eng. 2011
Information retrieval
query suggestion
0.112011
Interactive SQL query suggestion: Making databases user-friendly · ICDE 2011
Data models and query languages
SQL
0.112011
Interactive SQL query suggestion: Making databases user-friendly · ICDE 2011
Data stream processing › publish/subscribe
XML message brokering
0.112011
KEMB: A Keyword-Based XML Message Broker · IEEE Trans. Knowl. Data Eng. 2011
Query processing and optimization
interactive query processing
0.112010
Seaform: Search-As-You-Type in Forms · Proc. VLDB Endow. 2010
Information retrieval
search interfaces
0.112010
Seaform: Search-As-You-Type in Forms · Proc. VLDB Endow. 2010
Information retrieval
type-ahead search
0.112010
Seaform: Search-As-You-Type in Forms · Proc. VLDB Endow. 2010
Recommender systems › content recommendation
web page recommendation
0.112010
Personalizing Web Page Recommendation via Collaborative Filtering and Topic-Aware Markov Model · ICDM 2010
Data mining › structured data mining › graph mining
community detection
0.112009
Parallel community detection on large networks with propinquity dynamics · KDD 2009
Graph data management › graph similarity
graph edit distance
0.112009
Comparing Stars: On Approximating Graph Edit Distance · Proc. VLDB Endow. 2009
Graph data management
graph similarity
0.112009
Comparing Stars: On Approximating Graph Edit Distance · Proc. VLDB Endow. 2009

Methods — techniques the papers use, named apart from their topics

r-tree · 0.4assembly-based method · 0.4crowdsourcing · 0.3manifold learning · 0.3keyword propagation · 0.3pruning techniques · 0.2indexing · 0.2ranking · 0.2graph summarization · 0.2extended inverted index · 0.2template ranking · 0.1queryable template · 0.1progressive algorithm · 0.1upper and lower bounds · 0.1survey · 0.1prototype design · 0.1propinquity dynamics · 0.1polynomial-time approximation · 0.1
YearPublicationVenuePosition
2023 Pelive Floor Myofascisl Therapy is Associated with Improved VAS Pain Scores and FSFI Scores in Women with Dyspareunia 6 Months Post-partum
Huang Liu, Langchi Le, Ruihua Liu, Ren Jia, Liu Dan, Lizhu Zhou
Neural Process. Lett.7
2016 Welcome Message from DSE Managing Editor
abstract
an academic platform for sharing your achievements in data science and related research with scientists and engineers across the world.In recent years, the Chinese database community, represented by the China Computer Federation (CCF) Technical Committee on Databases (TCDB), has contributed a substantial number of publications to top conferences, including VLDB, SIGMOD, ICDE, WWW, and KDD.At the same time, collaboration between Chinese institutions and international partners in organizing international conferences and performing joint research has increased considerably.These are dramatic changes compared with the database research activities of the early years of this century.Similar trends have been observed within other technical committees of CCF, such as those focusing on the Internet, High-Performance Computing, and the Internet of Things.This progress in database research in China is continually being pushed forward by the worldwide development of the area of big data technologies and applications.Inspired by these exciting trends, we decided to develop a new journal so that scientists and engineers both inside and outside of China could share research results on the foundations and novel applications of big data technology.With this aim in mind, we formed
Lizhu Zhou
Data Sci. Eng.1
2015 G-Tree: An Efficient and Scalable Index for Spatial Search on Road Networks
abstract
In the recent decades, we have witnessed the rapidly growing popularity of location-based systems. Three types of location-based queries on road networks, single-pair shortest path query, k nearest neighbor (kNN) query, and keyword-based kNN query, are widely used in location-based systems. Inspired by R-tree, we propose a height-balanced and scalable index, namely G-tree, to efficiently support these queries. The space complexity of G-tree is O(|V|log|V|) where |V| is the number of vertices in the road network. Unlike previous works that support these queries separately, G-tree supports all these queries within one framework. The basis for this framework is an assembly-based method to calculate the shortest-path distances between two vertices. Based on the assembly-based method, efficient search algorithms to answer kNN queries and keyword-based kNN queries are developed. Experiment results show G-tree's theoretical and practical superiority over existing methods.
Ruicheng Zhong, Guoliang Li 0001, Kian-Lee Tan, Lizhu Zhou, Zhiguo Gong
IEEE Trans. Knowl. Data Eng.4
2013 G-tree: an efficient index for KNN search on road networks
abstract
In this paper we study the problem of kNN search on road networks. Given a query location and a set of candidate objects in a road network, the kNN search finds the k nearest objects to the query location. To address this problem, we propose a balanced search tree index, called G-tree. The G-tree of a road network is constructed by recursively partitioning the road network into sub-networks and each G-tree node corresponds to a sub-network. Inspired by classical kNN search on metric space, we introduce a best-first search algorithm on road networks, and propose an elaborately-designed assembly-based method to efficiently compute the minimum distance from a G-tree node to the query location. G-tree only takes O(|V|log|V|) space, where |V| is the number of vertices in a network, and thus can easily scale up to large road networks with more than 20 millions vertices. Experimental results on eight real-world datasets show that our method significantly outperforms state-of-the-art methods, even by 2-3 orders of magnitude.
Ruicheng Zhong, Guoliang Li 0001, Kian-Lee Tan, Lizhu Zhou
CIKM4
2013 Crowdsourcing-Assisted Query Structure Interpretation
Ju Fan, Lizhu Zhou
IJCAI3
2013 TsingNUS: a location-based service system towards live city
abstract
We present our system towards live city, called TsingNUS, aiming to provide users with more user-friendly location-aware search experiences. TsingNUS crawls location-based user-generated content from the Web (e.g., Foursquare and Twitter), cleans and integrates them to provide users with rich well-structured data. TsingNUS provides three user-friendly search paradigms: location-aware instant search, location-aware similarity search and direction-aware search. Instant search returns relevant answers instantly as users type in queries letter by letter, which can help users to save typing efforts significantly. Location-aware similarity search enables fuzzy matching between queries and the underlying data, which can tolerate typing errors. The two features boost the search performance and improve the experiences for mobile users who often misspell the keywords due to the limitation of the mobile phone's keyboard. In addition, users have direction-aware search requirements in many applications. For example, a driver on the highway wants to find the nearest gas station or restaurant. She has a search requirement that the answers should be in front of her driving direction. TsingNUS enables direction-aware search to address this problem and allows users to search in specific directions. Moreover, TsingNUS incorporates continuous search to efficiently support continuously moving queries in a client-server system which can reduce the number of queries submitted to the server and communication cost between the client and server. We have implemented and deployed a system which has been commonly used and widely accepted.
Guoliang Li 0001, Ruicheng Zhong, Weihuang Huang, Ju Fan, Kian-Lee Tan, Lizhu Zhou, Jianhua Feng
SIGMOD Conference8
2012 Location-aware instant search
abstract
Location-Based Services (LBS) have been widely accepted by mobile users recently. Existing LBS-based systems require users to type in complete keywords. However for mobile users it is rather difficult to type in complete keywords on mobile devices. To alleviate this problem, in this paper we study the location-aware instant search problem, which returns users location-aware answers as users type in queries letter by letter. The main challenge is to achieve high interactive speed. To address this challenge, in this paper we propose a novel index structure, prefix-region tree (called PR-Tree), to efficiently support location-aware instant search. PR-Tree is a tree-based index structure which seamlessly integrates the textual description and spatial information to index the spatial data. Using the PR-Tree, we develop efficient algorithms to support single prefix queries and multi-keyword queries. Experiments show that our method achieves high performance and significantly outperforms state-of-the-art methods.
Ruicheng Zhong, Ju Fan, Guoliang Li 0001, Kian-Lee Tan, Lizhu Zhou
CIKM5
2012 Keyword-propagation-based information enriching and noise removal for web news videos
abstract
The volume of Web videos have increased sharply through the past several years because of the evolvement of Web video sites.Enhanced algorithms on retrieval, classification and TDT (abbreviation of Topic Detection and Tracking) can bring lots of convenience to Web users as well as release tedious work from the administrators. Nevertheless, due to the the insufficiency of annotation keywords and the gap between video features and semantic concepts, it is still far away from satisfactory to implement them based on initial keywords and visual features. In this paper we utilize a keyword propagation algorithm based on manifold structure to enrich the keyword information and remove the noise for videos. Both text similarity and temporal similarity are employed to explore the relationship between any pair of videos and to construct the propagation model. We explore three applications, i.e., TDT, Retrieval and Classification based on a Web news video dataset obtained from a famous online video-distributing website, YouKu, and evaluate our approach. Experimental results demonstrate that they achieve satisfactory performance and always outperform the baseline methods.
Xiaoming Fan, Jianyong Wang 0001, Lizhu Zhou
KDD4
2012 Form-Based Instant Search and Query Autocompletion on Relational Data
Hao Wu 0010, Lizhu Zhou
WAIM2
2012 SEAL: Spatio-Textual Similarity Search
abstract
Location-based services (LBS) have become more and more ubiquitous recently. Existing methods focus on finding relevant points-of-interest (POIs) based on users' locations and query keywords. Nowadays, modern LBS applications generate a new kind of spatio-textual data, regions-of-interest (ROIs), containing region-based spatial information and textual description, e.g., mobile user profiles with active regions and interest tags. To satisfy search requirements on ROIs, we study a new research problem, called spatio-textual similarity search: Given a set of ROIs and a query ROI, we find the similar ROIs by considering spatial overlap and textual similarity. Spatio-textual similarity search has many important applications, e.g., social marketing in location-aware social networks. It calls for an efficient search method to support large scales of spatio-textual data in LBS systems. To this end, we introduce a filter-and-verification framework to compute the answers. In the filter step, we generate signatures for the ROIs and the query, and utilize the signatures to generate candidates whose signatures are similar to that of the query. In the verification step, we verify the candidates and identify the final answers. To achieve high performance, we generate effective high-quality signatures, and devise efficient filtering algorithms as well as pruning techniques. Experimental results on real and synthetic datasets show that our method achieves high performance.
Ju Fan, Guoliang Li 0001, Lizhu Zhou
Proc. VLDB Endow.3
2011 SDDB: A Self-Dependent and Data-Based Method for Constructing Bilingual Dictionary from the Web
Lizhu Zhou
APWeb2
2011 Measuring Similarity of Chinese Web Databases Based on Category Hierarchy
Ju Fan, Lizhu Zhou
APWeb3
2011 An Effective Approach for Searching Closest Sentence Translations from the Web
Ju Fan, Guoliang Li 0001, Lizhu Zhou
DASFAA (2)3
2011 Efficient Incremental Mining of Frequent Sequence Generators
Yukai He, Jianyong Wang 0001, Lizhu Zhou
DASFAA (1)3
2011 Interactive SQL query suggestion: Making databases user-friendly
abstract
SQL is a classical and powerful tool for querying relational databases. However, it is rather hard for inexperienced users to pose SQL queries, as they are required to be proficient in SQL syntax and have a thorough understanding of the underlying schema. To give users gratification, we propose SQLSUGG, an effective and user-friendly keyword-based method to help various users formulate SQL queries. SQLSUGG suggests SQL queries as users type in keywords, and can save users' typing efforts and help users avoid tedious SQL debugging. To achieve high suggestion effectiveness, we propose queryable templates to model the structures of SQL queries. We propose a template ranking model to suggest templates relevant to query keywords. We generate SQL queries from each suggested template based on the degree of matchings between keywords and attributes. For efficiency, we propose a progressive algorithm to compute top-k templates, and devise an efficient method to generate SQL queries from templates. We have implemented our methods on two real data sets, and the experimental results show that our method achieves high effectiveness and efficiency.
Ju Fan, Guoliang Li 0001, Lizhu Zhou
ICDE3
2011 An effective 3-in-1 keyword search method over heterogeneous data sources
Guoliang Li 0001, Jianhua Feng, Beng Chin Ooi, Jianyong Wang 0001, Lizhu Zhou
Inf. Syst.5
2011 KEMB: A Keyword-Based XML Message Broker
abstract
This paper studies the problem of XML message brokering with user subscribed profiles of keyword queries and presents a KEyword-based XML Message Broker (KEMB) to address this problem. In contrast to traditional-path-expressions-based XML message brokers, KEMB stores a large number of user profiles, in the form of keyword queries, which capture the data requirement of users/applications, as opposed to path expressions, such as XPath/XQuery expressions. KEMB brings new challenges: 1) how to effectively identify relevant answers of keyword queries in XML data streams; and 2) how to efficiently answer large numbers of concurrent keyword queries. We adopt compact lowest common ancestors (CLCAs) to effectively identify relevant answers. We devise an automaton-based method to process large numbers of queries and devise an effective optimization strategy to enhance performance and scalability. We have implemented and evaluated KEMB on various data sets. The experimental results show that KEMB achieves high performance and scales very well.
Guoliang Li 0001, Jianhua Feng, Jianyong Wang 0001, Lizhu Zhou
IEEE Trans. Knowl. Data Eng.4
2010 Suggesting Topic-Based Query Terms as You Type
abstract
Query term suggestion that interactively expands the queries is an indispensable technique to help users formulate high-quality queries and has attracted much attention in the community of web search. Existing methods usually suggest terms based on statistics in documents as well as query logs and external dictionaries, and they neglect the fact that the topic information is very crucial because it helps retrieve topically relevant documents. To give users gratification, we propose a novel term suggestion method: as the user types in queries letter by letter, we suggest the terms that are topically coherent with the query and could retrieve relevant documents instantly. For effectively suggesting highly relevant terms, we propose a generative model by incorporating the topical coherence of terms. The model learns the topics from the underlying documents based on Latent Dirichlet Allocation (LDA). For achieving the goal of instant query suggestion, we use a trie structure to index and access terms. We devise an efficient top-k algorithm to suggest terms as users type in queries. Experimental results show that our approach not only improves the effectiveness of term suggestion, but also achieves better efficiency and scalability.
Ju Fan, Hao Wu 0010, Guoliang Li 0001, Lizhu Zhou
APWeb4
2010 Personalizing Web Page Recommendation via Collaborative Filtering and Topic-Aware Markov Model
abstract
Web-page recommendation is to predict the next request of pages that Web users are potentially interested in when surfing the Web. This technique can guide Web users to find more useful pages without asking for them explicitly and has attracted much attention in the community of Web mining. However, few studies on Web page recommendation consider personalization, which is an indispensable feature to meet various preferences of users. In this paper, we propose a personalized Web page recommendation model called PIGEON (abbr. for PersonalIzed web paGe rEcommendatiON) via collaborative filtering and a topic-aware Markov model. We propose a graph-based iteration algorithm to discover users' interested topics, based on which user similarities are measured. To recommend topically coherent pages, we propose a topic-aware Markov model to learn users' navigation patterns which capture both temporal and topical relevance of pages. A thorough experimental evaluation conducted on a large real dataset demonstrates PIGEON's effectiveness and efficiency.
Qingyan Yang, Ju Fan, Jianyong Wang 0001, Lizhu Zhou
ICDM4
2010 Modeling Ontology of Folksonomy with Latent Semantics of Tags
abstract
Modeling ontology of folksonomy provides a way of learning light weight ontology's which is a hot topic investigated recently. Previous approaches for modeling ontology of folksonomy either ignores semantics (synonymy, hyponymy or polysemy) or do not simultaneously consider relationships between actors (users), concepts (tags) and instances(resources) or are based on the idea that title words are responsible for generating tags for resources. Latent semantics and user-tag dependencies instead of user-word dependencies however are extremely important. In this paper we address these problems by introducing a latent topic layer into the traditional tripartite Actor-Concept-Instance graph. We thus propose an Actor-Concept-Instance-Topic (ACIT) approach to model ontology from folksonomy in a unified way by directly using tags and users of resources. We illustrate on Bibsonomy dataset that our proposed approach ACIT outperforms title words based approaches Tag-Topic (TT) and (User-Word-Topic) UWT for modeling the ontology of folksonomy.
Ali Daud, Juan-Zi Li, Lizhu Zhou, Lei Zhang 0174, Ying Ding 0001, Faqir Muhammad
Web Intelligence3
2010 Knowledge discovery through directed probabilistic topic models: a survey
Ali Daud, Juan-Zi Li, Lizhu Zhou, Faqir Muhammad
Frontiers Comput. Sci. China3
2010 Finding and ranking compact connected trees for effective keyword proximity search in XML documents
Jianhua Feng, Guoliang Li 0001, Jianyong Wang 0001, Lizhu Zhou
Inf. Syst.4
2010 Modeling and Verifying Concurrent Programs with Finite Chu Spaces
Xutao Du, Chunxiao Xing, Lizhu Zhou
J. Comput. Sci. Technol.3
2010 Temporal expert finding through generalized time topic modeling
Ali Daud, Juan-Zi Li, Lizhu Zhou, Faqir Muhammad
Knowl. Based Syst.3
2010 Seaform: Search-As-You-Type in Forms
abstract
Form-style interfaces have been widely used to allow users to access information. In this demonstration paper, we develop a new search paradigm in form-style query interfaces, called Seaform (which stands for Search-As-You-Type in Forms), which computes answers on-the-fly as a user types in a query letter by letter and gives the user instant feedback. Seaform provides better user experiences compared with traditional form-based query systems by reducing the efforts for a user to compose a high-quality query to find relevant answers. Seaform can also enhance faceted search and allow users to on-the-fly explore the underlying data. This search paradigm requires high performance to achieve an interactive speed. We develop efficient techniques and use them to implement two systems on real datasets. We demonstrate the features of these systems.
Hao Wu 0010, Guoliang Li 0001, Chen Li 0001, Lizhu Zhou
Proc. VLDB Endow.4
2009 Exploiting Temporal Authors Interests via Temporal-Author-Topic Modeling
Ali Daud, Juan-Zi Li, Lizhu Zhou, Faqir Muhammad
ADMA3
2009 A Chu spaces semantics of control flow in BPEL
abstract
We present a Chu spaces semantics of typical control flow of BPEL including fault handling and link semantics. BPELcfis proposed as a simplification of this subset of BPEL. For the compositional modeling of BPEL, we present a Chu spaces process algebra consisting of seven operators. These operators allow faults to be thrown at any point of execution and take link-based synchronization into consideration. We present the abstract syntax of BPELcf, the semantic algebra, and the valuation functions for computing the Chu spaces denotation of BPELcfprograms. The valuation functions are straightforward because of the power of the Chu spaces process algebra.
Xutao Du, Chunxiao Xing, Lizhu Zhou
APSCC3
2009 Effective Fuzzy Keyword Search over Uncertain Data
Xiaoming Song, Guoliang Li 0001, Jianhua Feng, Lizhu Zhou
DASFAA4
2009 FOGGER: an algorithm for graph generator discovery
abstract
To our best knowledge, all existing graph pattern mining algorithms can only mine either closed, maximal or the complete set of frequent subgraphs instead of graph generators which are preferable to the closed subgraphs according to the Minimum Description Length principle in some applications. In this paper, we study a new problem of frequent subgraph mining, called frequent connected graph generator mining, which poses significant challenges due to the underlying complexity associated with frequent subgraph mining as well as the absence of Apriori property for graph generators. Whereas, we still present an efficient solution FOGGER for this new problem. By exploring some properties of graph generators, two effective pruning techniques, backward edge pruning and forward edge pruning, are proposed to prune the branches of the well-known DFS code enumeration tree that do not contain graph generators. To further improve the efficiency, an effective index structure, ADI++, is also devised to facilitate the subgraph isomorphism checking. We experimentally evaluate various aspects of FOGGER using both real and synthetic datasets. Our results demonstrate that the two pruning techniques are effective in pruning the unpromising parts of search space, and FOGGER is efficient and scalable in terms of the base size of input databases. Meanwhile, the performance study for graph generator-based classification model shows that generator-based model is much simpler and can achieve almost the same accuracy for classifying chemical compounds in comparison with closed subgraph-based model.
Zhiping Zeng, Jianyong Wang 0001, Lizhu Zhou
EDBT4
2009 Parallel community detection on large networks with propinquity dynamics
abstract
Graphs or networks can be used to model complex systems. Detecting community structures from large network data is a classic and challenging task. In this paper, we propose a novel community detection algorithm, which utilizes a dynamic process by contradicting the network topology and the topology-based propinquity, where the propinquity is a measure of the probability for a pair of nodes involved in a coherent community structure. Through several rounds of mutual reinforcement between topology and propinquity, the community structures are expected to naturally emerge. The overlapping vertices shared between communities can also be easily identified by an additional simple postprocessing. To achieve better efficiency, the propinquity is incrementally calculated. We implement the algorithm on a vertex-oriented bulk synchronous parallel(BSP) model so that the mining load can be distributed on thousands of machines. We obtained interesting experimental results on several real network data.
Jianyong Wang 0001, Yi Wang 0008, Lizhu Zhou
KDD4
2009 Conference Mining via Generalized Topic Modeling
Ali Daud, Juan-Zi Li, Lizhu Zhou, Faqir Muhammad
ECML/PKDD (1)3
2009 A gauss function based approach for unbalanced ontology matching
abstract
Ontology matching, aiming to obtain semantic correspondences between two ontologies, has played a key role in data exchange, data integration and metadata management. Among numerous matching scenarios, especially the applications cross multiple domains, we observe an important problem, denoted as unbalanced ontology matching which requires to find the matches between an ontology describing a local domain knowledge and another ontology covering the information over multiple domains, is not well studied in the community.
Qian Zhong, Juan-Zi Li, Guo Tong Xie, Jie Tang 0001, Lizhu Zhou
SIGMOD Conference6
2009 Discovering the staring people from social networks
abstract
In this paper, we study a novel problem of staring people discovery from social networks, which is concerned with finding people who are not only authoritative but also sociable in the social network. We formalize this problem as an optimization programming problem. Taking the co-author network as a case study, we define three objective functions and propose two methods to combine these objective functions. A genetic algorithm based method is further presented to solve this problem. Experimental results show that the proposed solution can effectively find the staring people from social networks.
Dewei Chen, Jie Tang 0001, Juan-Zi Li, Lizhu Zhou
WWW4
2009 Interactive search in XML data
abstract
In a traditional keyword-search system over XML data, a user composes a keyword query, submits it to the system, and retrieves relevant subtrees. In the case where the user has limited knowledge about the data, often the user feels "left in the dark" when issuing queries, and has to use a try-and-see approach for finding information. In this paper, we study a new information-access paradigm for XML data, called "Inks," in which the system searches on the underlying data "on the fly" as the user types in query keywords. Inks extends existing XML keyword search methods by interactively answering queries. We propose effective indices, early-termination techniques, and efficient search algorithms to achieve a high interactive speed. We have implemented our algorithm, and the experimental results show that our method achieves high search efficiency and result quality.
Guoliang Li 0001, Jianhua Feng, Lizhu Zhou
WWW3
2009 Incremental sequence-based frequent query pattern mining from XML queries
Guoliang Li 0001, Jianhua Feng, Jianyong Wang 0001, Lizhu Zhou
Data Min. Knowl. Discov.4
2009 CONTOUR: an efficient algorithm for discovering discriminating subsequences
Jianyong Wang 0001, Lizhu Zhou, George Karypis, Charu C. Aggarwal
Data Min. Knowl. Discov.3
2009 SAIL: Structure-aware indexing for effective and progressive top-k keyword search over XML documents
Guoliang Li 0001, Chen Li 0001, Jianhua Feng, Lizhu Zhou
Inf. Sci.4
2009 Self-Switching Classification Framework for Titled Documents
Lizhu Zhou
J. Comput. Sci. Technol.2
2009 Comparing Stars: On Approximating Graph Edit Distance
abstract
Graph data have become ubiquitous and manipulating them based on similarity is essential for many applications. Graph edit distance is one of the most widely accepted measures to determine similarities between graphs and has extensive applications in the fields of pattern recognition, computer vision etc. Unfortunately, the problem of graph edit distance computation is NP-Hard in general. Accordingly, in this paper we introduce three novel methods to compute the upper and lower bounds for the edit distance between two graphs in polynomial time. Applying these methods, two algorithms AppFull and AppSub are introduced to perform different kinds of graph search on graph databases. Comprehensive experimental studies are conducted on both real and synthetic datasets to examine various aspects of the methods for bounding graph edit distance. Result shows that these methods achieve good scalability in terms of both the number of graphs and the size of graphs. The effectiveness of these algorithms also confirms the usefulness of using our bounds in filtering and searching of graphs.
Zhiping Zeng, Anthony K. H. Tung, Jianyong Wang 0001, Jianhua Feng, Lizhu Zhou
Proc. VLDB Endow.5
2009 Learning in an Ambient Intelligent World: Enabling Technologies and Practices
abstract
The rapid evolution of information and communication technology opens a wide spectrum of opportunities to change our surroundings into an Ambient Intelligent (AmI) world. AmI is a vision of future information society, where people are surrounded by a digital environment that is sensitive to their needs, personalized to their requirements, anticipatory of their behavior, and responsive to their presence. It emphasizes on greater user friendliness, user empowerment, and more effective service support, with an aim to bring information and communication technology to everyone, every home, every business, and every school, thus improving the quality of human life. AmI unprecedentedly enhances learning experiences by endowing the users with the opportunities of learning in context, a breakthrough from the traditional education settings. In this survey paper, we examine some major characteristics of an AmI learning environment. To deliver a feasible and effective solution to ambient learning, we overview a few latest developed enabling technologies in context awareness and interactive learning. Associated practices are meanwhile reported. We also describe our experience in designing and implementing a smart class prototype, which allows teachers to simultaneously instruct both local and remote students in a context-aware and natural way.
Lizhu Zhou, Yuanchun Shi
IEEE Trans. Knowl. Data Eng.3
2008 High Confidence Fragment-Based Classification Rule Mining for Imbalanced HIV Data
Bing Lv, Jianyong Wang 0001, Lizhu Zhou
APWeb3
2008 Information Presentation on Mobile Devices: Techniques and Practices
Lin Qiao, Lizhu Zhou
APWeb3
2008 GHOST: an effective graph-based framework for name distinction
abstract
Name ambiguity stems from the fact that many people or objects share identical names. In this paper, we focus on investigating the problem in digital libraries to distinguish publications written by authors with identical names. We present an effective graph-based framework, GHOST (abbr. GrapH-based framewOrk for name diStincTion), to solve the problem systematically. We evaluated the framework on the real DBLP dataset, and the experimental results show that GHOST outperforms the state-of-the-art method.
Xiaoming Fan, Jianyong Wang 0001, Bing Lv, Lizhu Zhou
CIKM4
2008 Retune: Retrieving and Materializing Tuple Units for Effective Keyword Search over Relational Databases
Guoliang Li 0001, Jianhua Feng, Lizhu Zhou
ER3
2008 SESQ: A Model-Driven Method for Building Object Level Vertical Search Engines
Yukai He, Ju Fan, Lizhu Zhou
ER5
2008 Abstract Reachability Graph for Verifying Web Service Interfaces
Xutao Du, Chunxiao Xing, Lizhu Zhou
ICSR3
2008 Parallel mining of closed quasi-cliques
abstract
Graph structure can model the relationships among a set of objects. Mining quasi-clique patterns from large dense graph data makes sense with respect to both statistic and applications. The applications of frequent quasi-cliques include stock price correlation discovery, gene function prediction and protein molecular analysis. Although the graph mining community has devised many skills to accelerate the discovery process, mining time is always unacceptable, especially on large dense graph data with low support threshold. Therefore, parallel algorithms are desirable on mining quasi-clique patterns. Message passing is one of the most widely used parallel framework. In this paper, we parallelize the state-of-the-art closed quasi-clique mining algorithm called Cocain using message passing. The parallelized version of Cocain can achieve 30+ fold speedup on 32 processors in a cluster of SMPs. The techniques proposed in this work can be applied to parallelize other pattern-growth based frequent pattern mining algorithms.
Jianyong Wang 0001, Zhiping Zeng, Lizhu Zhou
IPDPS4
2008 Efficient Mining of Minimal Distinguishing Subgraph Patterns from Graph Databases
Zhiping Zeng, Jianyong Wang 0001, Lizhu Zhou
PAKDD3
2008 EASE: an effective 3-in-1 keyword search method for unstructured, semi-structured and structured data
abstract
Conventional keyword search engines are restricted to a given data model and cannot easily adapt to unstructured, semi-structured or structured data. In this paper, we propose an efficient and adaptive keyword search method, called EASE, for indexing and querying large collections of heterogenous data. To achieve high efficiency in processing keyword queries, we first model unstructured, semi-structured and structured data as graphs, and then summarize the graphs and construct graph indices instead of using traditional inverted indices. We propose an extended inverted index to facilitate keyword-based search, and present a novel ranking mechanism for enhancing search effectiveness. We have conducted an extensive experimental study using real datasets, and the results show that EASE achieves both high search efficiency and high accuracy, and outperforms the existing approaches significantly.
Guoliang Li 0001, Beng Chin Ooi, Jianhua Feng, Jianyong Wang 0001, Lizhu Zhou
SIGMOD Conference5
2008 Efficient Similarity Search for Tree-Structured Data
Guoliang Li 0001, Xuhui Liu, Jianhua Feng, Lizhu Zhou
SSDBM4
2008 Self-Switching Classification Framework for Titled Documents
abstract
Ambiguous words refer to words that have different meanings such as apple, window, etc. In text classification they are usually removed by feature reduction methods like information gain. Sometimes there are too many ambiguous words in the corpus that we cannot simply throw them away, especially when classifying documents from the Web. In this paper we look for a method to classify titled documents with the help of ambiguous words. Titled documents are a kind of documents that have a simple structure containing a title and an excerpt. News, messages, and paper abstracts with titles are such examples. Instead of introducing another feature reduction method, we describe a framework to make the best of ambiguous words in the titled documents. The framework improves the performance of traditional bag-of-words classifier with the help of a bag-of-word-pairs classifier. We implement the framework using one of the most popular classifiers, multinomial naive Bayes (MNB), as a case in point. The experiments with three real life datasets show that in our framework the MNB model performs much better than traditional MNB classifier and the naive weighting algorithm, which simply puts more weight on the title words.
Lizhu Zhou
WAIM2
2008 Effective Indices for Efficient Approximate String Search and Similarity Join
abstract
Data collections often have inconsistencies that arise due to a variety of reasons, and it is desirable to be able to identify and resolve them efficiently. Similarity queries are commonly used in data cleaning for matching similar data. In this work we concentrate on the following problem of approximate string matching based on edit distance: from a collection of strings, how to find those strings similar to a given string, or the strings in another collection of strings with similarity greater than some threshold? We propose an NFA-based (nondeterministic finite-state automation) method for effective approximate string search. We model strings as a trie and construct an NFA on top of the trie. We identify the similar strings by running the NFA based on the tree automata theory. Moreover, we propose grouped trie to further improve the performance of similarity search by incorporating some effective pruning techniques. We have implemented our method and the experimental results show that our approach achieves high performance and out performs the existing state-of-the-art methods by orders of magnitude.
Xuhui Liu, Guoliang Li 0001, Jianhua Feng, Lizhu Zhou
WAIM4
2008 Path Similarity Based Directory Ontology Matching
abstract
In a directory ontology, only concepts and hypernym/hyponym relationships between them are defined. Directory ontologies are widely used in many real-world applications such as product catalogues and Web directories. The widely used heterogeneous directories also bring a big challenge for integration of them. Previously, little attentions have been paid attention to the problem in the research community. In this paper, we propose a path similarity based approach for directory ontology matching. We define the concept's local label and the concept's path label. Next, we propose a path similarity method, which combines the local label and path label information. Then, a top-down similarity flooding is utilized to improve the matching result. Our experimental results on the music genre ontologies show that the proposed method can achieve a precision of 78% and a recall of 66%, significantly outperforming the baseline method.
Qian Zhong, Juan-Zi Li, Jie Tang 0001, Lizhu Zhou
WAIM5
2008 Efficient mining of frequent sequence generators
abstract
Sequential pattern mining has raised great interest in data mining research field in recent years. However, to our best knowledge, no existing work studies the problem of frequent sequence generator mining. In this paper we present a novel algorithm, FEAT (abbr. Frequent sEquence generATor miner), to perform this task. Ex-perimental results show that FEAT is more efficient than traditional sequential pattern mining algorithms but generates more concise re-sult set, and is very effective for classifying Web product reviews.
Chuancong Gao, Jianyong Wang 0001, Yukai He, Lizhu Zhou
WWW4
2008 Efficient vectorial operators for processing xml twig queries
abstract
This paper proposes several vectorial operators for processing XML twig queries, which are easy to be performed and inherently efficient for both Ancestor-Descendant (A-D) and Parent-Child (P-C) relationships. We develop optimizations on the vectorial operators to improve the efficiency of answering twig queries in holistic. We propose an algorithm to answer GTP queries based on our vectorial operators.
Guoliang Li 0001, Jianhua Feng, Jianyong Wang 0001, Lizhu Zhou
WWW5
2008 Sailer: an effective search engine for unified retrieval of heterogeneous xml and web documents
abstract
This paper studies the problem of unified ranked retrieval of heterogeneous XML documents and Web data. We propose an effective search engine called Sailer to adaptively and versatilely answer keyword queries over the heterogenous data. We model the Web pages and XML documents as graphs. We propose the concept of pivotal trees to effectively answer keyword queries and present an effective method to identify the top-k pivotal trees with the highest ranks from the graphs. Moreover, we propose effective indexes to facilitate the effective unified ranked retrieval. We have conducted an extensive experimental study using real datasets, and the experimental results show that Sailer achieves both high search efficiency and accuracy, and outperforms the existing approaches significantly.
Guoliang Li 0001, Jianhua Feng, Jianyong Wang 0001, Xiaoming Song, Lizhu Zhou
WWW5
2008 An effective and versatile keyword search engine on heterogenous data sources
abstract
We present EASE, an effective and versatile keyword search engine that enables users to easily access the heterogenous data composed of unstructured, semi-structured and structured data, without the need of learning XPath/XQuery or SQL languages. EASE addresses a challenge in keyword search that has been neglected in the literature: how to efficiently and adaptively process keyword queries on the heterogenous data. To provide such capability, EASE models unstructured, semi-structured and structured data as graphs, summarizes the graphs, and constructs graph indices instead of using traditional inverted indices for effective keyword search. EASE adopts an extended inverted index to facilitate keyword-based search, and employs a novel ranking mechanism for enhancing search effectiveness.
Guoliang Li 0001, Jianhua Feng, Jianyong Wang 0001, Lizhu Zhou
Proc. VLDB Endow.4
2007 A Framework for Titled Document Categorization with Modified Multinomial Naivebayes Classifier
Lizhu Zhou
ADMA2
2007 Effective keyword search for valuable lcas over xml documents
abstract
In this paper, we study the problem of effective keyword search over XML documents. We begin by introducing the notion of Valuable Lowest Common Ancestor (VLCA) to accurately and effectively answer keyword queries over XML documents. We then propose the concept of Compact VLCA (CVLCA) and compute the meaningful compact connected trees rooted as CVLCAs as the answers of keyword queries. To efficiently compute CVLCAs, we devise an effective optimization strategy for speeding up the computation, and exploit the key properties of CVLCA in the design of the stack-based algorithm for answering keyword queries. We have conducted an extensive experimental study and the experimental results show that our proposed approach achieves both high efficiency and effectiveness when compared with existing proposals.
Guoliang Li 0001, Jianhua Feng, Jianyong Wang 0001, Lizhu Zhou
CIKM4
2007 Efficient Holistic Twig Joins in Leaf-to-Root Combining with Root-to-Leaf Way
Guoliang Li 0001, Jianhua Feng, Yong Zhang 0002, Lizhu Zhou
DASFAA4
2007 Schema Mapping in P2P Networks Based on Classification and Probing
Guoliang Li 0001, Beng Chin Ooi, Bei Yu 0003, Lizhu Zhou
DASFAA4
2007 Classifying and Ranking: The First Step Towards Mining Inside Vertical Search Engines
Lizhu Zhou
DEXA3
2007 Mining Naturally Smooth Evolution of Clusters from Dynamic Data
abstract
Many clustering algorithms have been proposed to partition a set of static data points into groups. In this paper, we consider an evolutionary clustering problem where the input data points may move, disappeare, and emerge. Generally, these changes should result in a smooth evolution of the clusters. Mining this naturally smooth evolution is valuable for providing an aggregated view of the numerous individual behaviors. We solve this novel and generalized form of clustering problem by converting it into a Bayesian learning problem. Analogous to that the EM clustering algorithm clusters static data points by learning a Gaussian mixture model, our method mines the evolution of clusters from dynamic data points by learning a hidden semi-Markov model (HSMM). By utilizing characteristics of the evolutionary clustering problem, we derive a new unsupervised learning algorithm which is much more efficient than the algorithms used to learn traditional variable-duration HSMMs. Because the HSMM models the probabilistic relationship between the dynamic data set and corresponding evolving clusters, we can interpret the learned parameters as the evolving clusters intuitively using the Viterbi filtering technique. Because learning an HSMM is in fact learning an optimal Viterbi filter, the learned cluster evolutions are smooth and fit well with the data. We demonstrate the effectiveness of this method by experiments on both synthetic data and real data.
Yi Wang 0008, Shi-Xia Liu, Jianhua Feng, Lizhu Zhou
SDM4
2007 Discriminating Subsequence Discovery for Sequence Clustering
abstract
In this paper, we explore the discriminating subsequence-based clustering problem. First, several effective optimization techniques are proposed to accelerate the sequence mining process and a new algorithm, CONTOUR, is developed to efficiently and directly mine a subset of discriminating frequent subsequences which can be used to cluster the input sequences. Second, an accurate hierarchical clustering algorithm, SSC, is constructed based on the result of CONTOUR. The performance study evaluates the efficiency and scalability of CONTOUR, and the clustering quality of SSC.
Jianyong Wang 0001, Lizhu Zhou, George Karypis, Charu C. Aggarwal
SDM3
2007 Leveraging Webpage Classification for Data Object Recognition
abstract
Data-rich webpages are providing an increasingly important data source for web applications. While the problem of data object recognition is intensively discussed, it is mostly addressed as a separated process from the frontier task of relevant webpage identification. In this paper, we propose a method to leverage the classification result of data-rich webpages for efficient and scalable data object recognition. A novel context information is proposed, which can be inferred from the webpage classification and exploited in the bottom-up data object recognition. Experimental results show that the context information brings a 19% improvement in the running efficiency of the bottom- up data object recognition.
Lizhu Zhou
Web Intelligence2
2007 Efficient Mining of Frequent Closed XML Query Pattern
Jianhua Feng, Jianyong Wang 0001, Lizhu Zhou
J. Comput. Sci. Technol.4
2007 Out-of-core coherent closed quasi-clique mining from large dense graph databases
abstract
Due to the ability of graphs to represent more generic and more complicated relationships among different objects, graph mining has played a significant role in data mining, attracting increasing attention in the data mining community. In addition, frequent coherent subgraphs can provide valuable knowledge about the underlying internal structure of a graph database, and mining frequently occurring coherent subgraphs from large dense graph databases has witnessed several applications and received considerable attention in the graph mining community recently. In this article, we study how to efficiently mine the complete set of coherent closed quasi-cliques from large dense graph databases, which is an especially challenging task due to the fact that the downward-closure property no longer holds. By fully exploring some properties of quasi-cliques, we propose several novel optimization techniques which can prune the unpromising and redundant subsearch spaces effectively. Meanwhile, we devise an efficient closure checking scheme to facilitate the discovery of closed quasi-cliques only. Since large databases cannot be held in main memory, we also design an out-of-core solution with efficient index structures for mining coherent closed quasi-cliques from large dense graph databases. We call this Cocain*. Thorough performance study shows that Cocain* is very efficient and scalable for large dense graph databases.
Zhiping Zeng, Jianyong Wang 0001, Lizhu Zhou, George Karypis
ACM Trans. Database Syst.3
2006 SESQ: A Novel System for Building Domain Specific Web Search Engines
Lizhu Zhou
APWeb2
2006 Hidden Conditioned Homomorphism for XPath Fragment Containment
Yuguo Liao, Jianhua Feng, Yong Zhang 0002, Lizhu Zhou
DASFAA4
2006 Exploit Sequencing to Accelerate XML Twig Query Answering
Jianhua Feng, Jianyong Wang 0001, Lizhu Zhou
DASFAA4
2006 Segmented Document Classification: Problem and Solution
Lizhu Zhou
DEXA2
2006 CLAN: An Algorithm for Mining Closed Cliques from Large Dense Graph Databases
abstract
Most previously proposed frequent graph mining algorithms are intended to find the complete set of all frequent, closed subgraphs. However, in many cases only a subset of the frequent subgraphs with a certain topology is of special interest. Thus, the method of mining the complete set of all frequent subgraphs is not suitable for mining these frequent subgraphs of special interest as it wastes considerable computing power and space on uninteresting subgraphs. In this paper we develop a new algorithm, CLAN, to mine the frequent closed cliques, the most coherent structures in the graph setting. By exploring some properties of the clique pattern, we can simplify the canonical label design and the corresponding clique (or subclique) isomorphism testing. Several effective pruning methods are proposed to prune the search space, while the clique closure checking scheme is used to remove the non-closed clique patterns. Our empirical results show that CLAN is very efficient for large dense graph databases with which the traditional graph mining algorithms fail. The novelty of our method is further demonstrated by the application of CLAN in mining highly correlated stocks from large stock market data.
Jianyong Wang 0001, Zhiping Zeng, Lizhu Zhou
ICDE3
2006 Incremental Mining of Frequent Query Patterns from XML Queries for Caching
abstract
Existing studies for mining frequent XML query patterns mainly introduce a straightforward candidate generate-and-test strategy and compute frequencies of candidate query patterns from scratch periodically by checking the entire transaction database, which consists of XML query patterns transformed from user queries. However, it is nontrivial to maintain such discovered frequent patterns in real XML databases because there may incur frequent updates that may not only invalidate some existing frequent query patterns but also generate some new frequent ones. Accordingly, existing proposals are inefficient for the evolution of the transaction database. To address these problems, this paper presents an efficient algorithm IPS-FXQPMiner for mining frequent XML query patterns without candidate maintenance and costly tree-containment checking. We transform XML queries into sequences through a one- to-one mapping and then mine the frequent sequences to generate frequent XML query patterns. More importantly, based on IPS-FXQPMiner, an efficient incremental algorithm, Incre-FXQPMiner is proposed to incrementally mine frequent XML query patterns, which can minimize the I/O and computation requirements for handling incremental updates. Our experimental study on various real-life datasets demonstrates the efficiency and scalability of our algorithms over previous known alternatives.
Guoliang Li 0001, Jianhua Feng, Jianyong Wang 0001, Yong Zhang 0002, Lizhu Zhou
ICDM5
2006 Mining Complex Time-Series Data by Learning Markovian Models
abstract
In this paper, we propose a novel and general approach for time-series data mining. As an alternative to traditional ways of designing specific algorithm to mine certain kind of pattern directly from the data, our approach extracts the temporal structure of the time-series data by learning Markovian models, and then uses well established methods to efficiently mine a wide variety of patterns from the topology graph of the learned models. We consolidate the approach by explaining the use of some well-known Markovian models on mining several kinds of patterns. We then present a novel high-order hidden Markov model, the variable-length hidden Markov model (VLHMM), which combines the advantages of well- known Markovian models and has the superiority in both efficiency and accuracy. Therefore, it can mine a much wider variety of patterns than each of prior Markovian models. We demonstrate the power of VLHMM by mining four kinds of interesting patterns from 3D motion capture data, which is typical for the high-dimensionality and complex dynamics.
Yi Wang 0008, Lizhu Zhou, Jianhua Feng, Jianyong Wang 0001
ICDM2
2006 Real-Time Synthesis of 3D Animations by Learning Parametric Gaussians Using Self-Organizing Mixture Networks
Yi Wang 0008, Hujun Yin, Lizhu Zhou
ICONIP (2)3
2006 Coherent closed quasi-clique discovery from large dense graph databases
abstract
Frequent coherent subgraphs can provide valuable knowledge about the underlying internal structure of a graph database, and mining frequently occurring coherent subgraphs from large dense graph databases has been witnessed several applications and received considerable attention in the graph mining community recently. In this paper, we study how to efficiently mine the complete set of coherent closed quasi-cliques from large dense graph databases, which is an especially challenging task due to the downward-closure property no longer holds. By fully exploring some properties of quasi-cliques, we propose several novel optimization techniques, which can prune the unpromising and redundant sub-search spaces effectively. Meanwhile, we devise an efficient closure checking scheme to facilitate the discovery of only closed quasi-cliques. We also develop a coherent closed quasi-clique mining algorithm, Cocain1 Thorough performance study shows that Cocain is very efficient and scalable for large dense graph databases.
Zhiping Zeng, Jianyong Wang 0001, Lizhu Zhou, George Karypis
KDD3
2006 Learning Style-directed Dynamics of Human Motion for Automatic Motion Synthesis
abstract
This paper presents a new model, the HMM/Mix-SDTG, which describes Markov processes under control of a global vector variable called style variable. We present an EM learning algorithm to learn an HMM/Mix-SDTG from one or more 3D motion capture sequences labelled by their style values. Because each dimension of the style variable has explicit physical meaning, with the presented synthesis algorithm, we are able to generate arbitrarily new motion with style exactly as demand by specifying a style value. The output densities of HMM/Mix-SDTG is represented by mixtures of stylized decomposable triangulated graphs (Mix-SDTG), which, in addition to parameterizing the Markov process with the style variable, also achieve more numerical robustness and preventing common artifacts of 3D motion synthesis.
Yi Wang 0008, Lizhu Zhou
SMC3
2006 Supervised Learning of Motion Style for Real-time Synthesis of 3D Character Animations
abstract
In this paper, we present a supervised learning framework to learn a probabilistic mapping from values of a low-dimensionalstyle variable, which defines the characteristics of a certain kind of 3D human motion such as walking or boxing, to high-dimensional vectors defining 3D poses. All possible values of the style variable span an Euclidean space called style space. The supervised learning framework guarantees that each dimension of style space corresponds to a certain aspect of the motion characteristics, such as body height and pace length, so the user can precisely define a 3D pose by locating a point in the style space. Moreover, every curve in the Euclidean style space corresponds to a smooth motion sequence. We developed a graphical user interface program, with which, users simply points mouse cursor in the style space to define a 3D pose and drags mouse cursor to synthesis 3D animations in real-time.
Yi Wang 0008, Lei Xie 0001, Lizhu Zhou
SMC4
2006 The SOMN-HMM Model and Its Application to Automatic Synthesis of 3D Character Animations
abstract
Learning HMM from motion capture data for automatic 3D character animation synthesis is becoming a hot spot in research areas of computer graphics and machine learning. To ensure realistic synthesis, the model must be learned to fit the real distribution of human motion. Usually the fitness is measured by likelihood. In this paper, we present a new HMM learning algorithm, which incorporates stochastic optimization technique within the expectation-maximization (EM) learning framework. This algorithm is less prone to be trapped in local optimal and converges faster than traditional Baum-Welch learning algorithm. We apply the new algorithm to learning 3D motion under control of a style variable, which encodes the mood or personality of the performer. Given new style value, motions with corresponding style can be generated from the learned model.
Yi Wang 0008, Lei Xie 0001, Lizhu Zhou
SMC4
2006 Counting Graph Matches with Adaptive Statistics Collection
Jianhua Feng, Yuguo Liao, Lizhu Zhou
WAIM4
2006 SCEND: An Efficient Semantic Cache to Adequately Explore Answerability of Views
Guoliang Li 0001, Jianhua Feng, Na Ta 0001, Yong Zhang 0002, Lizhu Zhou
WISE5
2006 Meta-search Based Web Resource Discovery for Object-Level Vertical Search
Lizhu Zhou
WISE3
2006 2D/3D Web Visualization on Mobile Devices
Yi Wang 0008, Lizhu Zhou, Jianhua Feng, Lei Xie 0001, Chun Yuan 0003
WISE2
2006 Building a Domain Independent Platform for Collecting Domain Specific Data from the Web
Lizhu Zhou
WISE1
2006 Key-styling: learning motion style for real-time synthesis of 3D animation
abstract
Abstract In this paper, we present a novel real‐time motion synthesis approach that can generate 3D character animation with required style. The effectiveness of our approach comes from learning captured 3D human motion as a self‐organizing mixture network (SOMN); of parametric Gaussians.The learned model describes the motion under the control of a vector variable calledstyle variable, and acts as a probabilistic mapping from the low‐dimensional style values to the high‐dimensional 3D poses. We design a pose synthesis algorithm to allow the user to generate poses by specifying new style values. We also propose a novel motion synthesis method, the key‐styling, which accepts a sparse sequence of key style values and interpolates a dense sequence of style values to synthesize an animation. Key‐styling is able to produce animations that are more realistic and natural‐looking than those synthesized with the traditional key‐keyframing technique. Copyright © 2006 John Wiley & Sons, Ltd.
Yi Wang 0008, Lizhu Zhou
Comput. Animat. Virtual Worlds3
2005 BBTC: A New Update-Supporting Coding Scheme for XML Documents
Jianhua Feng, Guoliang Li 0001, Lizhu Zhou, Na Ta 0001, Yuguo Liao
WAIM3
2005 Extracting, Presenting and Browsing of Web Social Information
Yi Wang 0008, Lizhu Zhou
WAIM2
2005 Comparisons and computation of well-founded semantics for disjunctive logic programs
abstract
Much work has been done on extending the well-founded semantics to general disjunctive logic programs and various approaches have been proposed. However, these semantics are different from each other and no consensus is reached about which semantics is the most intended. In this article, we look at disjunctive well-founded reasoning from different angles. We show that there is an intuitive form of the well-founded reasoning in disjunctive logic programming which can be characterized by slightly modifying some existing approaches to defining disjunctive well-founded semantics, including program transformations, argumentation, unfounded sets (and resolution-like procedure). By employing the techniques developed by Brass and Dix in their transformation-based approach, we also provide a bottom-up procedure for this semantics. The significance of our work is not only in clarifying the relationship among different approaches, but also shed some light on what is an intended well-founded semantics for disjunctive logic programs.
Kewen Wang 0001, Lizhu Zhou
ACM Trans. Comput. Log.2
2004 A Highly Adaptable Web Information Extractor Using Graph Data Model
Lizhu Zhou, Jianhua Feng
APWeb2
2004 METS-Based Cataloging Toolkit for Digital Library Management System
Li Dong 0001, Bei Zhang 0001, Chunxiao Xing, Lizhu Zhou
Dublin Core Conference4
2004 A Cache Replacement Algorithm in Hierarchical Storage of Continuous Media Object
Yaoqiang Xu, Chunxiao Xing, Lizhu Zhou
WAIM3
2003 An Ontology-Based Method for Querying the Web Data
abstract
Nowadays large collections of documents of various types of information are available on the Internet, and finding relevant information on the World Wide Web is becoming a critical issue. Available search engines mostly rely on information retrieval techniques in finding information. They lack the desired power in query formulation which is beyond the keywords search. We introduce a method for querying the Web data that integrates both search engine and database techniques to support more complex search requests and get more precise results.
Chunxiao Xing, Lizhu Zhou, Jianhua Feng
AINA3
2003 A Hybrid Method for Web Data Extraction
abstract
Web data extraction refers to the technology that helps people find wanted information from the Web. We first classify existing data extraction algorithms into two classes: top-down and bottom-up, and then analyze their strengths and weaknesses in terms of extraction accuracy. On the basis of this analysis, we present a hybrid algorithm: bi-direction data extraction (BiDDE for short), which takes the full strengths of both top-down and bottom-up algorithms and yet avoid their weaknesses. The experimental results show that BiDDE has not only higher accuracy than top-down algorithm and bottom-up algorithm, but satisfactory performance.
Yu Wang 0011, Lizhu Zhou
Web Intelligence2
2003 Similarity measure and instance selection for collaborative filtering
abstract
Collaborative filtering has been very successful in both research and applications such as information filtering and E-commerce. The k-Nearest Neighbor (KNN) method is a popular way for its realization. Its key technique is to find k nearest neighbors for a given user to predict his interests. However, this method suffers from two fundamental problems: sparsity and scalability. In this paper, we present our solutions for these two problems. We adopt two techniques: a matrix conversion method for similarity measure and an instance selection method. And then we present an improved collaborative filtering algorithm based on these two methods. In contrast with existing collaborative algorithms, our method shows its satisfactory accuracy and performance.
Chun Zeng, Chunxiao Xing, Lizhu Zhou
WWW3
2002 WebME-Web mining environment
abstract
WebME is a Web mining environment (system) developed by Tsinghua University. It integrates many functions including: customization-based gathering, rough and subtle filtering, storing, indexing, recognizing, schema and data extracting, classifying, clustering and summarizing of Web pages, intelligent search engine, Web navigation, recommendation of Web information based on association of concept etc. We aim to utilize it as an experimental platform of Web mining on which we can design, implement, test and evaluate various algorithms of Web mining. In the paper, its architecture, main function and running environment are described.
Mingyu Lu, Shuying Pang, Yuchang Lu, Lizhu Zhou
SMC (2)5
2002 OLAP Query Processing Algorithm Based on Relational Storage
Jianhua Feng, Li Chao, Xudong Jiang 0002, Lizhu Zhou
WAIM4
2001 A New CSCW Prototype System for Color CRT CAD/CAM
abstract
A new computer supported collaborative work (CSCW) prototype system is developed to meet the needs of color CRT CAD/CAM. This paper discusses the characteristics and defines the key elements of the CSCW prototype system. The cooperation relation and cooperation work flow of color CRT CAD/CAM are provided. Compared with the traditional color CRT CAD/CAM system, the new CSCW prototype system is more suitable for the rapid design of color CRT. The efficiency of the new CSCW prototype is validated. A general CSCW architecture to support integrated product/process design and development is presented. A case study and evaluation of the architecture is also discussed. The integrated architecture will facilitate information access, sharing, and analysis among design team members using the open World Wide Web platforms and resources to design CRT products.
Yuxiu Cao, Chunxiao Xing, Lizhu Zhou
CSCWD4
2001 Closed World Assumption for Disjunctive Reasoning
Kewen Wang 0001, Lizhu Zhou
J. Comput. Sci. Technol.2
2001 An Extension to GCWA and Query Evaluation for Disjunctive Deductive Databases
Kewen Wang 0001, Lizhu Zhou
J. Intell. Inf. Syst.2