VLDB 2026 Research / reviewers in the wild / expert
Dongqing Yang
dblp:65/661
· DBLP profile ↗
100ranked-venue papers
1as first author
0since 2021 · last 2014
0000-0002-9610-2276ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 81 · 1 first-authorArtificial intelligence and machine learning · 25Applied, interdisciplinary, general and emerging computing · 8Systems, architecture and hardware · 2Security and privacy · 2Software engineering, systems software and programming languages · 1Human-computer interaction and ubiquitous computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
13 papers |
Web and social media mining · 21% Graph data management · 19% Data mining · 15% | |
| Artificial intelligence
2 papers |
Information extraction and text analysis · 92% Knowledge representation and reasoning · 8% | |
| Theoretical computer science
2 papers |
Graph algorithms and graph theory · 74% Mathematical optimization · 26% |
Topics — the 30 heaviest of 40, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Graph data management
shortest distance query |
0.3 | 2 | 2013 | Outsourcing shortest distance computing with privacy protection · VLDB J. 2013 Neighborhood-privacy protected shortest distance computing in cloud · SIGMOD Conference 2011 |
Natural language and speech › Information extraction and text analysis
emotion recognition |
0.2 | 1 | 2014 | Exploiting Community Emotion for Microblog Event Detection · EMNLP 2014 |
Natural language and speech › Information extraction and text analysis › event extraction
event detection |
0.2 | 1 | 2014 | Exploiting Community Emotion for Microblog Event Detection · EMNLP 2014 |
Natural language and speech › Information extraction and text analysis
sentiment analysis |
0.2 | 1 | 2014 | Exploiting Community Emotion for Microblog Event Detection · EMNLP 2014 |
Web and social media mining › social network analysis
influence maximization |
0.2 | 1 | 2013 | Influence maximization with limit cost in social network · Sci. China Inf. Sci. 2013 |
Web and social media mining
social network analysis |
0.2 | 1 | 2013 | Influence maximization with limit cost in social network · Sci. China Inf. Sci. 2013 |
Distributed and cloud data management
data placement |
0.1 | 1 | 2012 | MBA: A market-based approach to data allocation and dynamic migration for cloud database · Sci. China Inf. Sci. 2012 |
Graph data management
graph query processing |
0.1 | 1 | 2012 | Holistic Top-k Simple Shortest Path Join in Graphs · IEEE Trans. Knowl. Data Eng. 2012 |
Graph algorithms and graph theory
shortest path |
0.1 | 1 | 2012 | Holistic Top-k Simple Shortest Path Join in Graphs · IEEE Trans. Knowl. Data Eng. 2012 |
Indexing and storage engines › storage management
sparse data storage |
0.1 | 1 | 2010 | Exploring Correlated Subspaces for Efficient Query Processing in Sparse Databases · IEEE Trans. Knowl. Data Eng. 2010 |
Data mining › clustering › user clustering
customer segmentation |
0.1 | 1 | 2009 | MobileMiner: a real world case study of data mining in mobile communication · SIGMOD Conference 2009 |
Web and social media mining › community analysis
social network community analysis |
0.1 | 1 | 2009 | MobileMiner: a real world case study of data mining in mobile communication · SIGMOD Conference 2009 |
Web and social media mining › user behavior analysis
user behavior profiling |
0.1 | 1 | 2009 | MobileMiner: a real world case study of data mining in mobile communication · SIGMOD Conference 2009 |
Information retrieval › text analysis
chinese text processing |
0.1 | 1 | 2008 | A new similarity computing method based on concept similarity in Chinese text processing · Sci. China Ser. F Inf. Sci. 2008 |
Data mining › pattern mining › sequential pattern mining
closed sequential pattern mining |
0.1 | 1 | 2008 | SeqStream: Mining Closed Sequential Patterns over Stream Sliding Windows · ICDM 2008 |
Data mining
pattern mining |
0.1 | 1 | 2008 | SeqStream: Mining Closed Sequential Patterns over Stream Sliding Windows · ICDM 2008 |
Data mining › pattern mining
sequential pattern mining |
0.1 | 1 | 2008 | SeqStream: Mining Closed Sequential Patterns over Stream Sliding Windows · ICDM 2008 |
Data stream processing › continuous query processing
sliding window |
0.1 | 1 | 2008 | SeqStream: Mining Closed Sequential Patterns over Stream Sliding Windows · ICDM 2008 |
Data stream processing
stream mining |
0.1 | 1 | 2008 | SeqStream: Mining Closed Sequential Patterns over Stream Sliding Windows · ICDM 2008 |
Query processing and optimization › query execution
in-memory query processing |
0.1 | 1 | 2007 | EaseDB: a cache-oblivious in-memory query processor · SIGMOD Conference 2007 |
Data stream processing › complex event processing
pattern detection |
0.1 | 1 | 2007 | Effective variation management for pseudo periodical streams · SIGMOD Conference 2007 |
Privacy and data protection
query privacy |
0.0 | 1 | 2013 | Outsourcing shortest distance computing with privacy protection · VLDB J. 2013 |
Graph data management › graph query processing
graph join |
0.0 | 1 | 2012 | Holistic Top-k Simple Shortest Path Join in Graphs · IEEE Trans. Knowl. Data Eng. 2012 |
Cloud and datacenter computing
resource management |
0.0 | 1 | 2012 | MBA: A market-based approach to data allocation and dynamic migration for cloud database · Sci. China Inf. Sci. 2012 |
Privacy and data protection › anonymization
graph anonymization |
0.0 | 1 | 2011 | Neighborhood-privacy protected shortest distance computing in cloud · SIGMOD Conference 2011 |
Cloud and datacenter computing › cloud data management
cloud graph processing |
0.0 | 1 | 2011 | Neighborhood-privacy protected shortest distance computing in cloud · SIGMOD Conference 2011 |
Database theory
query answering |
0.0 | 1 | 2002 | COMMIX: towards effective web information extraction, integration and query answering · SIGMOD Conference 2002 |
Query processing and optimization › query rewriting
query answering using views |
0.0 | 1 | 2002 | COMMIX: towards effective web information extraction, integration and query answering · SIGMOD Conference 2002 |
Data integration and cleaning › data extraction
web data extraction |
0.0 | 1 | 2002 | COMMIX: towards effective web information extraction, integration and query answering · SIGMOD Conference 2002 |
Data integration and cleaning › semi-structured data integration
XML data integration |
0.0 | 1 | 2002 | COMMIX: towards effective web information extraction, integration and query answering · SIGMOD Conference 2002 |
Methods — techniques the papers use, named apart from their topics
greedy algorithm · 0.4pruning · 0.3market-based approach · 0.3candidate path searching · 0.3community emotion aggregation · 0.2burst detection · 0.2similarity computing · 0.2vertical partitioning · 0.1dimension correlation analysis · 0.1SQL query rewriting · 0.1data mining techniques · 0.1wave-pattern matching · 0.1cache-oblivious algorithms · 0.1cache-oblivious algorithm · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2014 | An Adaptive Skew Insensitive Join Algorithm for Large Scale Data Analytics
Wenjing Liao, Tengjiao Wang 0003, Hongyan Li 0002, Dongqing Yang, Kai Lei |
APWeb | 4 |
| 2014 | Topic-Based Sentiment Analysis Incorporating User Interactions
Wei Chen 0021, Gaoyan Ou, Tengjiao Wang 0003, Dongqing Yang, Kai Lei |
APWeb | 5 |
| 2014 | CLUSM: An Unsupervised Model for Microblog Sentiment Analysis Incorporating Link Information
Gaoyan Ou, Wei Chen 0021, Binyang Li, Tengjiao Wang 0003, Dongqing Yang, Kam-Fai Wong |
DASFAA (1) | 5 |
| 2014 | Exploiting Community Emotion for Microblog Event DetectionabstractMicroblog has become a major platform for information about real-world events.Automatically discovering realworld events from microblog has attracted the attention of many researchers.However, most of existing work ignore the importance of emotion information for event detection.We argue that people's emotional reactions immediately reflect the occurring of real-world events and should be important for event detection.In this study, we focus on the problem of communityrelated event detection by community emotions.To address the problem, we propose a novel framework which include the following three key components: microblog emotion classification, community emotion aggregation and community emotion burst detection.We evaluate our approach on real microblog data sets.Experimental results demonstrate the effectiveness of the proposed framework. Gaoyan Ou, Wei Chen 0021, Tengjiao Wang 0003, Zhongyu Wei, Binyang Li, Dongqing Yang, Kam-Fai Wong |
EMNLP | 6 |
| 2014 | BF-Matrix: A Secondary Index for the Cloud Storage
Hongyan Li 0002, Yue Wang 0014, Tengjiao Wang 0003, Dongqing Yang |
WAIM | 5 |
| 2014 | Sarcasm Detection in Social Media Based on Imbalanced Classification
Wei Chen 0021, Gaoyan Ou, Tengjiao Wang 0003, Dongqing Yang, Kai Lei |
WAIM | 5 |
| 2013 | Massively parallel learning of Bayesian networks with MapReduce for factor relationship analysisabstractBayesian Network (BN) is one of the most popular models in data mining technologies. Most of the algorithms of BN structure learning are developed for the centralized datasets, where all the data are gathered into a single computer node. They are often too costly or impractical for learning BN structures from large scale data. Through a simple interface with two functions, map and reduce, MapReduce facilitates parallel implementation of many real-world tasks such as data processing for search engines and machine learning. In this paper, we present a parallel algorithm for BN structure leaning from large-scale dateset by using a MapReduce cluster. We discuss the benefits of using MapReduce for BN structure learning, and demonstrate the performance of this approach by applying it to a real world financial factor relationships learning task from the domain of financial analysis. Wei Chen 0021, Tengjiao Wang 0003, Dongqing Yang, Kai Lei, Yueqin Liu |
IJCNN | 3 |
| 2013 | Incremental Local Evolutionary Outlier Detection for Dynamic Social Networks
Tengfei Ji, Dongqing Yang, Jun Gao 0003 |
ECML/PKDD (2) | 2 |
| 2013 | Aspect-Specific Polarity-Aware Summarization of Online Reviews
Gaoyan Ou, Wei Chen 0021, Tengjiao Wang 0003, Dongqing Yang, Kai Lei, Yueqin Liu |
WAIM | 5 |
| 2013 | Influence maximization with limit cost in social network
Yue Wang 0014, Weijing Huang, Lang Zong, Tengjiao Wang 0003, Dongqing Yang |
Sci. China Inf. Sci. | 5 |
| 2013 | Outsourcing shortest distance computing with privacy protection
Jun Gao 0003, Jeffrey Xu Yu, Ruoming Jin, Jiashuai Zhou, Tengjiao Wang 0003, Dongqing Yang |
VLDB J. | 6 |
| 2012 | A Scalable Algorithm for Detecting Community Outliers in Social Networks
Tengfei Ji, Jun Gao 0003, Dongqing Yang |
WAIM | 3 |
| 2012 | MBA: A market-based approach to data allocation and dynamic migration for cloud database
Tengjiao Wang 0003, Ziyu Lin, Bishan Yang, Jun Gao 0003, Allen Huang, Dongqing Yang, Shiwei Tang, Jinzhong Niu |
Sci. China Inf. Sci. | 6 |
| 2012 | Holistic Top-k Simple Shortest Path Join in GraphsabstractMotivated by the needs such as group relationship analysis, this paper introduces a new operation on graphs, named top-k path join, which discovers the top-k simple shortest paths between two given node sets. Rather than discovering the top-k simple paths between each node pair, this paper proposes a holistic join method which answers the top-k path join by finding constrained top-k simple shortest paths between two nodes, and then devises an efficient method to handle the latter problem. Specifically, we transform the graph by encoding the precomputed shortest paths to the target node, and use the transformed graph in the candidate path searching. We show that the candidate path searching on the transformed graph not only has the same result as that on the original graph but also can be terminated much earlier with the aid of precomputed results. We also discuss two other optimization strategies, including considering the join constraint in the candidate path generation as early as possible, and pruning search space in each candidate path generation with an adaptively determined threshold. The final extensive experimental results also show that our method offers a significant performance improvement over existing ones. Jun Gao 0003, Jeffrey Xu Yu, Huida Qiu, Tengjiao Wang 0003, Dongqing Yang |
IEEE Trans. Knowl. Data Eng. | 6 |
| 2011 | FXProj - A Fuzzy XML Documents Projected Clustering Based on Structure and Content
Tengfei Ji, Xiaoyuan Bao, Dongqing Yang |
ADMA (1) | 3 |
| 2011 | Efficient Subject-Oriented Evaluating and Mining Methods for Data with Schema Uncertainty
Yue Wang 0014, Changjie Tang, Tengjiao Wang 0003, Dongqing Yang |
ADMA (1) | 4 |
| 2011 | Influence Maximizing and Local Influenced Community Detection Based on Multiple Spread Model
Qiuling Yan, Shaosong Guo, Dongqing Yang |
ADMA (2) | 3 |
| 2011 | An Empirical Study of Massively Parallel Bayesian Networks Learning for Sentiment Extraction from Unstructured Text
Wei Chen 0021, Lang Zong, Weijing Huang, Gaoyan Ou, Yue Wang 0014, Dongqing Yang |
APWeb | 6 |
| 2011 | Neighborhood-privacy protected shortest distance computing in cloudabstractWith the advent of cloud computing, it becomes desirable to utilize cloud computing to efficiently process complex operations on large graphs without compromising their sensitive information. This paper studies shortest distance computing in the cloud, which aims at the following goals: i) preventing outsourced graphs from neighborhood attack, ii) preserving shortest distances in outsourced graphs, iii) minimizing overhead on the client side. The basic idea of this paper is to transform an original graph G into a link graph Gl kept locally and a set of outsourced graphs Go. Each outsourced graph should meet the requirement of a new security model called 1-neighborhood-d-radius. In addition, the shortest distance query can be answered using Gl and Go. Our objective is to minimize the space cost on the client side when both security and utility requirements are satisfied. We devise a greedy method to produce Gl and Go, which can exactly answer the shortest distance queries. We also develop an efficient transformation method to support approximate shortest distance answering under a given additive error bound. The final experimental results illustrate the effectiveness and efficiency of our method. Jun Gao 0003, Jeffrey Xu Yu, Ruoming Jin, Jiashuai Zhou, Tengjiao Wang 0003, Dongqing Yang |
SIGMOD Conference | 6 |
| 2011 | Informed Prediction with Incremental Core-Based Friend Cycle Discovering
Yue Wang 0014, Weijing Huang, Wei Chen 0021, Tengjiao Wang 0003, Dongqing Yang |
WAIM | 5 |
| 2011 | Worst-case VaR and robust portfolio optimization with interval random uncertainty set
Wei Chen 0021, Shaohua Tan, Dongqing Yang |
Expert Syst. Appl. | 3 |
| 2011 | DHC: Distributed, Hierarchical Clustering in Sensor Networks
Xiuli Ma, Haifeng Hu 0003, Shuangfeng Li, Hongmei Xiao, Qiong Luo 0001, Dongqing Yang, Shiwei Tang |
J. Comput. Sci. Technol. | 6 |
| 2011 | Efficient approaches for summarizing subspace clusters into k representatives
Xiuli Ma, Dongqing Yang, Shiwei Tang, Meng Shuai, Kunqing Xie |
Soft Comput. | 3 |
| 2010 | A General Multi-relational Classification Approach Using Feature Generation and Selection
Miao Zou, Tengjiao Wang 0003, Hongyan Li 0002, Dongqing Yang |
ADMA (2) | 4 |
| 2010 | Fast top-k simple shortest paths discovery in graphsabstractWith the wide applications of large scale graph data such as social networks, the problem of finding the top-k shortest paths attracts increasing attention. This paper focuses on the discovery of the top-k simple shortest paths (paths without loops). The well known algorithm for this problem is due to Yen, and the provided worstcase bound O(kn(m + nlogn)), which comes from O(n) times single-source shortest path discovery for each of k shortest paths, remains unbeaten for 30 years, where n is the number of nodes and m is the number of edges. In this paper, we observe that there are shared sub-paths among O(kn) single-source shortest paths. The basic idea behind our method is to pre-compute the shortest paths to the target node, and utilize them to reduce the discovery cost at running time. Specifically, we transform the original graph by encoding the pre-computed paths, and prove that the shortest path discovered over the transformed graph is equivalent to that in the original graph. Most importantly, the path discovery over the transformed graph can be terminated much earlier than before. In addition, two optimization strategies are presented. One is to reduce the total iteration times for shortest path discovery, and the other is to prune the search space in each iteration with an adaptively-determined threshold. Although the worst-case complexity cannot be lowered, our method is proven to be much more efficient in a general case. The final extensive experimental results (on both real and synthetic graphs) also show that our method offers a significant performance improvement over the existing ones. Jun Gao 0003, Huida Qiu, Tengjiao Wang 0003, Dongqing Yang |
CIKM | 5 |
| 2010 | Multiple Sensitive Association Protection in the Outsourced Database
Jun Gao 0003, Tengjiao Wang 0003, Dongqing Yang |
DASFAA (2) | 4 |
| 2010 | Efficient evaluation of query rewriting plan over materialized XML view
Jun Gao 0003, Jiaheng Lu, Tengjiao Wang 0003, Dongqing Yang |
J. Syst. Softw. | 4 |
| 2010 | Exploring Correlated Subspaces for Efficient Query Processing in Sparse DatabasesabstractSparse data are becoming increasingly common and available in many real-life applications. However, relatively little attention has been paid to effectively model the sparse data and existing approaches such as the conventional "horizontal¿ and "vertical¿ representations fail to provide satisfactory performance for both storage and query processing, as such approaches are too rigid and generally do not consider the dimension correlations. In this paper, we propose a new approach, named HoVer, to store and conduct query for sparse data sets in an unmodified RDBMS, where HoVer stands for horizontal representation over vertically partitioned subspaces. According to the dimension correlations of sparse data sets, a novel mechanism has been developed to vertically partition a high-dimensional sparse data set into multiple lower-dimensional subspaces, and all the dimensions are highly correlated intrasubspace and highly unrelated intersubspace, respectively. Therefore, original data objects can be represented by the horizontal format in respective subspaces. With the novel HoVer representation, users can write SQL queries over the original horizontal view, which can be easily rewritten into queries over the subspace tables. Experiments over synthetic and real-life data sets show that our approach is effective in finding correlated subspaces and yields superior performance for the storage and query of sparse data. Bin Cui 0001, Jiakui Zhao, Dongqing Yang |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2009 | Smart UI: A user interactive model in collaborative service environmentsabstractIn collaborative service environments, a Web site will provide more powerful function with lower cost just according to make some external Web services collaborate together. In these Web sites, how to get user requirement to arrange external Web services is very important. In this paper, Smart UI, a new smart user interactive model in service collaborative environment different with current simply and isolated HTML forms in Web sites is presented. In the help of interaction ontology, this model is used to interact with user friendly and intelligently, as well as it is very easy to find satisfied Web services using interaction results. We believe that most users are satisfied for this interaction model, and a Web site using this model will get much more chance to success. Qi Sui, Dongqing Yang |
CSCWD | 2 |
| 2009 | MobileMiner: a real world case study of data mining in mobile communicationabstractMobile communication data analysis has been often used as a background application to motivate many data mining problems. However, very few data mining researchers have a chance to see a working data mining system on real mobile communication data. In this demo, we showcase our new system MobileMiner on a real mobile communication data set, which presents a case study of business solutions using state-of-the-art data mining techniques. MobileMiner adaptively profiles users' behavior from their calling and moving record streams. Customer segmentation and social community analysis can be conducted based on user profiles. We show how data mining techniques can help in mobile communication data analysis. Moreover, we also show some interesting observations which still cannot be mined by the current techniques, and thus may motivate new research and development. Tengjiao Wang 0003, Bishan Yang, Jun Gao 0003, Dongqing Yang, Shiwei Tang, Kedong Liu, Jian Pei 0001 |
SIGMOD Conference | 4 |
| 2009 | A Bipartite Graph Framework for Summarizing High-Dimensional Binary, Categorical and Numeric Data
Xiuli Ma, Dongqing Yang, Shiwei Tang, Meng Shuai |
SSDBM | 3 |
| 2009 | Efficient algorithms for incremental maintenance of closed sequential patterns in large databases
Lei Chang, Tengjiao Wang 0003, Dongqing Yang, Hua Luan, Shiwei Tang |
Data Knowl. Eng. | 3 |
| 2009 | Accelerating sequence searching: dimensionality reduction method
Guojie Song, Bin Cui 0001, Baihua Zheng, Kunqing Xie, Dongqing Yang |
Knowl. Inf. Syst. | 5 |
| 2008 | T-rotation: Multiple Publications of Privacy Preserving Data Sequence
Youdong Tao, Yunhai Tong, Shaohua Tan, Shiwei Tang, Dongqing Yang |
ADMA | 5 |
| 2008 | Squeezing Long Sequence Data for Efficient Similarity Search
Guojie Song, Bin Cui 0001, Baihua Zheng, Kunqing Xie, Dongqing Yang |
APWeb | 5 |
| 2008 | Effective Data Distribution and Reallocation Strategies for Fast Query Response in Distributed Query-Intensive Data Environments
Tengjiao Wang 0003, Bishan Yang, Jun Gao 0003, Dongqing Yang |
APWeb | 4 |
| 2008 | Protecting the Publishing Identity in Multiple Tuples
Youdong Tao, Yunhai Tong, Shaohua Tan, Shiwei Tang, Dongqing Yang |
DBSec | 5 |
| 2008 | Effective Skyline Cardinality Estimation on Data Streams
Jiakui Zhao, Lijun Chen 0002, Bin Cui 0001, Dongqing Yang |
DEXA | 5 |
| 2008 | SeqStream: Mining Closed Sequential Patterns over Stream Sliding WindowsabstractPrevious studies have shown mining closed patterns provides more benefits than mining the complete set of frequent patterns, since closed pattern mining leads to more compact results and more efficient algorithms. It is quite useful in a data stream environment where memory and computation power are major concerns. This paper studies the problem of mining closed sequential patterns over data stream sliding windows. A synopsis structure IST (Inverse Closed Sequence Tree) is designed to keep inverse closed sequential patterns in current window. An efficient algorithm SeqStream is developed to mine closed sequential patterns in stream windows incrementally, and various novel strategies are adopted in SeqStream to prune search space aggressively. Extensive experiments on both real and synthetic data sets show that SeqStream outperforms PrefixSpan, CloSpan and BIDE by a factor of about one to two orders of magnitude. Lei Chang, Tengjiao Wang 0003, Dongqing Yang, Hua Luan |
ICDM | 3 |
| 2008 | Road Network Based Adaptive Query Evaluation in VANETabstractIn the Vehicle Ad-hoc NETwork (VANET), moving vehicles organize into a mobile wireless Ad-hoc network to share online traffic information. Each vehicle can issue a declarative query for aggregating the traffic information from others in order to facilitate the navigation and avoid traffic jam. Existing query methods suffer from high latency, incomplete results, and large messages due to the movement of the vehicles in VANET. In this paper, we propose an adaptive query evaluation method based on the road network. In order to overcome the problems incurred by the movement, a relative static query evaluation plan is constructed based on the road network, and each vehicle can participate in the query evaluation plan autonomously. We also introduce control messages to notify the changed location of the query originator to other vehicles involved in the evaluation plan. In addition, we propose an one time message transferring based results collecting method to reduce the message cost. The optimization over the multiple queries is also discussed to reduce the messages further. We evaluate the performance of our method by extensive simulations. Experimental results show that our method can provide complete results within a short response time and small traffic overhead. Jun Gao 0003, Jinsong Han, Dongqing Yang, Tengjiao Wang 0003 |
MDM | 3 |
| 2008 | BOAI: Fast Alternating Decision Tree Induction Based on Bottom-Up Evaluation
Bishan Yang, Tengjiao Wang 0003, Dongqing Yang, Lei Chang |
PAKDD | 3 |
| 2008 | A new similarity computing method based on concept similarity in Chinese text processing
Dongqing Yang, Shiwei Tang, Tengjiao Wang 0003, Jun Gao 0003 |
Sci. China Ser. F Inf. Sci. | 2 |
| 2008 | XFlat: Query-friendly encrypted XML view publishing
Jun Gao 0003, Tengjiao Wang 0003, Dongqing Yang |
Inf. Sci. | 3 |
| 2008 | PGG: An Online Pattern Based Approach for Stream Variation Management
Lu-An Tang, Bin Cui 0001, Hongyan Li 0002, Gaoshan Miao, Dongqing Yang, Xinbiao Zhou |
J. Comput. Sci. Technol. | 5 |
| 2007 | A Coding Hierarchy Computing Based Clustering Algorithm
Jing Peng 0002, Changjie Tang, Dongqing Yang, An-long Chen, Lei Duan |
ADMA | 3 |
| 2007 | CACS: A Novel Classification Algorithm Based on Concept Similarity
Jing Peng 0002, Dongqing Yang, Changjie Tang, Jianjun Hu |
ADMA | 2 |
| 2007 | A general framework for improving query processing performance on multi-level memory hierarchiesabstractWe propose a general framework for improving the query processing performance on multi-level memory hierarchies. Our motivation is that (1) the memory hierarchy is an important performance factor for query processing, (2) both the memory hierarchy and database systems are becoming increasingly complex and diverse, and (3) increasing the amount of tuning does not always improve the performance. Therefore, we categorize multiple levels of memory performance tuning and quantify their performance impacts. As a case study, we use this framework to improve the in-memory performance of storage models, B+-trees, nested-loop joins and hash joins. Our empirical evaluation verifies the usefulness of the proposed framework. Bingsheng He, Qiong Luo 0001, Dongqing Yang |
DaMoN | 4 |
| 2007 | Web Service Composition Based on Message Schema Analysis
Aiqiang Gao, Dongqing Yang, Shiwei Tang |
DASFAA | 2 |
| 2007 | Optimizing Moving Queries over Moving Object Data Streams
Dan Lin 0001, Bin Cui 0001, Dongqing Yang |
DASFAA | 3 |
| 2007 | CLAIM: An Efficient Method for Relaxed Frequent Closed Itemsets Mining over Stream Data
Guojie Song, Dongqing Yang, Bin Cui 0001, Baihua Zheng, Kunqing Xie |
DASFAA | 2 |
| 2007 | An Optimized Process Neural Network Model
Guojie Song, Dongqing Yang, Bin Cui 0001, Kunqing Xie |
DASFAA | 2 |
| 2007 | Evaluating MAX and MIN over Sliding Windows with Various Size Using the Exemplary Sketch
Jiakui Zhao, Dongqing Yang, Bin Cui 0001, Lijun Chen 0002, Jun Gao 0003 |
DASFAA | 2 |
| 2007 | MQTree Based Query Rewriting over Multiple XML Views
Jun Gao 0003, Tengjiao Wang 0003, Dongqing Yang |
DEXA | 3 |
| 2007 | A New Text Clustering Method Using Hidden Markov Model
Dongqing Yang, Shiwei Tang, Tengjiao Wang 0003, Aiqiang Gao |
NLDB | 2 |
| 2007 | EaseDB: a cache-oblivious in-memory query processorabstractWe propose to demonstrate EaseDB, the first cache-oblivious queryprocessor for in-memory relational query processing. The cache-oblivious notion from the theory community refers to the property that no parameters in an algorithm or a data structure need to be tuned for a specific memory hierarchy for optimality. As a result, EaseDB automatically optimizes the cache performance as well as the overall performance of query processing on any memory hierarchy. We have developed a visualization interface to show the detailed performance of EaseDB in comparison with its cache-conscious counterpart, with both the parameters in the cache-conscious algorithms and the hardware platforms varied. Bingsheng He, Qiong Luo 0001, Dongqing Yang |
SIGMOD Conference | 4 |
| 2007 | Effective variation management for pseudo periodical streamsabstractMany database applications require the analysis and processing of data streams. In such systems, huge amounts of data arrive rapidly and their values change over time. The variations on streams typically imply some fundamental changes of the underlying objects and possess significant domain meanings. In some data streams, successive events seem to recur in a certain time interval, but the data indeed evolves with tiny differences as time elapses. This feature is called pseudo periodicity, which poses a non-trivial challenge to variation management in data streams. This paper presents our research effort in online variation management over such streams, and the idea can be applied to the problem domain of medical applications, such as patient vital signal monitoring. We propose a new method named Pattern Growth Graph (PGG) to detect and manage variations over pseudo periodical streams. PGG adopts the wave-pattern to capture the major information of data evolution and represent them compactly. With the help of wave-pattern matching algorithm, PGG detects the stream variations in a single pass over the stream data. PGG only stores the different segments of the pattern for incoming stream, and hence it can substantially compress the data without losing important information. The statistical information of PGG helps to distinguish meaningful data changes from noise and to reconstruct the stream with acceptable accuracy. Extensive experiments on real datasets containing millions of data items demonstrate the feasibility and effectiveness of the proposed scheme. Lv-an Tang, Bin Cui 0001, Hongyan Li 0002, Gaoshan Miao, Dongqing Yang, Xinbiao Zhou |
SIGMOD Conference | 5 |
| 2007 | An Adaptive Approach to Schema Classification for Data Warehouse Modeling
Hongding Wang, Yunhai Tong, Shaohua Tan, Shiwei Tang, Dongqing Yang, Guohui Sun |
J. Comput. Sci. Technol. | 5 |
| 2006 | Mining Compressed Sequential Patterns
Lei Chang, Dongqing Yang, Shiwei Tang, Tengjiao Wang 0003 |
ADMA | 2 |
| 2006 | XFlat: Query Friendly Encrypted XML View Publishing
Jun Gao 0003, Tengjiao Wang 0003, Dongqing Yang |
APWeb | 3 |
| 2006 | QoS-Driven Web Service Composition with Inter Service Conflicts
Aiqiang Gao, Dongqing Yang, Shiwei Tang, Ming Zhang 0004 |
APWeb | 2 |
| 2006 | WISE: A Prototype for Ontology Driven Development of Web Information Systems
Lv-an Tang, Hongyan Li 0002, Baojun Qiu, Meimei Li, Dongqing Yang, Shiwei Tang |
APWeb | 8 |
| 2006 | Mining Models of Composite Web Services for Performance Analysis
Aiqiang Gao, Dongqing Yang, Shiwei Tang, Ming Zhang 0004 |
DASFAA | 2 |
| 2006 | Triple-driven data modeling methodology in data warehousing: a case studyabstractIn this paper, we present a useful data modeling methodology in data warehousing which integrates three existing approaches normally used in isolation: goal-driven, data-driven and user-driven. It comprises of four stages. Goal-driven stage produces subjects and KPIs(Key Performance Indicators) of main business fields. Data-driven stage produces subject oriented enterprise data schema. User-driven stage yields analytical requirements represented by measures and dimensions of each subject. Combination stage combines the triple-driven results. By triple-driven, we can get a more complete, more structured and more layered data model of a data warehouse. We illustrate each stage step by step using examples in our case study. Yuhong Guo, Shiwei Tang, Yunhai Tong, Dongqing Yang |
DOLAP | 4 |
| 2006 | KD3 Scheme for Privacy Preserving Data Mining
Yunhai Tong, Shiwei Tang, Dongqing Yang |
ISI | 4 |
| 2006 | Mining Maximal Correlated Member Clusters in High Dimensional Database
Lizheng Jiang, Dongqing Yang, Shiwei Tang, Xiuli Ma, Dehui Zhang |
PAKDD | 2 |
| 2006 | CCWrapper: Adaptive Predefined Schema Guided Web Extraction
Jun Gao 0003, Dongqing Yang, Tengjiao Wang 0003 |
WAIM | 2 |
| 2006 | KCAM: Concentrating on Structural Similarity for XML Fragments
Lingbo Kong, Shiwei Tang, Dongqing Yang, Tengjiao Wang 0003, Jun Gao 0003 |
WAIM | 3 |
| 2006 | Binary Search Join between an IR System and an RDBMSabstractIntegrating relational database technologies into Web information retrieval enables users to ask complex queries beyond traditional keyword searches over Web pages. One approach to this integration is to have a software layer on top of an information retrieval (IR) system and an RDBMS (relational database management system). A core operation in this top layer is to join the intermediate results from the two underlying systems (called the IR results and the DB results correspondingly) in order to produce the final ranked results for each query. Unfortunately, most conventional join algorithms are inefficient for this operation. In this paper, we propose one simple join algorithm called binary search join (BSJ) for the operation of joining the IR results and the DB results. This algorithm takes advantage of the fact that the IR results are already ranked by relevance and that the DB results are already sorted by the join attribute. It scans the IR results and for each IR result tuple performs a binary search over the DB results. We analytically and empirically study the performance of BSJ in comparison with several conventional join algorithms on a repository of Chinese news Web pages. The experiment results prove that BSJ works best in most cases Ernest Dawei Wang, Qiong Luo 0001, Dongqing Yang, Shiwei Tang |
Web Intelligence | 3 |
| 2006 | Cardinality Computing: A New Step Towards Fully Representing Multi-sets by Bloom Filters
Jiakui Zhao, Dongqing Yang, Lijun Chen 0002, Jun Gao 0003, Tengjiao Wang 0003 |
WISE | 2 |
| 2005 | Privacy Preserving Naive Bayes Classification
Yunhai Tong, Shiwei Tang, Dongqing Yang |
ADMA | 4 |
| 2005 | Finding Hidden Semantics Behind Reference Linkages : An Ontological Approach for Scientific Digital Libraries
Peixiang Zhao 0002, Ming Zhang 0004, Dongqing Yang, Shiwei Tang |
DASFAA | 3 |
| 2005 | Avoiding Error-Prone Reordering Optimization During Legal Systems Migration
Youlin Fang, Dongqing Yang |
DEXA | 3 |
| 2005 | Mining Association Rules from Distorted Data for Privacy Preservation
Yunhai Tong, Shiwei Tang, Dongqing Yang |
KES (3) | 4 |
| 2005 | Mining Most General Multidimensional Summarization of Probably Groups in Data Warehouses
Jian Pei 0001, Shiwei Tang, Dongqing Yang |
SSDBM | 4 |
| 2005 | Web Service Composition Using Markov Decision Processes
Aiqiang Gao, Dongqing Yang, Shiwei Tang, Ming Zhang 0004 |
WAIM | 2 |
| 2005 | Understanding User Operations on Web Page in WISE
Hongyan Li 0002, Ming Xue, Shiwei Tang, Dongqing Yang |
WAIM | 5 |
| 2005 | An Ontology Based Approach to Construct Behaviors in Web Information Systems
Lv-an Tang, Hongyan Li 0002, Zhiyong Pan, Dongqing Yang, Meimei Li, Shiwei Tang, Ying Ying |
WAIM | 4 |
| 2005 | Validating key constraints over XML document using XPath and structure checking
Dongqing Yang, Shiwei Tang, Tengjiao Wang 0003, Jun Gao 0003 |
Future Gener. Comput. Syst. | 2 |
| 2004 | A Comparative Study on Feature Weight in Text Categorization
Zhi-Hong Deng 0001, Shiwei Tang, Dongqing Yang, Ming Zhang 0004, Liyu Li, Kunqing Xie |
APWeb | 3 |
| 2004 | WIEAS: Helping to Discover Web Information Sources and Extract Data from Them
Liyu Li, Shiwei Tang, Dongqing Yang, Tengjiao Wang 0003, Zhi-Hong Deng 0001, Zhihua Su |
APWeb | 3 |
| 2004 | EGA: An Algorithm for Automatic Semi-structured Web Documents Extraction
Liyu Li, Shiwei Tang, Dongqing Yang, Tengjiao Wang 0003, Zhihua Su |
DASFAA | 3 |
| 2004 | Incremental Maintenance of Discovered Mobile User Maximal Moving Sequential Patterns
Shuai Ma 0001, Shiwei Tang, Dongqing Yang, Tengjiao Wang 0003, Chanjun Yang |
DASFAA | 3 |
| 2004 | Discovering and Generating Materialized XML Views in Data Integration Systems
Tengjiao Wang 0003, Dongqing Yang, Shiwei Tang |
IDEAS | 2 |
| 2004 | Online Mining in Sensor Networks
Xiuli Ma, Dongqing Yang, Shiwei Tang, Qiong Luo 0001, Dehui Zhang, Shuangfeng Li |
NPC | 2 |
| 2004 | Combining Clustering with Moving Sequential Pattern Mining: A Novel and Efficient Technique
Shuai Ma 0001, Shiwei Tang, Dongqing Yang, Tengjiao Wang 0003, Jinqiang Han |
PAKDD | 3 |
| 2004 | QReduction: Synopsizing XPath Query Set Efficiently under Resource Constraint
Jun Gao 0003, Xiuli Ma, Dongqing Yang, Tengjiao Wang 0003, Shiwei Tang |
WAIM | 3 |
| 2004 | Extracting Key Value and Checking Structural Constraints for Validating XML Key Constraints
Dongqing Yang, Shiwei Tang, Tengjiao Wang 0003, Jun Gao 0003 |
WAIM | 2 |
| 2004 | Discovery of Frequent XML Query Patterns with DTD Cardinality Constraints
Dongqing Yang, Shiwei Tang, Tengjiao Wang 0003, Jun Gao 0003 |
WAIM | 2 |
| 2004 | Efficient Incremental Maintenance of Frequent Patterns with FP-Tree
Xiuli Ma, Yunhai Tong, Shiwei Tang, Dongqing Yang |
J. Comput. Sci. Technol. | 4 |
| 2003 | CyberETL: Towards Visual Debugging Transformations in Data Integration
Youlin Fang, Dongqing Yang, Shiwei Tang, Yunhai Tong, Libo Yu |
WAIM | 2 |
| 2003 | A New Fast Clustering Algorithm Based on Reference and Density
Shuai Ma 0001, Tengjiao Wang 0003, Shiwei Tang, Dongqing Yang, Jun Gao 0003 |
WAIM | 4 |
| 2003 | Normalizing XML Element Trees as Well-Designed Document Structures for Data Integration
Wenbing Zhao 0002, Shaohua Tan, Dongqing Yang, Shiwei Tang |
WAIM | 3 |
| 2002 | COMMIX: towards effective web information extraction, integration and query answeringabstractAs WWW becomes more and more popular and powerful, how to search information on the web in database way becomes an important research topic. COMMIX, which is developed in the DB group in Peking University (China), is a system towards building very large database using data from the Web for information extraction, integration and query answering. COMMIX has some innovative features, such as ontology-based wrapper generation, XML-based information integration, view-based query answering, and QBE-style XML query interface. Tengjiao Wang 0003, Shiwei Tang, Dongqing Yang, Jun Gao 0003, Yuqing Wu, Jian Pei 0001 |
SIGMOD Conference | 3 |
| 2002 | Towards Efficient Re-mining of Frequent Patterns upon Threshold Changes
Xiuli Ma, Shiwei Tang, Dongqing Yang, Xiaoping Du |
WAIM | 3 |
| 2001 | H-Mine: Hyper-Structure Mining of Frequent Patterns in Large DatabasesabstractMethods for efficient mining of frequent patterns have been studied extensively by many researchers. However, the previously proposed methods still encounter some performance bottlenecks when mining databases with different data characteristics, such as dense vs. sparse, long vs. short patterns, memory-based vs. disk-based, etc. In this study, we propose a simple and novel hyper-linked data structure, H-struct and a new mining algorithm, H-mine, which takes advantage of this data structure and dynamically adjusts links in the mining process. A distinct feature of this method is that it has very limited and precisely predictable space overhead and runs really fast in memory-based setting. Moreover it can be scaled up to very large databases by database partitioning, and when the data set becomes dense, (conditional) FP-trees can be constructed dynamically as part of the mining process. Our study shows that H-mine has high performance in various kinds of data, outperforms the previously developed algorithms in different settings, and is highly scalable in mining large databases. This study also proposes a new data mining methodology, space-preserving mining, which may have strong impact in the future development of efficient and scalable data mining methods. Jian Pei 0001, Jiawei Han 0001, Hongjun Lu, Shojiro Nishio, Shiwei Tang, Dongqing Yang |
ICDM | 6 |
| 2001 | An XML Based Electronic Medical Record Integration System
Hongyan Li 0002, Shiwei Tang, Dongqing Yang |
WAIM | 3 |
| 2001 | A New Conceptual Graph Generated Algorithm for Semi-structured Databases
Kam-Fai Wong, Yat Fan Su, Dongqing Yang, Shiwei Tang |
Web Intelligence | 3 |
| 2001 | Extracting Local Schema from Semistructured Data Based on Graph-Oriented Semantic Model
Tengjiao Wang 0003, Shiwei Tang, Dongqing Yang |
J. Comput. Sci. Technol. | 3 |
| 1991 | Usage refinement for ER-to-relation design transformations
Toby J. Teorey, Dongqing Yang |
Inf. Sci. | 2 |
| 1987 | A practical approach to transforming extended ER diagrams into the relational model
Dongqing Yang, Toby J. Teorey, James P. Fry |
Inf. Sci. | 1 |