Dongqing Yang

dblp:65/661 · DBLP profile ↗
← Back
100ranked-venue papers
1as first author
0since 2021 · last 2014
0000-0002-9610-2276ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 81 · 1 first-authorArtificial intelligence and machine learning · 25Applied, interdisciplinary, general and emerging computing · 8Systems, architecture and hardware · 2Security and privacy · 2Software engineering, systems software and programming languages · 1Human-computer interaction and ubiquitous computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
13 papers
Web and social media mining · 21% Graph data management · 19% Data mining · 15%
Artificial intelligence
2 papers
Information extraction and text analysis · 92% Knowledge representation and reasoning · 8%
Theoretical computer science
2 papers
Graph algorithms and graph theory · 74% Mathematical optimization · 26%

Topics — the 30 heaviest of 40, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Graph data management
shortest distance query
0.322013
Outsourcing shortest distance computing with privacy protection · VLDB J. 2013
Neighborhood-privacy protected shortest distance computing in cloud · SIGMOD Conference 2011
Natural language and speech › Information extraction and text analysis
emotion recognition
0.212014
Exploiting Community Emotion for Microblog Event Detection · EMNLP 2014
Natural language and speech › Information extraction and text analysis › event extraction
event detection
0.212014
Exploiting Community Emotion for Microblog Event Detection · EMNLP 2014
Natural language and speech › Information extraction and text analysis
sentiment analysis
0.212014
Exploiting Community Emotion for Microblog Event Detection · EMNLP 2014
Web and social media mining › social network analysis
influence maximization
0.212013
Influence maximization with limit cost in social network · Sci. China Inf. Sci. 2013
Web and social media mining
social network analysis
0.212013
Influence maximization with limit cost in social network · Sci. China Inf. Sci. 2013
Distributed and cloud data management
data placement
0.112012
MBA: A market-based approach to data allocation and dynamic migration for cloud database · Sci. China Inf. Sci. 2012
Graph data management
graph query processing
0.112012
Holistic Top-k Simple Shortest Path Join in Graphs · IEEE Trans. Knowl. Data Eng. 2012
Graph algorithms and graph theory
shortest path
0.112012
Holistic Top-k Simple Shortest Path Join in Graphs · IEEE Trans. Knowl. Data Eng. 2012
Indexing and storage engines › storage management
sparse data storage
0.112010
Exploring Correlated Subspaces for Efficient Query Processing in Sparse Databases · IEEE Trans. Knowl. Data Eng. 2010
Data mining › clustering › user clustering
customer segmentation
0.112009
MobileMiner: a real world case study of data mining in mobile communication · SIGMOD Conference 2009
Web and social media mining › community analysis
social network community analysis
0.112009
MobileMiner: a real world case study of data mining in mobile communication · SIGMOD Conference 2009
Web and social media mining › user behavior analysis
user behavior profiling
0.112009
MobileMiner: a real world case study of data mining in mobile communication · SIGMOD Conference 2009
Information retrieval › text analysis
chinese text processing
0.112008
A new similarity computing method based on concept similarity in Chinese text processing · Sci. China Ser. F Inf. Sci. 2008
Data mining › pattern mining › sequential pattern mining
closed sequential pattern mining
0.112008
SeqStream: Mining Closed Sequential Patterns over Stream Sliding Windows · ICDM 2008
Data mining
pattern mining
0.112008
SeqStream: Mining Closed Sequential Patterns over Stream Sliding Windows · ICDM 2008
Data mining › pattern mining
sequential pattern mining
0.112008
SeqStream: Mining Closed Sequential Patterns over Stream Sliding Windows · ICDM 2008
Data stream processing › continuous query processing
sliding window
0.112008
SeqStream: Mining Closed Sequential Patterns over Stream Sliding Windows · ICDM 2008
Data stream processing
stream mining
0.112008
SeqStream: Mining Closed Sequential Patterns over Stream Sliding Windows · ICDM 2008
Query processing and optimization › query execution
in-memory query processing
0.112007
EaseDB: a cache-oblivious in-memory query processor · SIGMOD Conference 2007
Data stream processing › complex event processing
pattern detection
0.112007
Effective variation management for pseudo periodical streams · SIGMOD Conference 2007
Privacy and data protection
query privacy
0.012013
Outsourcing shortest distance computing with privacy protection · VLDB J. 2013
Graph data management › graph query processing
graph join
0.012012
Holistic Top-k Simple Shortest Path Join in Graphs · IEEE Trans. Knowl. Data Eng. 2012
Cloud and datacenter computing
resource management
0.012012
MBA: A market-based approach to data allocation and dynamic migration for cloud database · Sci. China Inf. Sci. 2012
Privacy and data protection › anonymization
graph anonymization
0.012011
Neighborhood-privacy protected shortest distance computing in cloud · SIGMOD Conference 2011
Cloud and datacenter computing › cloud data management
cloud graph processing
0.012011
Neighborhood-privacy protected shortest distance computing in cloud · SIGMOD Conference 2011
Database theory
query answering
0.012002
COMMIX: towards effective web information extraction, integration and query answering · SIGMOD Conference 2002
Query processing and optimization › query rewriting
query answering using views
0.012002
COMMIX: towards effective web information extraction, integration and query answering · SIGMOD Conference 2002
Data integration and cleaning › data extraction
web data extraction
0.012002
COMMIX: towards effective web information extraction, integration and query answering · SIGMOD Conference 2002
Data integration and cleaning › semi-structured data integration
XML data integration
0.012002
COMMIX: towards effective web information extraction, integration and query answering · SIGMOD Conference 2002

Methods — techniques the papers use, named apart from their topics

greedy algorithm · 0.4pruning · 0.3market-based approach · 0.3candidate path searching · 0.3community emotion aggregation · 0.2burst detection · 0.2similarity computing · 0.2vertical partitioning · 0.1dimension correlation analysis · 0.1SQL query rewriting · 0.1data mining techniques · 0.1wave-pattern matching · 0.1cache-oblivious algorithms · 0.1cache-oblivious algorithm · 0.1
YearPublicationVenuePosition
2014 An Adaptive Skew Insensitive Join Algorithm for Large Scale Data Analytics
Wenjing Liao, Tengjiao Wang 0003, Hongyan Li 0002, Dongqing Yang, Kai Lei
APWeb4
2014 Topic-Based Sentiment Analysis Incorporating User Interactions
Wei Chen 0021, Gaoyan Ou, Tengjiao Wang 0003, Dongqing Yang, Kai Lei
APWeb5
2014 CLUSM: An Unsupervised Model for Microblog Sentiment Analysis Incorporating Link Information
Gaoyan Ou, Wei Chen 0021, Binyang Li, Tengjiao Wang 0003, Dongqing Yang, Kam-Fai Wong
DASFAA (1)5
2014 Exploiting Community Emotion for Microblog Event Detection
abstract
Microblog has become a major platform for information about real-world events.Automatically discovering realworld events from microblog has attracted the attention of many researchers.However, most of existing work ignore the importance of emotion information for event detection.We argue that people's emotional reactions immediately reflect the occurring of real-world events and should be important for event detection.In this study, we focus on the problem of communityrelated event detection by community emotions.To address the problem, we propose a novel framework which include the following three key components: microblog emotion classification, community emotion aggregation and community emotion burst detection.We evaluate our approach on real microblog data sets.Experimental results demonstrate the effectiveness of the proposed framework.
Gaoyan Ou, Wei Chen 0021, Tengjiao Wang 0003, Zhongyu Wei, Binyang Li, Dongqing Yang, Kam-Fai Wong
EMNLP6
2014 BF-Matrix: A Secondary Index for the Cloud Storage
Hongyan Li 0002, Yue Wang 0014, Tengjiao Wang 0003, Dongqing Yang
WAIM5
2014 Sarcasm Detection in Social Media Based on Imbalanced Classification
Wei Chen 0021, Gaoyan Ou, Tengjiao Wang 0003, Dongqing Yang, Kai Lei
WAIM5
2013 Massively parallel learning of Bayesian networks with MapReduce for factor relationship analysis
abstract
Bayesian Network (BN) is one of the most popular models in data mining technologies. Most of the algorithms of BN structure learning are developed for the centralized datasets, where all the data are gathered into a single computer node. They are often too costly or impractical for learning BN structures from large scale data. Through a simple interface with two functions, map and reduce, MapReduce facilitates parallel implementation of many real-world tasks such as data processing for search engines and machine learning. In this paper, we present a parallel algorithm for BN structure leaning from large-scale dateset by using a MapReduce cluster. We discuss the benefits of using MapReduce for BN structure learning, and demonstrate the performance of this approach by applying it to a real world financial factor relationships learning task from the domain of financial analysis.
Wei Chen 0021, Tengjiao Wang 0003, Dongqing Yang, Kai Lei, Yueqin Liu
IJCNN3
2013 Incremental Local Evolutionary Outlier Detection for Dynamic Social Networks
Tengfei Ji, Dongqing Yang, Jun Gao 0003
ECML/PKDD (2)2
2013 Aspect-Specific Polarity-Aware Summarization of Online Reviews
Gaoyan Ou, Wei Chen 0021, Tengjiao Wang 0003, Dongqing Yang, Kai Lei, Yueqin Liu
WAIM5
2013 Influence maximization with limit cost in social network
Yue Wang 0014, Weijing Huang, Lang Zong, Tengjiao Wang 0003, Dongqing Yang
Sci. China Inf. Sci.5
2013 Outsourcing shortest distance computing with privacy protection
Jun Gao 0003, Jeffrey Xu Yu, Ruoming Jin, Jiashuai Zhou, Tengjiao Wang 0003, Dongqing Yang
VLDB J.6
2012 A Scalable Algorithm for Detecting Community Outliers in Social Networks
Tengfei Ji, Jun Gao 0003, Dongqing Yang
WAIM3
2012 MBA: A market-based approach to data allocation and dynamic migration for cloud database
Tengjiao Wang 0003, Ziyu Lin, Bishan Yang, Jun Gao 0003, Allen Huang, Dongqing Yang, Shiwei Tang, Jinzhong Niu
Sci. China Inf. Sci.6
2012 Holistic Top-k Simple Shortest Path Join in Graphs
abstract
Motivated by the needs such as group relationship analysis, this paper introduces a new operation on graphs, named top-k path join, which discovers the top-k simple shortest paths between two given node sets. Rather than discovering the top-k simple paths between each node pair, this paper proposes a holistic join method which answers the top-k path join by finding constrained top-k simple shortest paths between two nodes, and then devises an efficient method to handle the latter problem. Specifically, we transform the graph by encoding the precomputed shortest paths to the target node, and use the transformed graph in the candidate path searching. We show that the candidate path searching on the transformed graph not only has the same result as that on the original graph but also can be terminated much earlier with the aid of precomputed results. We also discuss two other optimization strategies, including considering the join constraint in the candidate path generation as early as possible, and pruning search space in each candidate path generation with an adaptively determined threshold. The final extensive experimental results also show that our method offers a significant performance improvement over existing ones.
Jun Gao 0003, Jeffrey Xu Yu, Huida Qiu, Tengjiao Wang 0003, Dongqing Yang
IEEE Trans. Knowl. Data Eng.6
2011 FXProj - A Fuzzy XML Documents Projected Clustering Based on Structure and Content
Tengfei Ji, Xiaoyuan Bao, Dongqing Yang
ADMA (1)3
2011 Efficient Subject-Oriented Evaluating and Mining Methods for Data with Schema Uncertainty
Yue Wang 0014, Changjie Tang, Tengjiao Wang 0003, Dongqing Yang
ADMA (1)4
2011 Influence Maximizing and Local Influenced Community Detection Based on Multiple Spread Model
Qiuling Yan, Shaosong Guo, Dongqing Yang
ADMA (2)3
2011 An Empirical Study of Massively Parallel Bayesian Networks Learning for Sentiment Extraction from Unstructured Text
Wei Chen 0021, Lang Zong, Weijing Huang, Gaoyan Ou, Yue Wang 0014, Dongqing Yang
APWeb6
2011 Neighborhood-privacy protected shortest distance computing in cloud
abstract
With the advent of cloud computing, it becomes desirable to utilize cloud computing to efficiently process complex operations on large graphs without compromising their sensitive information. This paper studies shortest distance computing in the cloud, which aims at the following goals: i) preventing outsourced graphs from neighborhood attack, ii) preserving shortest distances in outsourced graphs, iii) minimizing overhead on the client side. The basic idea of this paper is to transform an original graph G into a link graph Gl kept locally and a set of outsourced graphs Go. Each outsourced graph should meet the requirement of a new security model called 1-neighborhood-d-radius. In addition, the shortest distance query can be answered using Gl and Go. Our objective is to minimize the space cost on the client side when both security and utility requirements are satisfied. We devise a greedy method to produce Gl and Go, which can exactly answer the shortest distance queries. We also develop an efficient transformation method to support approximate shortest distance answering under a given additive error bound. The final experimental results illustrate the effectiveness and efficiency of our method.
Jun Gao 0003, Jeffrey Xu Yu, Ruoming Jin, Jiashuai Zhou, Tengjiao Wang 0003, Dongqing Yang
SIGMOD Conference6
2011 Informed Prediction with Incremental Core-Based Friend Cycle Discovering
Yue Wang 0014, Weijing Huang, Wei Chen 0021, Tengjiao Wang 0003, Dongqing Yang
WAIM5
2011 Worst-case VaR and robust portfolio optimization with interval random uncertainty set
Wei Chen 0021, Shaohua Tan, Dongqing Yang
Expert Syst. Appl.3
2011 DHC: Distributed, Hierarchical Clustering in Sensor Networks
Xiuli Ma, Haifeng Hu 0003, Shuangfeng Li, Hongmei Xiao, Qiong Luo 0001, Dongqing Yang, Shiwei Tang
J. Comput. Sci. Technol.6
2011 Efficient approaches for summarizing subspace clusters into k representatives
Xiuli Ma, Dongqing Yang, Shiwei Tang, Meng Shuai, Kunqing Xie
Soft Comput.3
2010 A General Multi-relational Classification Approach Using Feature Generation and Selection
Miao Zou, Tengjiao Wang 0003, Hongyan Li 0002, Dongqing Yang
ADMA (2)4
2010 Fast top-k simple shortest paths discovery in graphs
abstract
With the wide applications of large scale graph data such as social networks, the problem of finding the top-k shortest paths attracts increasing attention. This paper focuses on the discovery of the top-k simple shortest paths (paths without loops). The well known algorithm for this problem is due to Yen, and the provided worstcase bound O(kn(m + nlogn)), which comes from O(n) times single-source shortest path discovery for each of k shortest paths, remains unbeaten for 30 years, where n is the number of nodes and m is the number of edges. In this paper, we observe that there are shared sub-paths among O(kn) single-source shortest paths. The basic idea behind our method is to pre-compute the shortest paths to the target node, and utilize them to reduce the discovery cost at running time. Specifically, we transform the original graph by encoding the pre-computed paths, and prove that the shortest path discovered over the transformed graph is equivalent to that in the original graph. Most importantly, the path discovery over the transformed graph can be terminated much earlier than before. In addition, two optimization strategies are presented. One is to reduce the total iteration times for shortest path discovery, and the other is to prune the search space in each iteration with an adaptively-determined threshold. Although the worst-case complexity cannot be lowered, our method is proven to be much more efficient in a general case. The final extensive experimental results (on both real and synthetic graphs) also show that our method offers a significant performance improvement over the existing ones.
Jun Gao 0003, Huida Qiu, Tengjiao Wang 0003, Dongqing Yang
CIKM5
2010 Multiple Sensitive Association Protection in the Outsourced Database
Jun Gao 0003, Tengjiao Wang 0003, Dongqing Yang
DASFAA (2)4
2010 Efficient evaluation of query rewriting plan over materialized XML view
Jun Gao 0003, Jiaheng Lu, Tengjiao Wang 0003, Dongqing Yang
J. Syst. Softw.4
2010 Exploring Correlated Subspaces for Efficient Query Processing in Sparse Databases
abstract
Sparse data are becoming increasingly common and available in many real-life applications. However, relatively little attention has been paid to effectively model the sparse data and existing approaches such as the conventional "horizontal¿ and "vertical¿ representations fail to provide satisfactory performance for both storage and query processing, as such approaches are too rigid and generally do not consider the dimension correlations. In this paper, we propose a new approach, named HoVer, to store and conduct query for sparse data sets in an unmodified RDBMS, where HoVer stands for horizontal representation over vertically partitioned subspaces. According to the dimension correlations of sparse data sets, a novel mechanism has been developed to vertically partition a high-dimensional sparse data set into multiple lower-dimensional subspaces, and all the dimensions are highly correlated intrasubspace and highly unrelated intersubspace, respectively. Therefore, original data objects can be represented by the horizontal format in respective subspaces. With the novel HoVer representation, users can write SQL queries over the original horizontal view, which can be easily rewritten into queries over the subspace tables. Experiments over synthetic and real-life data sets show that our approach is effective in finding correlated subspaces and yields superior performance for the storage and query of sparse data.
Bin Cui 0001, Jiakui Zhao, Dongqing Yang
IEEE Trans. Knowl. Data Eng.3
2009 Smart UI: A user interactive model in collaborative service environments
abstract
In collaborative service environments, a Web site will provide more powerful function with lower cost just according to make some external Web services collaborate together. In these Web sites, how to get user requirement to arrange external Web services is very important. In this paper, Smart UI, a new smart user interactive model in service collaborative environment different with current simply and isolated HTML forms in Web sites is presented. In the help of interaction ontology, this model is used to interact with user friendly and intelligently, as well as it is very easy to find satisfied Web services using interaction results. We believe that most users are satisfied for this interaction model, and a Web site using this model will get much more chance to success.
Qi Sui, Dongqing Yang
CSCWD2
2009 MobileMiner: a real world case study of data mining in mobile communication
abstract
Mobile communication data analysis has been often used as a background application to motivate many data mining problems. However, very few data mining researchers have a chance to see a working data mining system on real mobile communication data. In this demo, we showcase our new system MobileMiner on a real mobile communication data set, which presents a case study of business solutions using state-of-the-art data mining techniques. MobileMiner adaptively profiles users' behavior from their calling and moving record streams. Customer segmentation and social community analysis can be conducted based on user profiles. We show how data mining techniques can help in mobile communication data analysis. Moreover, we also show some interesting observations which still cannot be mined by the current techniques, and thus may motivate new research and development.
Tengjiao Wang 0003, Bishan Yang, Jun Gao 0003, Dongqing Yang, Shiwei Tang, Kedong Liu, Jian Pei 0001
SIGMOD Conference4
2009 A Bipartite Graph Framework for Summarizing High-Dimensional Binary, Categorical and Numeric Data
Xiuli Ma, Dongqing Yang, Shiwei Tang, Meng Shuai
SSDBM3
2009 Efficient algorithms for incremental maintenance of closed sequential patterns in large databases
Lei Chang, Tengjiao Wang 0003, Dongqing Yang, Hua Luan, Shiwei Tang
Data Knowl. Eng.3
2009 Accelerating sequence searching: dimensionality reduction method
Guojie Song, Bin Cui 0001, Baihua Zheng, Kunqing Xie, Dongqing Yang
Knowl. Inf. Syst.5
2008 T-rotation: Multiple Publications of Privacy Preserving Data Sequence
Youdong Tao, Yunhai Tong, Shaohua Tan, Shiwei Tang, Dongqing Yang
ADMA5
2008 Squeezing Long Sequence Data for Efficient Similarity Search
Guojie Song, Bin Cui 0001, Baihua Zheng, Kunqing Xie, Dongqing Yang
APWeb5
2008 Effective Data Distribution and Reallocation Strategies for Fast Query Response in Distributed Query-Intensive Data Environments
Tengjiao Wang 0003, Bishan Yang, Jun Gao 0003, Dongqing Yang
APWeb4
2008 Protecting the Publishing Identity in Multiple Tuples
Youdong Tao, Yunhai Tong, Shaohua Tan, Shiwei Tang, Dongqing Yang
DBSec5
2008 Effective Skyline Cardinality Estimation on Data Streams
Jiakui Zhao, Lijun Chen 0002, Bin Cui 0001, Dongqing Yang
DEXA5
2008 SeqStream: Mining Closed Sequential Patterns over Stream Sliding Windows
abstract
Previous studies have shown mining closed patterns provides more benefits than mining the complete set of frequent patterns, since closed pattern mining leads to more compact results and more efficient algorithms. It is quite useful in a data stream environment where memory and computation power are major concerns. This paper studies the problem of mining closed sequential patterns over data stream sliding windows. A synopsis structure IST (Inverse Closed Sequence Tree) is designed to keep inverse closed sequential patterns in current window. An efficient algorithm SeqStream is developed to mine closed sequential patterns in stream windows incrementally, and various novel strategies are adopted in SeqStream to prune search space aggressively. Extensive experiments on both real and synthetic data sets show that SeqStream outperforms PrefixSpan, CloSpan and BIDE by a factor of about one to two orders of magnitude.
Lei Chang, Tengjiao Wang 0003, Dongqing Yang, Hua Luan
ICDM3
2008 Road Network Based Adaptive Query Evaluation in VANET
abstract
In the Vehicle Ad-hoc NETwork (VANET), moving vehicles organize into a mobile wireless Ad-hoc network to share online traffic information. Each vehicle can issue a declarative query for aggregating the traffic information from others in order to facilitate the navigation and avoid traffic jam. Existing query methods suffer from high latency, incomplete results, and large messages due to the movement of the vehicles in VANET. In this paper, we propose an adaptive query evaluation method based on the road network. In order to overcome the problems incurred by the movement, a relative static query evaluation plan is constructed based on the road network, and each vehicle can participate in the query evaluation plan autonomously. We also introduce control messages to notify the changed location of the query originator to other vehicles involved in the evaluation plan. In addition, we propose an one time message transferring based results collecting method to reduce the message cost. The optimization over the multiple queries is also discussed to reduce the messages further. We evaluate the performance of our method by extensive simulations. Experimental results show that our method can provide complete results within a short response time and small traffic overhead.
Jun Gao 0003, Jinsong Han, Dongqing Yang, Tengjiao Wang 0003
MDM3
2008 BOAI: Fast Alternating Decision Tree Induction Based on Bottom-Up Evaluation
Bishan Yang, Tengjiao Wang 0003, Dongqing Yang, Lei Chang
PAKDD3
2008 A new similarity computing method based on concept similarity in Chinese text processing
Dongqing Yang, Shiwei Tang, Tengjiao Wang 0003, Jun Gao 0003
Sci. China Ser. F Inf. Sci.2
2008 XFlat: Query-friendly encrypted XML view publishing
Jun Gao 0003, Tengjiao Wang 0003, Dongqing Yang
Inf. Sci.3
2008 PGG: An Online Pattern Based Approach for Stream Variation Management
Lu-An Tang, Bin Cui 0001, Hongyan Li 0002, Gaoshan Miao, Dongqing Yang, Xinbiao Zhou
J. Comput. Sci. Technol.5
2007 A Coding Hierarchy Computing Based Clustering Algorithm
Jing Peng 0002, Changjie Tang, Dongqing Yang, An-long Chen, Lei Duan
ADMA3
2007 CACS: A Novel Classification Algorithm Based on Concept Similarity
Jing Peng 0002, Dongqing Yang, Changjie Tang, Jianjun Hu
ADMA2
2007 A general framework for improving query processing performance on multi-level memory hierarchies
abstract
We propose a general framework for improving the query processing performance on multi-level memory hierarchies. Our motivation is that (1) the memory hierarchy is an important performance factor for query processing, (2) both the memory hierarchy and database systems are becoming increasingly complex and diverse, and (3) increasing the amount of tuning does not always improve the performance. Therefore, we categorize multiple levels of memory performance tuning and quantify their performance impacts. As a case study, we use this framework to improve the in-memory performance of storage models, B+-trees, nested-loop joins and hash joins. Our empirical evaluation verifies the usefulness of the proposed framework.
Bingsheng He, Qiong Luo 0001, Dongqing Yang
DaMoN4
2007 Web Service Composition Based on Message Schema Analysis
Aiqiang Gao, Dongqing Yang, Shiwei Tang
DASFAA2
2007 Optimizing Moving Queries over Moving Object Data Streams
Dan Lin 0001, Bin Cui 0001, Dongqing Yang
DASFAA3
2007 CLAIM: An Efficient Method for Relaxed Frequent Closed Itemsets Mining over Stream Data
Guojie Song, Dongqing Yang, Bin Cui 0001, Baihua Zheng, Kunqing Xie
DASFAA2
2007 An Optimized Process Neural Network Model
Guojie Song, Dongqing Yang, Bin Cui 0001, Kunqing Xie
DASFAA2
2007 Evaluating MAX and MIN over Sliding Windows with Various Size Using the Exemplary Sketch
Jiakui Zhao, Dongqing Yang, Bin Cui 0001, Lijun Chen 0002, Jun Gao 0003
DASFAA2
2007 MQTree Based Query Rewriting over Multiple XML Views
Jun Gao 0003, Tengjiao Wang 0003, Dongqing Yang
DEXA3
2007 A New Text Clustering Method Using Hidden Markov Model
Dongqing Yang, Shiwei Tang, Tengjiao Wang 0003, Aiqiang Gao
NLDB2
2007 EaseDB: a cache-oblivious in-memory query processor
abstract
We propose to demonstrate EaseDB, the first cache-oblivious queryprocessor for in-memory relational query processing. The cache-oblivious notion from the theory community refers to the property that no parameters in an algorithm or a data structure need to be tuned for a specific memory hierarchy for optimality. As a result, EaseDB automatically optimizes the cache performance as well as the overall performance of query processing on any memory hierarchy. We have developed a visualization interface to show the detailed performance of EaseDB in comparison with its cache-conscious counterpart, with both the parameters in the cache-conscious algorithms and the hardware platforms varied.
Bingsheng He, Qiong Luo 0001, Dongqing Yang
SIGMOD Conference4
2007 Effective variation management for pseudo periodical streams
abstract
Many database applications require the analysis and processing of data streams. In such systems, huge amounts of data arrive rapidly and their values change over time. The variations on streams typically imply some fundamental changes of the underlying objects and possess significant domain meanings. In some data streams, successive events seem to recur in a certain time interval, but the data indeed evolves with tiny differences as time elapses. This feature is called pseudo periodicity, which poses a non-trivial challenge to variation management in data streams. This paper presents our research effort in online variation management over such streams, and the idea can be applied to the problem domain of medical applications, such as patient vital signal monitoring. We propose a new method named Pattern Growth Graph (PGG) to detect and manage variations over pseudo periodical streams. PGG adopts the wave-pattern to capture the major information of data evolution and represent them compactly. With the help of wave-pattern matching algorithm, PGG detects the stream variations in a single pass over the stream data. PGG only stores the different segments of the pattern for incoming stream, and hence it can substantially compress the data without losing important information. The statistical information of PGG helps to distinguish meaningful data changes from noise and to reconstruct the stream with acceptable accuracy. Extensive experiments on real datasets containing millions of data items demonstrate the feasibility and effectiveness of the proposed scheme.
Lv-an Tang, Bin Cui 0001, Hongyan Li 0002, Gaoshan Miao, Dongqing Yang, Xinbiao Zhou
SIGMOD Conference5
2007 An Adaptive Approach to Schema Classification for Data Warehouse Modeling
Hongding Wang, Yunhai Tong, Shaohua Tan, Shiwei Tang, Dongqing Yang, Guohui Sun
J. Comput. Sci. Technol.5
2006 Mining Compressed Sequential Patterns
Lei Chang, Dongqing Yang, Shiwei Tang, Tengjiao Wang 0003
ADMA2
2006 XFlat: Query Friendly Encrypted XML View Publishing
Jun Gao 0003, Tengjiao Wang 0003, Dongqing Yang
APWeb3
2006 QoS-Driven Web Service Composition with Inter Service Conflicts
Aiqiang Gao, Dongqing Yang, Shiwei Tang, Ming Zhang 0004
APWeb2
2006 WISE: A Prototype for Ontology Driven Development of Web Information Systems
Lv-an Tang, Hongyan Li 0002, Baojun Qiu, Meimei Li, Dongqing Yang, Shiwei Tang
APWeb8
2006 Mining Models of Composite Web Services for Performance Analysis
Aiqiang Gao, Dongqing Yang, Shiwei Tang, Ming Zhang 0004
DASFAA2
2006 Triple-driven data modeling methodology in data warehousing: a case study
abstract
In this paper, we present a useful data modeling methodology in data warehousing which integrates three existing approaches normally used in isolation: goal-driven, data-driven and user-driven. It comprises of four stages. Goal-driven stage produces subjects and KPIs(Key Performance Indicators) of main business fields. Data-driven stage produces subject oriented enterprise data schema. User-driven stage yields analytical requirements represented by measures and dimensions of each subject. Combination stage combines the triple-driven results. By triple-driven, we can get a more complete, more structured and more layered data model of a data warehouse. We illustrate each stage step by step using examples in our case study.
Yuhong Guo, Shiwei Tang, Yunhai Tong, Dongqing Yang
DOLAP4
2006 KD3 Scheme for Privacy Preserving Data Mining
Yunhai Tong, Shiwei Tang, Dongqing Yang
ISI4
2006 Mining Maximal Correlated Member Clusters in High Dimensional Database
Lizheng Jiang, Dongqing Yang, Shiwei Tang, Xiuli Ma, Dehui Zhang
PAKDD2
2006 CCWrapper: Adaptive Predefined Schema Guided Web Extraction
Jun Gao 0003, Dongqing Yang, Tengjiao Wang 0003
WAIM2
2006 KCAM: Concentrating on Structural Similarity for XML Fragments
Lingbo Kong, Shiwei Tang, Dongqing Yang, Tengjiao Wang 0003, Jun Gao 0003
WAIM3
2006 Binary Search Join between an IR System and an RDBMS
abstract
Integrating relational database technologies into Web information retrieval enables users to ask complex queries beyond traditional keyword searches over Web pages. One approach to this integration is to have a software layer on top of an information retrieval (IR) system and an RDBMS (relational database management system). A core operation in this top layer is to join the intermediate results from the two underlying systems (called the IR results and the DB results correspondingly) in order to produce the final ranked results for each query. Unfortunately, most conventional join algorithms are inefficient for this operation. In this paper, we propose one simple join algorithm called binary search join (BSJ) for the operation of joining the IR results and the DB results. This algorithm takes advantage of the fact that the IR results are already ranked by relevance and that the DB results are already sorted by the join attribute. It scans the IR results and for each IR result tuple performs a binary search over the DB results. We analytically and empirically study the performance of BSJ in comparison with several conventional join algorithms on a repository of Chinese news Web pages. The experiment results prove that BSJ works best in most cases
Ernest Dawei Wang, Qiong Luo 0001, Dongqing Yang, Shiwei Tang
Web Intelligence3
2006 Cardinality Computing: A New Step Towards Fully Representing Multi-sets by Bloom Filters
Jiakui Zhao, Dongqing Yang, Lijun Chen 0002, Jun Gao 0003, Tengjiao Wang 0003
WISE2
2005 Privacy Preserving Naive Bayes Classification
Yunhai Tong, Shiwei Tang, Dongqing Yang
ADMA4
2005 Finding Hidden Semantics Behind Reference Linkages : An Ontological Approach for Scientific Digital Libraries
Peixiang Zhao 0002, Ming Zhang 0004, Dongqing Yang, Shiwei Tang
DASFAA3
2005 Avoiding Error-Prone Reordering Optimization During Legal Systems Migration
Youlin Fang, Dongqing Yang
DEXA3
2005 Mining Association Rules from Distorted Data for Privacy Preservation
Yunhai Tong, Shiwei Tang, Dongqing Yang
KES (3)4
2005 Mining Most General Multidimensional Summarization of Probably Groups in Data Warehouses
Jian Pei 0001, Shiwei Tang, Dongqing Yang
SSDBM4
2005 Web Service Composition Using Markov Decision Processes
Aiqiang Gao, Dongqing Yang, Shiwei Tang, Ming Zhang 0004
WAIM2
2005 Understanding User Operations on Web Page in WISE
Hongyan Li 0002, Ming Xue, Shiwei Tang, Dongqing Yang
WAIM5
2005 An Ontology Based Approach to Construct Behaviors in Web Information Systems
Lv-an Tang, Hongyan Li 0002, Zhiyong Pan, Dongqing Yang, Meimei Li, Shiwei Tang, Ying Ying
WAIM4
2005 Validating key constraints over XML document using XPath and structure checking
Dongqing Yang, Shiwei Tang, Tengjiao Wang 0003, Jun Gao 0003
Future Gener. Comput. Syst.2
2004 A Comparative Study on Feature Weight in Text Categorization
Zhi-Hong Deng 0001, Shiwei Tang, Dongqing Yang, Ming Zhang 0004, Liyu Li, Kunqing Xie
APWeb3
2004 WIEAS: Helping to Discover Web Information Sources and Extract Data from Them
Liyu Li, Shiwei Tang, Dongqing Yang, Tengjiao Wang 0003, Zhi-Hong Deng 0001, Zhihua Su
APWeb3
2004 EGA: An Algorithm for Automatic Semi-structured Web Documents Extraction
Liyu Li, Shiwei Tang, Dongqing Yang, Tengjiao Wang 0003, Zhihua Su
DASFAA3
2004 Incremental Maintenance of Discovered Mobile User Maximal Moving Sequential Patterns
Shuai Ma 0001, Shiwei Tang, Dongqing Yang, Tengjiao Wang 0003, Chanjun Yang
DASFAA3
2004 Discovering and Generating Materialized XML Views in Data Integration Systems
Tengjiao Wang 0003, Dongqing Yang, Shiwei Tang
IDEAS2
2004 Online Mining in Sensor Networks
Xiuli Ma, Dongqing Yang, Shiwei Tang, Qiong Luo 0001, Dehui Zhang, Shuangfeng Li
NPC2
2004 Combining Clustering with Moving Sequential Pattern Mining: A Novel and Efficient Technique
Shuai Ma 0001, Shiwei Tang, Dongqing Yang, Tengjiao Wang 0003, Jinqiang Han
PAKDD3
2004 QReduction: Synopsizing XPath Query Set Efficiently under Resource Constraint
Jun Gao 0003, Xiuli Ma, Dongqing Yang, Tengjiao Wang 0003, Shiwei Tang
WAIM3
2004 Extracting Key Value and Checking Structural Constraints for Validating XML Key Constraints
Dongqing Yang, Shiwei Tang, Tengjiao Wang 0003, Jun Gao 0003
WAIM2
2004 Discovery of Frequent XML Query Patterns with DTD Cardinality Constraints
Dongqing Yang, Shiwei Tang, Tengjiao Wang 0003, Jun Gao 0003
WAIM2
2004 Efficient Incremental Maintenance of Frequent Patterns with FP-Tree
Xiuli Ma, Yunhai Tong, Shiwei Tang, Dongqing Yang
J. Comput. Sci. Technol.4
2003 CyberETL: Towards Visual Debugging Transformations in Data Integration
Youlin Fang, Dongqing Yang, Shiwei Tang, Yunhai Tong, Libo Yu
WAIM2
2003 A New Fast Clustering Algorithm Based on Reference and Density
Shuai Ma 0001, Tengjiao Wang 0003, Shiwei Tang, Dongqing Yang, Jun Gao 0003
WAIM4
2003 Normalizing XML Element Trees as Well-Designed Document Structures for Data Integration
Wenbing Zhao 0002, Shaohua Tan, Dongqing Yang, Shiwei Tang
WAIM3
2002 COMMIX: towards effective web information extraction, integration and query answering
abstract
As WWW becomes more and more popular and powerful, how to search information on the web in database way becomes an important research topic. COMMIX, which is developed in the DB group in Peking University (China), is a system towards building very large database using data from the Web for information extraction, integration and query answering. COMMIX has some innovative features, such as ontology-based wrapper generation, XML-based information integration, view-based query answering, and QBE-style XML query interface.
Tengjiao Wang 0003, Shiwei Tang, Dongqing Yang, Jun Gao 0003, Yuqing Wu, Jian Pei 0001
SIGMOD Conference3
2002 Towards Efficient Re-mining of Frequent Patterns upon Threshold Changes
Xiuli Ma, Shiwei Tang, Dongqing Yang, Xiaoping Du
WAIM3
2001 H-Mine: Hyper-Structure Mining of Frequent Patterns in Large Databases
abstract
Methods for efficient mining of frequent patterns have been studied extensively by many researchers. However, the previously proposed methods still encounter some performance bottlenecks when mining databases with different data characteristics, such as dense vs. sparse, long vs. short patterns, memory-based vs. disk-based, etc. In this study, we propose a simple and novel hyper-linked data structure, H-struct and a new mining algorithm, H-mine, which takes advantage of this data structure and dynamically adjusts links in the mining process. A distinct feature of this method is that it has very limited and precisely predictable space overhead and runs really fast in memory-based setting. Moreover it can be scaled up to very large databases by database partitioning, and when the data set becomes dense, (conditional) FP-trees can be constructed dynamically as part of the mining process. Our study shows that H-mine has high performance in various kinds of data, outperforms the previously developed algorithms in different settings, and is highly scalable in mining large databases. This study also proposes a new data mining methodology, space-preserving mining, which may have strong impact in the future development of efficient and scalable data mining methods.
Jian Pei 0001, Jiawei Han 0001, Hongjun Lu, Shojiro Nishio, Shiwei Tang, Dongqing Yang
ICDM6
2001 An XML Based Electronic Medical Record Integration System
Hongyan Li 0002, Shiwei Tang, Dongqing Yang
WAIM3
2001 A New Conceptual Graph Generated Algorithm for Semi-structured Databases
Kam-Fai Wong, Yat Fan Su, Dongqing Yang, Shiwei Tang
Web Intelligence3
2001 Extracting Local Schema from Semistructured Data Based on Graph-Oriented Semantic Model
Tengjiao Wang 0003, Shiwei Tang, Dongqing Yang
J. Comput. Sci. Technol.3
1991 Usage refinement for ER-to-relation design transformations
Toby J. Teorey, Dongqing Yang
Inf. Sci.2
1987 A practical approach to transforming extended ER diagrams into the relational model
Dongqing Yang, Toby J. Teorey, James P. Fry
Inf. Sci.1