VLDB 2026 Research / reviewers in the wild / expert
Dongdong Shan
dblp:00/8701
· DBLP profile ↗
8ranked-venue papers
3as first author
0since 2021 · last 2015
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 8 · 3 first-authorArtificial intelligence and machine learning · 7 · 3 first-authorApplied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
3 papers |
Information retrieval · 60% Indexing and storage engines · 13% Data mining · 9% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Hardware accelerators and domain-specific architectures · 100% |
Topics — the 12 heaviest of 12, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Information retrieval › indexing
index compression |
0.2 | 1 | 2015 | A General SIMD-Based Approach to Accelerating Compression Algorithms · ACM Trans. Inf. Syst. 2015 |
Indexing and storage engines › data compression
integer compression |
0.2 | 1 | 2015 | A General SIMD-Based Approach to Accelerating Compression Algorithms · ACM Trans. Inf. Syst. 2015 |
Hardware accelerators and domain-specific architectures › data-parallel accelerator
SIMD accelerator |
0.2 | 1 | 2015 | A General SIMD-Based Approach to Accelerating Compression Algorithms · ACM Trans. Inf. Syst. 2015 |
Information retrieval › indexing › inverted index
block-max index |
0.1 | 1 | 2012 | Optimized top-k processing with global page scores on block-max indexes · WSDM 2012 |
Web and social media mining › event detection
burst detection |
0.1 | 1 | 2012 | EventSearch: a system for event discovery and retrieval on multi-type historical data · KDD 2012 |
Information retrieval › query processing
dynamic pruning |
0.1 | 1 | 2012 | Optimized top-k processing with global page scores on block-max indexes · WSDM 2012 |
Information retrieval › document retrieval › temporal information retrieval
event retrieval |
0.1 | 1 | 2012 | EventSearch: a system for event discovery and retrieval on multi-type historical data · KDD 2012 |
Information retrieval
query processing |
0.1 | 1 | 2012 | Optimized top-k processing with global page scores on block-max indexes · WSDM 2012 |
Information retrieval
ranking |
0.1 | 1 | 2012 | Optimized top-k processing with global page scores on block-max indexes · WSDM 2012 |
Data mining › text mining
temporal text mining |
0.1 | 1 | 2012 | EventSearch: a system for event discovery and retrieval on multi-type historical data · KDD 2012 |
Query processing and optimization
top-k query processing |
0.1 | 1 | 2012 | Optimized top-k processing with global page scores on block-max indexes · WSDM 2012 |
Information retrieval › indexing › inverted index
document identifier assignment |
0.0 | 1 | 2012 | Optimized top-k processing with global page scores on block-max indexes · WSDM 2012 |
Methods — techniques the papers use, named apart from their topics
SIMD · 0.4maxscore · 0.1burst model · 0.1block-max index · 0.1WAND · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2015 | A General SIMD-Based Approach to Accelerating Compression AlgorithmsabstractCompression algorithms are important for data-oriented tasks, especially in the era of “Big Data.” Modern processors equipped with powerful SIMD instruction sets provide us with an opportunity for achieving better compression performance. Previous research has shown that SIMD-based optimizations can multiply decoding speeds. Following these pioneering studies, we propose a general approach to accelerate compression algorithms. By instantiating the approach, we have developed several novel integer compression algorithms, called Group-Simple, Group-Scheme, Group-AFOR, and Group-PFD, and implemented their corresponding vectorized versions. We evaluate the proposed algorithms on two public TREC datasets, a Wikipedia dataset, and a Twitter dataset. With competitive compression ratios and encoding speeds, our SIMD-based algorithms outperform state-of-the-art nonvectorized algorithms with respect to decoding speeds. Wayne Xin Zhao, Daniel Lemire, Dongdong Shan, Jian-Yun Nie, Hongfei Yan, Ji-Rong Wen |
ACM Trans. Inf. Syst. | 4 |
| 2013 | Group-Scheme: SIMD-based compression algorithms for web text dataabstractCompression algorithms have been quite important for data oriented tasks, especially in the era of Big Data. The rapid development of modern processors facilitates us with powerful SIMD instruction sets, which provides an opportunity for better performance. Although SIMD based optimization on compression have been explored in some studies [2, 7], these studies usually focus on modifying the existing algorithms to fit into the SIMD instruction. In this paper, we propose a compression framework with a novel storage layout format, which aims to improve instruction-level parallelizability of compression algorithms. By instantiating the framework, we design a novel compression algorithm family, called Group-Scheme, and present a parallelized version of Group-Scheme, called SIMD-Group-Scheme. We evaluate the proposed algorithms on two public TREC data sets. With very competitive performance on compression ratio and encoding speed, SIMD-Group-Scheme significantly outperforms the implementation without SIMD instructions and state-of-the-art algorithm (i.e. SIMD-G8IU [7]), w.r.t decoding speed. Wayne Xin Zhao, Dongdong Shan, Hongfei Yan |
IEEE BigData | 3 |
| 2012 | EventSearch: a system for event discovery and retrieval on multi-type historical dataabstractWe present EventSearch, a system for event extraction and retrieval on four types of news-related historical data, i.e., Web news articles, newspapers, TV news program, and micro-blog short messages. The system incorporates over 11 million web pages extracted from "Web InfoMall", the Chinese Web Archive since 2001. The newspaper and TV news video clips also span from 2001 to 2011. The system, upon a user query, returns a list of event snippets from multiple data sources. A novel burst model is used to discover events from time-stamped texts. In addition to offline event extraction, our system also provides online event extraction to further meet the user needs. EventSearch provides meaningful analytics that synthesize an accurate description of events. Users interact with the system by ranking the identified events using different criteria (scale, recency and relevance) and submitting their own information needs in different input fields. Dongdong Shan, Wayne Xin Zhao, Rishan Chen, Baihan Shu, Ziqi Wang 0002, Hongfei Yan, Xiaoming Li 0001 |
KDD | 1 |
| 2012 | Optimized top-k processing with global page scores on block-max indexesabstractLarge web search engines are facing formidable performance challenges because they have to process thousands of queries per second on tens of billions of documents, within interactive response time. Among many others, Top-k query processing (also called early termination or dynamic pruning) is an important class of optimization techniques that can improve the search efficiency and achieve faster query processing by avoiding the scoring of documents that are unlikely to be in the top results. One recent technique is using Block-Max index. In the Block-Max index, the posting lists are organized as blocks and the maximum score for each block is stored to improve the query efficiency. Although query processing speedup is achieved with Block-Max index, the ranking function for the Top-k results is the term-based approach. It is well known that documents' static scores are also important for a good ranking function. In this paper, we show that the performance of the state-of-the-art algorithms with the Block-Max index is degraded when the static score is added in the ranking function. Then we study efficient techniques for Top-k query processing in the case where a page's static score is given, such as PageRank, in addition to the term-based approach. In particular, we propose a set of new algorithms based on the WAND and MaxScore with Block-Max index using local score, which outperform the existing ones. Then we propose new techniques to estimate a better score upper bound for each block. We also study the search efficiency on different index structures where the document identifiers are assigned by URL sorting or by static document scores. Experiments on TREC GOV2 and ClueWeb09B show that considerable performance gains are achieved. Dongdong Shan, Shuai Ding 0006, Jing He 0010, Hongfei Yan, Xiaoming Li 0001 |
WSDM | 1 |
| 2011 | Recommending citations with translation modelabstractCitation Recommendation is useful for an author to find out the papers or books that can support the materials she is writing about. It is a challengeable problem since the vocabulary used in the content of papers and in the citation contexts are usually quite different. To address this problem, we propose to use translation model, which can bridge the gap between two heterogeneous languages. We conduct an experiment and find the translation model can provide much better candidates of citations than the state-of-the-art methods. Jing He 0010, Dongdong Shan, Hongfei Yan |
CIKM | 3 |
| 2011 | Efficient phrase querying with flat position indexabstractA large proportion of search engine queries contain phrases,namely a sequence of adjacent words. In this paper, we propose to use flat position index (a.k.a schema-independent index) for phrase query evaluation. In the flat position index, the entire document collection is viewed as a huge sequence of tokens. Each token is represented by one flat position, which is a unique position offset from the beginning of the collection. Each indexed term is associated with a list of the flat positions about that term in the sequence. To recover DocID from flat positions efficiently, we propose a novel cache sensitive look-up table (CSLT), which is much faster than existing search algorithms. Experiments on TREC GOV2 data collection show that flat position index can reduce the index size and speed up phrase querying substantially, compared with traditional word-level index. Dongdong Shan, Wayne Xin Zhao, Jing He 0010, Rui Yan 0001, Hongfei Yan, Xiaoming Li 0001 |
CIKM | 1 |
| 2011 | Citation count prediction: learning to estimate future citations for literatureabstractIn most of the cases, scientists depend on previous literature which is relevant to their research fields for developing new ideas. However, it is not wise, nor possible, to track all existed publications because the volume of literature collection grows extremely fast. Therefore, researchers generally follow, or cite merely a small proportion of publications which they are interested in. For such a large collection, it is rather interesting to forecast which kind of literature is more likely to attract scientists' response. In this paper, we use the citations as a measurement for the popularity among researchers and study the interesting problem of Citation Count Prediction (CCP) to examine the characteristics for popularity. Estimation of possible popularity is of great significance and is quite challenging. We have utilized several features of fundamental characteristics for those papers that are highly cited and have predicted the popularity degree of each literature in the future. We have implemented a system which takes a series of features of a particular publication as input and produces as output the estimated citation counts of that article after a given time period. We consider several regression models to formulate the learning process and evaluate their performance based on the coefficient of determination (R-square). Experimental results on a real-large data set show that the best predictive model achieves a mean average predictive performance of 0.740 measured in R-square, which significantly outperforms several alternative algorithms. Rui Yan 0001, Jie Tang 0001, Dongdong Shan, Xiaoming Li 0001 |
CIKM | 4 |
| 2010 | Context modeling for ranking and tagging bursty features in text streamsabstractBursty features in text streams are very useful in many text mining applications. Most existing studies detect bursty features based purely on term frequency changes without taking into account the semantic contexts of terms, and as a result the detected bursty features may not always be interesting or easy to interpret. In this paper we propose to model the contexts of bursty features using a language modeling approach. We then propose a novel topic diversity-based metric using the context models to find newsworthy bursty features. We also propose to use the context models to automatically assign meaningful tags to bursty features. Using a large corpus of a stream of news articles, we quantitatively show that the proposed context language models for bursty features can effectively help rank bursty features based on their newsworthiness and to assign meaningful tags to annotate bursty features. Wayne Xin Zhao, Jing Jiang 0001, Jing He 0010, Dongdong Shan, Hongfei Yan, Xiaoming Li 0001 |
CIKM | 4 |