VLDB 2026 Research / reviewers in the wild / expert
Fan Zhang 0092
dblp:21/3626-92
· DBLP profile ↗
13ranked-venue papers
3as first author
2since 2021 · last 2024
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 8 · 1 first-authorArtificial intelligence and machine learning · 5 · 2 first-authorSystems, architecture and hardware · 2 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
5 papers |
Data mining · 51% Information retrieval · 39% Web and social media mining · 10% | |
| Artificial intelligence
1 paper |
Information extraction and text analysis · 50% Knowledge representation and reasoning · 50% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Computational finance and economics · 100% |
Topics — the 14 heaviest of 17, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Data mining
anomaly detection |
0.3 | 1 | 2017 | Mining Suspicious Tax Evasion Groups in Big Data · ICDE 2017 |
Data mining › pattern mining
graph pattern mining |
0.3 | 1 | 2017 | Mining Suspicious Tax Evasion Groups in Big Data · ICDE 2017 |
Data mining › anomaly detection
graph anomaly detection |
0.2 | 1 | 2016 | Mining Suspicious Tax Evasion Groups in Big Data · IEEE Trans. Knowl. Data Eng. 2016 |
Data mining
pattern mining |
0.2 | 1 | 2016 | Mining Suspicious Tax Evasion Groups in Big Data · IEEE Trans. Knowl. Data Eng. 2016 |
Web and social media mining
user identity linkage |
0.2 | 1 | 2013 | What's in a name?: an unsupervised approach to link users across communities · WSDM 2013 |
Natural language and speech › Information extraction and text analysis
relation extraction |
0.1 | 1 | 2011 | Nonlinear Evidence Fusion and Propagation for Hyponymy Relation Mining · ACL 2011 |
Information retrieval › indexing
index compression |
0.1 | 1 | 2011 | Efficient Parallel Lists Intersection and Index Compression Algorithms using Graphics Processing Units · Proc. VLDB Endow. 2011 |
Information retrieval › indexing › index compression
inverted index compression |
0.1 | 1 | 2011 | Efficient Parallel Lists Intersection and Index Compression Algorithms using Graphics Processing Units · Proc. VLDB Endow. 2011 |
Information retrieval › query processing
web query processing |
0.1 | 1 | 2011 | Efficient Parallel Lists Intersection and Index Compression Algorithms using Graphics Processing Units · Proc. VLDB Endow. 2011 |
Information retrieval
document retrieval |
0.1 | 1 | 2010 | Revisiting globally sorted indexes for efficient document retrieval · WSDM 2010 |
Information retrieval › query processing
early termination |
0.1 | 1 | 2010 | Revisiting globally sorted indexes for efficient document retrieval · WSDM 2010 |
Information retrieval
indexing |
0.1 | 1 | 2010 | Revisiting globally sorted indexes for efficient document retrieval · WSDM 2010 |
Information retrieval › indexing
inverted index |
0.1 | 1 | 2010 | Revisiting globally sorted indexes for efficient document retrieval · WSDM 2010 |
GPUs and heterogeneous computing
GPU query processing |
0.0 | 1 | 2011 | Efficient Parallel Lists Intersection and Index Compression Algorithms using Graphics Processing Units · Proc. VLDB Endow. 2011 |
Methods — techniques the papers use, named apart from their topics
pattern tree matching · 0.8graph mining · 0.5graph-based method · 0.3colored network model · 0.3linear regression · 0.2hash segmentation · 0.2d-gap compression · 0.2binary search · 0.2unsupervised learning · 0.2n-gram probability · 0.2nonlinear evidence fusion · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | SkillNet-X: A Multilingual Multitask Model with Sparsely Activated SkillsabstractTraditional multitask learning methods typically can only leverage shared knowledge within specific tasks or languages, resulting in a loss of either cross-language or cross-task knowledge. This paper proposes a general multilingual multitask model, named SkillNet-X, which enables a single model to tackle many different tasks from different languages. To this end, we define several language-specific skills and task-specific skills, each of which corresponds to a skill module. SkillNet-X sparsely activates parts of the skill modules which are relevant to eitherthe target task or the target language. Acting as knowledge transit hubs, skill modules are capable of absorbing task-related knowledge and language-related knowledge consecutively. We evaluate SkillNet-X on eleven natural language understanding datasets in four languages. Results show that SkillNet-X performs better than task-specific and two multitask learning baselines.To investigate the generalization of our model, we conduct experiments on two new tasks and find that SkillNet-X significantly outperforms baselines. Zhangyin Feng, Yong Dai 0001, Fan Zhang 0092, Duyu Tang, Shuangzhi Wu, Bing Qin 0001, Yunbo Cao, Shuming Shi 0001 |
ICASSP | 3 |
| 2023 | Skillnet-NLG: General-Purpose Natural Language Generation with a Sparsely Activated ApproachabstractWe present SkillNet-NLG, a sparsely activated approach that handles many natural language generation tasks with one model. Different from traditional dense models that always activate all the parameters, SkillNet-NLG selectively activates relevant parts of the parameters to accomplish a task, where the relevance is controlled by a set of predefined skills. The strength of such model design is that it provides an opportunity to precisely adapt relevant skills to learn new tasks effectively. We evaluate on Chinese natural language generation tasks. Results show that, with only one model file, SkillNet-NLG outperforms previous best performance methods on four of five tasks. SkillNet-NLG performs better than two multitask learning baselines (a dense model and a Mixture-of-Expert model) and achieves comparable performance to task-specific models. Lastly, SkillNet-NLG surpasses baseline systems when adapted to new tasks. Junwei Liao, Duyu Tang, Fan Zhang 0092, Shuming Shi 0001 |
ICASSP | 3 |
| 2017 | Mining Suspicious Tax Evasion Groups in Big DataabstractThere is evidence that an increasing number of enterprises plot together to evade tax in an unperceived way. At the same time, the taxation information related data is a classic kind of big data. These issues challenge the effectiveness of traditional data mining-based tax evasion detection methods. To address this problem, we first investigate the classic tax evasion cases, and employ a graph-based method to characterize their property that describes two suspicious relationship trails with a same antecedent node behind an Interest-Affiliated Transaction (IAT). Next, we propose a Colored Network-Based Model (CNBM) for characterizing economic behaviors, social relationships and the IATs between taxpayers, and generating a Taxpayer Interest Interacted Network (TPIIN). To accomplish the tax evasion detection task by discovering suspicious groups in a TPIIN, methods for building a patterns tree and matching component patterns are introduced and the completeness of the methods based on graph theory is presented. Then, we describe an experiment based on real data and a simulated network. The experimental results show that our proposed method greatly improves the efficiency of tax evasion detection, as well as provides a clear explanation of the tax evasion behaviors of taxpayer groups. Feng Tian 0002, Kuo-Ming Chao, Nick Godwin, Nazaraf Shah, Fan Zhang 0092 |
ICDE | 7 |
| 2016 | Mining Suspicious Tax Evasion Groups in Big DataabstractThere is evidence that an increasing number of enterprises plot together to evade tax in an unperceived way. At the same time, the taxation information related data is a classic kind of big data. These issues challenge the effectiveness of traditional data mining-based tax evasion detection methods. To address this problem, we first investigate the classic tax evasion cases, and employ a graph-based method to characterize their property that describes two suspicious relationship trails with a same antecedent node behind an Interest-Affiliated Transaction (IAT). Next, we propose a Colored Network-Based Model (CNBM) for characterizing economic behaviors, social relationships, and the IATs between taxpayers, and generating a Taxpayer Interest Interacted Network (TPIIN). To accomplish the tax evasion detection task by discovering suspicious groups in a TPIIN, methods for building a patterns tree and matching component patterns are introduced and the completeness of the methods based on graph theory is presented. Then, we describe an experiment based on real data and a simulated network. The experimental results show that our proposed method greatly improves the efficiency of tax evasion detection, as well as provides a clear explanation of the tax evasion behaviors of taxpayer groups. Feng Tian 0002, Kuo-Ming Chao, Nick Godwin, Nazaraf Shah, Fan Zhang 0092 |
IEEE Trans. Knowl. Data Eng. | 7 |
| 2013 | Exploiting query term correlation for list caching in web search enginesabstractCaching technologies have been widely employed to boost the performance of Web search engines. Motivated by the correlation between terms in query logs from a commercial search engine, we explore the idea of a caching scheme based on pairs of terms, rather than individual terms (which is the typical approach used by search engines today). We propose an inverted list caching policy, based on the Least Recently Used method, in which the co-occurring correlation between terms in the query stream is accounted for when deciding on which terms to keep in the cache. We consider not only the term co-occurrence within the same query but also the co-occurrence between separate queries. Experimental results show that the proposed approach can improve not only the cache hit ratio but also the overall throughput of the system when compared to existing list caching algorithms. Jiancong Tong, Gang Wang 0001, Douglas S. Stones, Shizhao Sun, Xiaoguang Liu 0001, Fan Zhang 0092 |
CIKM | 6 |
| 2013 | What's in a name?: an unsupervised approach to link users across communitiesabstractIn this paper, we consider the problem of linking users across multiple online communities. Specifically, we focus on the alias-disambiguation step of this user linking task, which is meant to differentiate users with the same usernames. We start quantitatively analyzing the importance of the alias-disambiguation step by conducting a survey on 153 volunteers and an experimental analysis on a large dataset of About.me (75,472 users). The analysis shows that the alias-disambiguation solution can address a major part of the user linking problem in terms of the coverage of true pairwise decisions (46.8%). To the best of our knowledge, this is the first study on human behaviors with regards to the usages of online usernames. We then cast the alias-disambiguation step as a pairwise classification problem and propose a novel unsupervised approach. The key idea of our approach is to automatically label training instances based on two observations: (a) rare usernames are likely owned by a single natural person, e.g. pennystar88 as a positive instance; (b) common usernames are likely owned by different natural persons, e.g. tank as a negative instance. We propose using the n-gram probabilities of usernames to estimate the rareness or commonness of usernames. Moreover, these two observations are verified by using the dataset of Yahoo! Answers. The empirical evaluations on 53 forums verify: (a) the effectiveness of the classifiers with the automatically generated training data and (b) that the rareness and commonness of usernames can help user linking. We also analyze the cases where the classifiers fail. Jing Liu 0022, Fan Zhang 0092, Xinying Song, Young-In Song, Chin-Yew Lin, Hsiao-Wuen Hon |
WSDM | 2 |
| 2013 | Mining subtopics from text fragments for a web query
Qinglei Wang, Ya-nan Qian, Ruihua Song, Zhicheng Dou, Fan Zhang 0092, Tetsuya Sakai |
Inf. Retr. | 5 |
| 2011 | Nonlinear Evidence Fusion and Propagation for Hyponymy Relation Mining
Fan Zhang 0092, Shuming Shi 0001, Jing Liu 0022, Shu-Qi Sun, Chin-Yew Lin |
ACL | 1 |
| 2011 | Efficient Parallel Lists Intersection and Index Compression Algorithms using Graphics Processing UnitsabstractMajor web search engines answer thousands of queries per second requesting information about billions of web pages. The data sizes and query loads are growing at an exponential rate. To manage the heavy workload, we consider techniques for utilizing a Graphics Processing Unit (GPU). We investigate new approaches to improve two important operations of search engines -- lists intersection and index compression. For lists intersection, we develop techniques for efficient implementation of the binary search algorithm for parallel computation. We inspect some representative real-world datasets and find that a sufficiently long inverted list has an overall linear rate of increase. Based on this observation, we propose Linear Regression and Hash Segmentation techniques for contracting the search range. For index compression, the traditional d-gap based compression schemata are not well-suited for parallel computation, so we propose a Linear Regression Compression schema which has an inherent parallel structure. We further discuss how to efficiently intersect the compressed lists on a GPU. Our experimental results show significant improvements in the query processing throughput on several datasets. Naiyong Ao, Fan Zhang 0092, Di Wu 0036, Douglas S. Stones, Gang Wang 0001, Xiaoguang Liu 0001, Jing Liu 0010, Sheng Lin 0002 |
Proc. VLDB Endow. | 2 |
| 2010 | Efficient term proximity search with term-pair indexesabstractThere has been a large amount of research on early termination techniques in web search and information retrieval. Such techniques return the top-k documents without scanning and evaluating the full inverted lists of the query terms. Thus, they can greatly improve query processing efficiency. However, only a limited amount of efficient top-k processing work considers the impact of term proximity, i.e., the distance between term occurrences in a document, which has recently been integrated into a number of retrieval models to improve effectiveness. Shuming Shi 0001, Fan Zhang 0092, Torsten Suel, Ji-Rong Wen |
CIKM | 3 |
| 2010 | An Improved Parallel MEMS Processing-Level Simulation Implementation Using Graphic Processing Unit
Yupeng Guo, Xiaoguang Liu 0001, Gang Wang 0001, Fan Zhang 0092 |
ICA3PP (2) | 4 |
| 2010 | Revisiting globally sorted indexes for efficient document retrievalabstractThere has been a large amount of research on efficient document retrieval in both IR and web search areas. One important technique to improve retrieval efficiency is early termination, which speeds up query processing by avoiding scanning the entire inverted lists. Most early termination techniques first build new inverted indexes by sorting the inverted lists in the order of either the term-dependent information, e.g., term frequencies or term IR scores, or the term-independent information, e.g., static rank of the document; and then apply appropriate retrieval strategies on the resulting indexes. Although the methods based only on the static rank have been shown to be ineffective for the early termination, there are still many advantages of using the methods based on term-independent information. In this paper, we propose new techniques to organize inverted indexes based on the term-independent information beyond static rank and study the new retrieval strategies on the resulting indexes. We perform a detailed experimental evaluation on our new techniques and compare them with the existing approaches. Our results on the TREC GOV and GOV2 data sets show that our techniques can improve query efficiency significantly. Fan Zhang 0092, Shuming Shi 0001, Ji-Rong Wen |
WSDM | 1 |
| 2008 | An Improved Parallel Implementation of 3-D DRIE Simulation on Multi-core ProcessorsabstractDeep reactive ion etching (DRIE) technique is a new and powerful tool in micro-electro-mechanical systems (MEMS) fabrication. A 3-D DRIE simulation can help researcher understand the time-evolution of Bosch process used in DRIE. Due to the high complexity of the algorithm used in the simulation, it is necessary to develop an algorithm that can speedup the simulation. This paper presents a parallel implementation of the 3-D DRIE simulation based on multi-core processor. The algorithm is based on data partition. We examine four different data partition strategies and find two-dimensional block-cyclic distribution can obtain perfect load balance. The experimental results show that the parallel algorithm obtains a substantial speedup over the serial algorithm on an Intel quad-core computer. Fan Zhang 0092, Gang Wang 0001, Xiaoguang Liu 0001, Guangyi Sun, Jing Liu 0010, Guizhang Lu |
HPCC | 1 |