Fan Zhang 0092

dblp:21/3626-92 · DBLP profile ↗
← Back
13ranked-venue papers
3as first author
2since 2021 · last 2024
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 8 · 1 first-authorArtificial intelligence and machine learning · 5 · 2 first-authorSystems, architecture and hardware · 2 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
5 papers
Data mining · 51% Information retrieval · 39% Web and social media mining · 10%
Artificial intelligence
1 paper
Information extraction and text analysis · 50% Knowledge representation and reasoning · 50%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Computational finance and economics · 100%

Topics — the 14 heaviest of 17, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Data mining
anomaly detection
0.312017
Mining Suspicious Tax Evasion Groups in Big Data · ICDE 2017
Data mining › pattern mining
graph pattern mining
0.312017
Mining Suspicious Tax Evasion Groups in Big Data · ICDE 2017
Data mining › anomaly detection
graph anomaly detection
0.212016
Mining Suspicious Tax Evasion Groups in Big Data · IEEE Trans. Knowl. Data Eng. 2016
Data mining
pattern mining
0.212016
Mining Suspicious Tax Evasion Groups in Big Data · IEEE Trans. Knowl. Data Eng. 2016
Web and social media mining
user identity linkage
0.212013
What's in a name?: an unsupervised approach to link users across communities · WSDM 2013
Natural language and speech › Information extraction and text analysis
relation extraction
0.112011
Nonlinear Evidence Fusion and Propagation for Hyponymy Relation Mining · ACL 2011
Information retrieval › indexing
index compression
0.112011
Efficient Parallel Lists Intersection and Index Compression Algorithms using Graphics Processing Units · Proc. VLDB Endow. 2011
Information retrieval › indexing › index compression
inverted index compression
0.112011
Efficient Parallel Lists Intersection and Index Compression Algorithms using Graphics Processing Units · Proc. VLDB Endow. 2011
Information retrieval › query processing
web query processing
0.112011
Efficient Parallel Lists Intersection and Index Compression Algorithms using Graphics Processing Units · Proc. VLDB Endow. 2011
Information retrieval
document retrieval
0.112010
Revisiting globally sorted indexes for efficient document retrieval · WSDM 2010
Information retrieval › query processing
early termination
0.112010
Revisiting globally sorted indexes for efficient document retrieval · WSDM 2010
Information retrieval
indexing
0.112010
Revisiting globally sorted indexes for efficient document retrieval · WSDM 2010
Information retrieval › indexing
inverted index
0.112010
Revisiting globally sorted indexes for efficient document retrieval · WSDM 2010
GPUs and heterogeneous computing
GPU query processing
0.012011
Efficient Parallel Lists Intersection and Index Compression Algorithms using Graphics Processing Units · Proc. VLDB Endow. 2011

Methods — techniques the papers use, named apart from their topics

pattern tree matching · 0.8graph mining · 0.5graph-based method · 0.3colored network model · 0.3linear regression · 0.2hash segmentation · 0.2d-gap compression · 0.2binary search · 0.2unsupervised learning · 0.2n-gram probability · 0.2nonlinear evidence fusion · 0.1
YearPublicationVenuePosition
2024 SkillNet-X: A Multilingual Multitask Model with Sparsely Activated Skills
abstract
Traditional multitask learning methods typically can only leverage shared knowledge within specific tasks or languages, resulting in a loss of either cross-language or cross-task knowledge. This paper proposes a general multilingual multitask model, named SkillNet-X, which enables a single model to tackle many different tasks from different languages. To this end, we define several language-specific skills and task-specific skills, each of which corresponds to a skill module. SkillNet-X sparsely activates parts of the skill modules which are relevant to eitherthe target task or the target language. Acting as knowledge transit hubs, skill modules are capable of absorbing task-related knowledge and language-related knowledge consecutively. We evaluate SkillNet-X on eleven natural language understanding datasets in four languages. Results show that SkillNet-X performs better than task-specific and two multitask learning baselines.To investigate the generalization of our model, we conduct experiments on two new tasks and find that SkillNet-X significantly outperforms baselines.
Zhangyin Feng, Yong Dai 0001, Fan Zhang 0092, Duyu Tang, Shuangzhi Wu, Bing Qin 0001, Yunbo Cao, Shuming Shi 0001
ICASSP3
2023 Skillnet-NLG: General-Purpose Natural Language Generation with a Sparsely Activated Approach
abstract
We present SkillNet-NLG, a sparsely activated approach that handles many natural language generation tasks with one model. Different from traditional dense models that always activate all the parameters, SkillNet-NLG selectively activates relevant parts of the parameters to accomplish a task, where the relevance is controlled by a set of predefined skills. The strength of such model design is that it provides an opportunity to precisely adapt relevant skills to learn new tasks effectively. We evaluate on Chinese natural language generation tasks. Results show that, with only one model file, SkillNet-NLG outperforms previous best performance methods on four of five tasks. SkillNet-NLG performs better than two multitask learning baselines (a dense model and a Mixture-of-Expert model) and achieves comparable performance to task-specific models. Lastly, SkillNet-NLG surpasses baseline systems when adapted to new tasks.
Junwei Liao, Duyu Tang, Fan Zhang 0092, Shuming Shi 0001
ICASSP3
2017 Mining Suspicious Tax Evasion Groups in Big Data
abstract
There is evidence that an increasing number of enterprises plot together to evade tax in an unperceived way. At the same time, the taxation information related data is a classic kind of big data. These issues challenge the effectiveness of traditional data mining-based tax evasion detection methods. To address this problem, we first investigate the classic tax evasion cases, and employ a graph-based method to characterize their property that describes two suspicious relationship trails with a same antecedent node behind an Interest-Affiliated Transaction (IAT). Next, we propose a Colored Network-Based Model (CNBM) for characterizing economic behaviors, social relationships and the IATs between taxpayers, and generating a Taxpayer Interest Interacted Network (TPIIN). To accomplish the tax evasion detection task by discovering suspicious groups in a TPIIN, methods for building a patterns tree and matching component patterns are introduced and the completeness of the methods based on graph theory is presented. Then, we describe an experiment based on real data and a simulated network. The experimental results show that our proposed method greatly improves the efficiency of tax evasion detection, as well as provides a clear explanation of the tax evasion behaviors of taxpayer groups.
Feng Tian 0002, Kuo-Ming Chao, Nick Godwin, Nazaraf Shah, Fan Zhang 0092
ICDE7
2016 Mining Suspicious Tax Evasion Groups in Big Data
abstract
There is evidence that an increasing number of enterprises plot together to evade tax in an unperceived way. At the same time, the taxation information related data is a classic kind of big data. These issues challenge the effectiveness of traditional data mining-based tax evasion detection methods. To address this problem, we first investigate the classic tax evasion cases, and employ a graph-based method to characterize their property that describes two suspicious relationship trails with a same antecedent node behind an Interest-Affiliated Transaction (IAT). Next, we propose a Colored Network-Based Model (CNBM) for characterizing economic behaviors, social relationships, and the IATs between taxpayers, and generating a Taxpayer Interest Interacted Network (TPIIN). To accomplish the tax evasion detection task by discovering suspicious groups in a TPIIN, methods for building a patterns tree and matching component patterns are introduced and the completeness of the methods based on graph theory is presented. Then, we describe an experiment based on real data and a simulated network. The experimental results show that our proposed method greatly improves the efficiency of tax evasion detection, as well as provides a clear explanation of the tax evasion behaviors of taxpayer groups.
Feng Tian 0002, Kuo-Ming Chao, Nick Godwin, Nazaraf Shah, Fan Zhang 0092
IEEE Trans. Knowl. Data Eng.7
2013 Exploiting query term correlation for list caching in web search engines
abstract
Caching technologies have been widely employed to boost the performance of Web search engines. Motivated by the correlation between terms in query logs from a commercial search engine, we explore the idea of a caching scheme based on pairs of terms, rather than individual terms (which is the typical approach used by search engines today). We propose an inverted list caching policy, based on the Least Recently Used method, in which the co-occurring correlation between terms in the query stream is accounted for when deciding on which terms to keep in the cache. We consider not only the term co-occurrence within the same query but also the co-occurrence between separate queries. Experimental results show that the proposed approach can improve not only the cache hit ratio but also the overall throughput of the system when compared to existing list caching algorithms.
Jiancong Tong, Gang Wang 0001, Douglas S. Stones, Shizhao Sun, Xiaoguang Liu 0001, Fan Zhang 0092
CIKM6
2013 What's in a name?: an unsupervised approach to link users across communities
abstract
In this paper, we consider the problem of linking users across multiple online communities. Specifically, we focus on the alias-disambiguation step of this user linking task, which is meant to differentiate users with the same usernames. We start quantitatively analyzing the importance of the alias-disambiguation step by conducting a survey on 153 volunteers and an experimental analysis on a large dataset of About.me (75,472 users). The analysis shows that the alias-disambiguation solution can address a major part of the user linking problem in terms of the coverage of true pairwise decisions (46.8%). To the best of our knowledge, this is the first study on human behaviors with regards to the usages of online usernames. We then cast the alias-disambiguation step as a pairwise classification problem and propose a novel unsupervised approach. The key idea of our approach is to automatically label training instances based on two observations: (a) rare usernames are likely owned by a single natural person, e.g. pennystar88 as a positive instance; (b) common usernames are likely owned by different natural persons, e.g. tank as a negative instance. We propose using the n-gram probabilities of usernames to estimate the rareness or commonness of usernames. Moreover, these two observations are verified by using the dataset of Yahoo! Answers. The empirical evaluations on 53 forums verify: (a) the effectiveness of the classifiers with the automatically generated training data and (b) that the rareness and commonness of usernames can help user linking. We also analyze the cases where the classifiers fail.
Jing Liu 0022, Fan Zhang 0092, Xinying Song, Young-In Song, Chin-Yew Lin, Hsiao-Wuen Hon
WSDM2
2013 Mining subtopics from text fragments for a web query
Qinglei Wang, Ya-nan Qian, Ruihua Song, Zhicheng Dou, Fan Zhang 0092, Tetsuya Sakai
Inf. Retr.5
2011 Nonlinear Evidence Fusion and Propagation for Hyponymy Relation Mining
Fan Zhang 0092, Shuming Shi 0001, Jing Liu 0022, Shu-Qi Sun, Chin-Yew Lin
ACL1
2011 Efficient Parallel Lists Intersection and Index Compression Algorithms using Graphics Processing Units
abstract
Major web search engines answer thousands of queries per second requesting information about billions of web pages. The data sizes and query loads are growing at an exponential rate. To manage the heavy workload, we consider techniques for utilizing a Graphics Processing Unit (GPU). We investigate new approaches to improve two important operations of search engines -- lists intersection and index compression. For lists intersection, we develop techniques for efficient implementation of the binary search algorithm for parallel computation. We inspect some representative real-world datasets and find that a sufficiently long inverted list has an overall linear rate of increase. Based on this observation, we propose Linear Regression and Hash Segmentation techniques for contracting the search range. For index compression, the traditional d-gap based compression schemata are not well-suited for parallel computation, so we propose a Linear Regression Compression schema which has an inherent parallel structure. We further discuss how to efficiently intersect the compressed lists on a GPU. Our experimental results show significant improvements in the query processing throughput on several datasets.
Naiyong Ao, Fan Zhang 0092, Di Wu 0036, Douglas S. Stones, Gang Wang 0001, Xiaoguang Liu 0001, Jing Liu 0010, Sheng Lin 0002
Proc. VLDB Endow.2
2010 Efficient term proximity search with term-pair indexes
abstract
There has been a large amount of research on early termination techniques in web search and information retrieval. Such techniques return the top-k documents without scanning and evaluating the full inverted lists of the query terms. Thus, they can greatly improve query processing efficiency. However, only a limited amount of efficient top-k processing work considers the impact of term proximity, i.e., the distance between term occurrences in a document, which has recently been integrated into a number of retrieval models to improve effectiveness.
Shuming Shi 0001, Fan Zhang 0092, Torsten Suel, Ji-Rong Wen
CIKM3
2010 An Improved Parallel MEMS Processing-Level Simulation Implementation Using Graphic Processing Unit
Yupeng Guo, Xiaoguang Liu 0001, Gang Wang 0001, Fan Zhang 0092
ICA3PP (2)4
2010 Revisiting globally sorted indexes for efficient document retrieval
abstract
There has been a large amount of research on efficient document retrieval in both IR and web search areas. One important technique to improve retrieval efficiency is early termination, which speeds up query processing by avoiding scanning the entire inverted lists. Most early termination techniques first build new inverted indexes by sorting the inverted lists in the order of either the term-dependent information, e.g., term frequencies or term IR scores, or the term-independent information, e.g., static rank of the document; and then apply appropriate retrieval strategies on the resulting indexes. Although the methods based only on the static rank have been shown to be ineffective for the early termination, there are still many advantages of using the methods based on term-independent information. In this paper, we propose new techniques to organize inverted indexes based on the term-independent information beyond static rank and study the new retrieval strategies on the resulting indexes. We perform a detailed experimental evaluation on our new techniques and compare them with the existing approaches. Our results on the TREC GOV and GOV2 data sets show that our techniques can improve query efficiency significantly.
Fan Zhang 0092, Shuming Shi 0001, Ji-Rong Wen
WSDM1
2008 An Improved Parallel Implementation of 3-D DRIE Simulation on Multi-core Processors
abstract
Deep reactive ion etching (DRIE) technique is a new and powerful tool in micro-electro-mechanical systems (MEMS) fabrication. A 3-D DRIE simulation can help researcher understand the time-evolution of Bosch process used in DRIE. Due to the high complexity of the algorithm used in the simulation, it is necessary to develop an algorithm that can speedup the simulation. This paper presents a parallel implementation of the 3-D DRIE simulation based on multi-core processor. The algorithm is based on data partition. We examine four different data partition strategies and find two-dimensional block-cyclic distribution can obtain perfect load balance. The experimental results show that the parallel algorithm obtains a substantial speedup over the serial algorithm on an Intel quad-core computer.
Fan Zhang 0092, Gang Wang 0001, Xiaoguang Liu 0001, Guangyi Sun, Jing Liu 0010, Guizhang Lu
HPCC1