Fangbo Tao

dblp:66/10768 · DBLP profile ↗
← Back
14ranked-venue papers
5as first author
0since 2021 · last 2018
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 11 · 5 first-authorArtificial intelligence and machine learning · 8 · 2 first-authorSecurity and privacy · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
11 papers
Data mining · 62% Information retrieval · 16% Knowledge graphs · 9%
Artificial intelligence
5 papers
Information extraction and text analysis · 63% Representation and self-supervised learning · 21% Graph learning · 16%

Topics — the 25 heaviest of 30, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Data mining › structured data mining › graph mining
heterogeneous information network
0.632017
Embedding Learning with Events in Heterogeneous Information Networks · IEEE Trans. Knowl. Data Eng. 2017
Research-insight: providing insight on research by publication network analysis · SIGMOD Conference 2013
AMETHYST: a system for mining and exploring topical hierarchies of heterogeneous data · KDD 2013
Data mining › representation learning
hyperedge-based embedding
0.522017
Embedding Learning with Events in Heterogeneous Information Networks · IEEE Trans. Knowl. Data Eng. 2017
Large-Scale Embedding Learning in Heterogeneous Event Data · ICDM 2016
Data mining › structured data mining › graph mining
information network analysis
0.422014
NewsNetExplorer: automatic construction and exploration of news information networks · SIGMOD Conference 2014
Research-insight: providing insight on research by publication network analysis · SIGMOD Conference 2013
Machine learning › Representation and self-supervised learning › text embedding › text representation learning
document representation
0.312018
Doc2Cube: Allocating Documents to Text Cube Without Labeled Data · ICDM 2018
Knowledge graphs
taxonomy construction
0.312018
TaxoGen: Unsupervised Topic Taxonomy Construction by Adaptive Term Embedding and Clustering · KDD 2018
Data mining
text mining
0.312018
Doc2Cube: Allocating Documents to Text Cube Without Labeled Data · ICDM 2018
Data mining
anomaly detection
0.312017
Identifying Semantically Deviating Outlier Documents · EMNLP 2017
Information retrieval
multimodal embedding
0.312017
ReAct: Online Multimodal Embedding for Recency-Aware Spatiotemporal Activity Modeling · SIGIR 2017
Data mining › structured data mining › graph mining
network embedding
0.312017
Embedding Learning with Events in Heterogeneous Information Networks · IEEE Trans. Knowl. Data Eng. 2017
Data mining
pattern mining
0.312017
Embedding Learning with Events in Heterogeneous Information Networks · IEEE Trans. Knowl. Data Eng. 2017
Machine learning › Graph learning
network embedding
0.212016
Large-Scale Embedding Learning in Heterogeneous Event Data · ICDM 2016
Data mining › structured data mining › graph mining › heterogeneous information network
heterogeneous information network mining
0.212016
Large-Scale Embedding Learning in Heterogeneous Event Data · ICDM 2016
Natural language and speech › Information extraction and text analysis › relation extraction
distant supervision
0.212015
ClusType: Effective Entity Recognition and Typing by Relation Phrase-Based Clustering · KDD 2015
Natural language and speech › Information extraction and text analysis
named entity recognition
0.212015
Leveraging Pattern Semantics for Extracting Entities in Enterprises · WWW 2015
Information retrieval › document retrieval › domain-specific retrieval
news retrieval
0.212014
NewsNetExplorer: automatic construction and exploration of news information networks · SIGMOD Conference 2014
Web and social media mining
citation network analysis
0.212013
Research-insight: providing insight on research by publication network analysis · SIGMOD Conference 2013
Data mining › structured data mining
graph mining
0.212013
AMETHYST: a system for mining and exploring topical hierarchies of heterogeneous data · KDD 2013
Information retrieval › search interfaces
search result organization
0.212013
AMETHYST: a system for mining and exploring topical hierarchies of heterogeneous data · KDD 2013
Information retrieval › document organization
topic hierarchy
0.212013
AMETHYST: a system for mining and exploring topical hierarchies of heterogeneous data · KDD 2013
Data integration and cleaning › heterogeneous data integration
unstructured and structured data integration
0.212013
EventCube: multi-dimensional search and mining of structured and text data · KDD 2013
Knowledge graphs
knowledge graph construction
0.112015
Leveraging Pattern Semantics for Extracting Entities in Enterprises · WWW 2015
Data mining › text mining
information extraction
0.112014
NewsNetExplorer: automatic construction and exploration of news information networks · SIGMOD Conference 2014
Knowledge graphs › ontology
domain ontology
0.012013
AMETHYST: a system for mining and exploring topical hierarchies of heterogeneous data · KDD 2013
Query processing and optimization
OLAP
0.012013
EventCube: multi-dimensional search and mining of structured and text data · KDD 2013
Knowledge graphs › ontology
ontology construction
0.012013
AMETHYST: a system for mining and exploring topical hierarchies of heterogeneous data · KDD 2013

Methods — techniques the papers use, named apart from their topics

weak supervision · 0.7joint embedding · 0.7similarity search · 0.4term embeddings · 0.3spherical clustering · 0.3hierarchical clustering · 0.3end-to-end framework · 0.3distant supervised learning · 0.3word embeddings · 0.3semi-supervised multimodal embedding · 0.3geographical topic models · 0.3generative model · 0.3hyperedge prediction · 0.2embedding learning · 0.2phrase mining · 0.2multi-view clustering · 0.2joint optimization · 0.2bootstrapping · 0.2
YearPublicationVenuePosition
2018 Guess Me if You Can: Acronym Disambiguation for Enterprises
abstract
Acronyms are abbreviations formed from the initial components of words or phrases.In enterprises, people often use acronyms to make communications more efficient.However, acronyms could be difficult to understand for people who are not familiar with the subject matter (new employees, etc.), thereby affecting productivity.To alleviate such troubles, we study how to automatically resolve the true meanings of acronyms in a given context.Acronym disambiguation for enterprises is challenging for several reasons.First, acronyms may be highly ambiguous since an acronym used in the enterprise could have multiple internal and external meanings.Second, there are usually no comprehensive knowledge bases such as Wikipedia available in enterprises.Finally, the system should be generic to work for any enterprise.In this work we propose an end-to-end framework to tackle all these challenges.The framework takes the enterprise corpus as input and produces a high-quality acronym disambiguation system as output.Our disambiguation models are trained via distant supervised learning, without requiring any manually labeled training examples.Therefore, our proposed framework can be deployed to any enterprise to support highquality acronym disambiguation.Experimental results on real world data justified the effectiveness of our system.
Yang Li 0150, Bo Zhao 0001, Ariel Fuxman, Fangbo Tao
ACL (1)4
2018 Doc2Cube: Allocating Documents to Text Cube Without Labeled Data
abstract
Data cube is a cornerstone architecture in multidimensional analysis of structured datasets. It is highly desirable to conduct multidimensional analysis on text corpora with cube structures for various text-intensive applications in healthcare, business intelligence, and social media analysis. However, one bottleneck to constructing text cube is to automatically put millions of documents into the right cube cells so that quality multidimensional analysis can be conducted afterwards-it is too expensive to allocate documents manually or rely on massively labeled data. We propose Doc2Cube, a method that constructs a text cube from a given text corpus in an unsupervised way. Initially, only the label names (e.g., USA, China) of each dimension (e.g., location) are provided instead of any labeled data. Doc2Cube leverages label names as weak supervision signals and iteratively performs joint embedding of labels, terms, and documents to uncover their semantic similarities. To generate joint embeddings that are discriminative for cube construction, Doc2Cube learns dimension-tailored document representations by selectively focusing on terms that are highly label-indicative in each dimension. Furthermore, Doc2Cube alleviates label sparsity by propagating the information from label names to other terms and enriching the labeled term set. Our experiments on real data demonstrate the superiority of Doc2Cube over existing methods.
Fangbo Tao, Chao Zhang 0014, Xiusi Chen, Meng Jiang 0001, Tim Hanratty, Lance M. Kaplan, Jiawei Han 0001
ICDM1
2018 TaxoGen: Unsupervised Topic Taxonomy Construction by Adaptive Term Embedding and Clustering
abstract
Taxonomy construction is not only a fundamental task for semantic analysis of text corpora, but also an important step for applications such as information filtering, recommendation, and Web search. Existing pattern-based methods extract hypernym-hyponym term pairs and then organize these pairs into a taxonomy. However, by considering each term as an independent concept node, they overlook the topical proximity and the semantic correlations among terms. In this paper, we propose a method for constructing topic taxonomies, wherein every node represents a conceptual topic and is defined as a cluster of semantically coherent concept terms. Our method, TaxoGen, uses term embeddings and hierarchical clustering to construct a topic taxonomy in a recursive fashion. To ensure the quality of the recursive process, it consists of: (1) an adaptive spherical clustering module for allocating terms to proper levels when splitting a coarse topic into fine-grained ones; (2) a local embedding module for learning term embeddings that maintain strong discriminative power at different levels of the taxonomy. Our experiments on two real datasets demonstrate the effectiveness of TaxoGen compared with baseline methods.
Chao Zhang 0014, Fangbo Tao, Xiusi Chen, Meng Jiang 0001, Brian M. Sadler, Michelle Vanni, Jiawei Han 0001
KDD2
2017 Identifying Semantically Deviating Outlier Documents
abstract
A document outlier is a document that substantially deviates in semantics from the majority ones in a corpus.Automatic identification of document outliers can be valuable in many applications, such as screening health records for medical mistakes.In this paper, we study the problem of mining semantically deviating document outliers in a given corpus.We develop a generative model to identify frequent and characteristic semantic regions in the word embedding space to represent the given corpus, and a robust outlierness measure which is resistant to noisy content in documents.Experiments conducted on two real-world textual data sets show that our method can achieve an up to 135% improvement over baselines in terms of recall at top-1% of the outlier ranking.
Honglei Zhuang, Chi Wang 0001, Fangbo Tao, Lance M. Kaplan, Jiawei Han 0001
EMNLP3
2017 ReAct: Online Multimodal Embedding for Recency-Aware Spatiotemporal Activity Modeling
abstract
Spatiotemporalactivity modeling is an important task for applications like tour recommendation and place search. The recently developed geographical topic models have demonstrated compelling results in using geo-tagged social media (GTSM) for spatiotemporal activity modeling. Nevertheless, they all operate in batch and cannot dynamically accommodate the latest information in the GTSM stream to reveal up-to-date spatiotemporal activities. We propose ReAct, a method that processes continuous GTSM streams and obtains recency-aware spatiotemporal activity models on the fly. Distinguished from existing topic-based methods, ReAct embeds all the regions, hours, and keywords into the same latent space to capture their correlations. To generate high-quality embeddings, it adopts a novel semi-supervised multimodal embedding paradigm that leverages the activity category information to guide the embedding process. Furthermore, as new records arrive continuously, it employs strategies to effectively incorporate the new information while preserving the knowledge encoded in previous embeddings. Our experiments on the geo-tagged tweet streams in two major cities have shown that ReAct significantly outperforms existing methods for location and activity retrieval tasks.
Chao Zhang 0014, Keyang Zhang, Quan Yuan 0001, Fangbo Tao, Tim Hanratty, Jiawei Han 0001
SIGIR4
2017 Embedding Learning with Events in Heterogeneous Information Networks
abstract
In real-world applications, objects of multiple types are interconnected, formingHeterogeneous Information Networks. In such heterogeneous information networks, we make the key observation that many interactions happen due to someeventand the objects in each event form a complete semantic unit. By taking advantage of such a property, we propose a generic framework calledHyperEdge-BasedEmbedding(Hebe) to learn object embeddings with events in heterogeneous information networks, where ahyperedgeencompasses the objects participating in one event. TheHebeframework models the proximity among objects in each event with two methods: (1) predicting a target object given other participating objects in the event, and (2) predicting if the event can be observed given all the participating objects. Since each hyperedge encapsulates more information of a given event,Hebeis robust to data sparseness and noise. In addition,Hebeis scalable when the data size spirals. Extensive experiments on large-scale real-world datasets show the efficacy and robustness of the proposed framework.
Huan Gui, Fangbo Tao, Meng Jiang 0001, Brandon Norick, Lance M. Kaplan, Jiawei Han 0001
IEEE Trans. Knowl. Data Eng.3
2016 Large-Scale Embedding Learning in Heterogeneous Event Data
abstract
Heterogeneous events, which are defined as events connecting strongly-typed objects, are ubiquitous in the real world. We propose a HyperEdge-Based Embedding (Hebe) framework for heterogeneous event data, where a hyperedge represents the interaction among a set of involving objects in an event. The Hebe framework models the proximity among objects in an event by predicting a target object given the other participating objects in the event (hyperedge). Since each hyperedge encapsulates more information on a given event, Hebe is robust to data sparseness. In addition, Hebe is scalable when the data size spirals. Extensive experiments on large-scale real-world datasets demonstrate the efficacy and robustness of Hebe.
Huan Gui, Fangbo Tao, Meng Jiang 0001, Brandon Norick, Jiawei Han 0001
ICDM3
2015 ClusType: Effective Entity Recognition and Typing by Relation Phrase-Based Clustering
abstract
Entity recognition is an important but challenging research problem. In reality, many text collections are from specific, dynamic, or emerging domains, which poses significant new challenges for entity recognition with increase in name ambiguity and context sparsity, requiring entity detection without domain restriction. In this paper, we investigate entity recognition (ER) with distant-supervision and propose a novel relation phrase-based ER framework, called ClusType, that runs data-driven phrase mining to generate entity mention candidates and relation phrases, and enforces the principle that relation phrases should be softly clustered when propagating type information between their argument entities. Then we predict the type of each entity mention based on the type signatures of its co-occurring relation phrases and the type indicators of its surface name, as computed over the corpus. Specifically, we formulate a joint optimization problem for two tasks, type propagation with relation phrases and multi-view relation phrase clustering. Our experiments on multiple genres---news, Yelp reviews and tweets---demonstrate the effectiveness and robustness of ClusType, with an average of 37% improvement in F1 score over the best compared method.
Xiang Ren 0001, Ahmed El-Kishky, Chi Wang 0001, Fangbo Tao, Clare R. Voss, Jiawei Han 0001
KDD4
2015 Leveraging Pattern Semantics for Extracting Entities in Enterprises
abstract
Entity Extraction is a process of identifying meaningful entities from text documents. In enterprises, extracting entities improves enterprise efficiency by facilitating numerous applications, including search, recommendation, etc. However, the problem is particularly challenging on enterprise domains due to several reasons. First, the lack of redundancy of enterprise entities makes previous web-based systems like NELL and OpenIE not effective, since using only high-precision/low-recall patterns like those systems would miss the majority of sparse enterprise entities, while using more low-precision patterns in sparse setting also introduces noise drastically. Second, semantic drift is common in enterprises ("Blue" refers to "Windows Blue"), such that public signals from the web cannot be directly applied on entities. Moreover, many internal entities never appear on the web. Sparse internal signals are the only source for discovering them. To address these challenges, we propose an end-to-end framework for extracting entities in enterprises, taking the input of enterprise corpus and limited seeds to generate a high-quality entity collection as output. We introduce the novel concept of Semantic Pattern Graph to leverage public signals to understand the underlying semantics of lexical patterns, reinforce pattern evaluation using mined semantics, and yield more accurate and complete entities. Experiments on Microsoft enterprise data show the effectiveness of our approach.
Fangbo Tao, Bo Zhao 0001, Ariel Fuxman, Yang Li 0150, Jiawei Han 0001
WWW1
2014 NewsNetExplorer: automatic construction and exploration of news information networks
abstract
News data is one of the most abundant and familiar data sources. News data can be systematically utilized and ex- plored by database, data mining, NLP and information re- trieval researchers to demonstrate to the general public the power of advanced information technology. In our view, news data contains rich, inter-related and multi-typed data objects, forming one or a set of gigantic, interconnected, het- erogeneous information networks. Much knowledge can be derived and explored with such an information network if we systematically develop effective and scalable data-intensive information network analysis technologies. By further developing a set of information extraction, in- formation network construction, and information network mining methods, we extract types, topical hierarchies and other semantic structures from news data, construct a semi- structured news information network NewsNet. Further, we develop a set of news information network exploration and mining mechanisms that explore news in multi-dimensional space, which include (i) OLAP-based operations on the hierarchical dimensional and topical structures and rich-text, such as cell summary, single dimension analysis, and promo- tion analysis, (ii) a set of network-based operations, such as similarity search and ranking-based clustering, and (iii) a set of hybrid operations or network-OLAP operations, such as entity ranking at different granularity levels. These form the basis of our proposed NewsNetExplorer system. Although some of these functions have been studied in recent research, effective and scalable realization of such functions in large networks still poses multiple challenging research problems. Moreover, some functions are our on-going research tasks. By integrating these functions, NewsNetExplorer not only provides with us insightful recommendations in NewsNet exploration system but also helps us gain insight on how to perform effective information extraction, integration and mining in large unstructured datasets.
Fangbo Tao, George Brova, Jiawei Han 0001, Heng Ji 0001, Chi Wang 0001, Brandon Norick, Ahmed El-Kishky, Xiang Ren 0001, Yizhou Sun
SIGMOD Conference1
2013 AMETHYST: a system for mining and exploring topical hierarchies of heterogeneous data
abstract
In this demo we present AMETHYST, a system for exploring and analyzing a topical hierarchy constructed from a heterogeneous information network (HIN). HINs, composed of multiple types of entities and links are very common in the real world. Many have a text component, and thus can benefit from a high quality hierarchical organization of the topics in the network dataset. By organizing the topics into a hierarchy, AMETHYST helps understand search results in the context of an ontology, and explain entity relatedness at different granularities. The automatically constructed topical hierarchy reflects a domain-specific ontology, interacts with multiple types of linked entities, and can be tailored for both free text and OLAP queries.
Marina Danilevsky, Chi Wang 0001, Fangbo Tao, Nihit Desai, Jiawei Han 0001
KDD3
2013 EventCube: multi-dimensional search and mining of structured and text data
abstract
A large portion of real world data is either text or structured (e.g., relational) data. Moreover, such data objects are often linked together (e.g., structured specification of products linking with the corresponding product descriptions and customer comments). Even for text data such as news data, typed entities can be extracted with entity extraction tools. The EventCube project constructs TextCube and TopicCube from interconnected structured and text data (or from text data via entity extraction and dimension building), and performs multidimensional search and analysis on such datasets, in an informative, powerful, and user-friendly manner. This proposed EventCube demo will show the power of the system not only on the originally designed ASRS (Aviation Safety Report System) data sets, but also on news datasets collected from multiple news agencies, and academic datasets constructed from the DBLP and web data. The system has high potential to be extended in many powerful ways and serve as a general platform for search, OLAP (online analytical processing) and data mining on integrated text and structured data. After the system demo in the conference, the system will be put on the web for public access and evaluation.
Fangbo Tao, Kin Hou Lei, Jiawei Han 0001, ChengXiang Zhai, Marina Danilevsky, Nihit Desai, Bolin Ding, Heng Ji 0001, Rucha Kanade, Anne Kao, Qi Li 0014, Yanen Li, Cindy Xide Lin, Nikunj C. Oza, Ashok N. Srivastava, Rodney Tjoelker, Chi Wang 0001, Duo Zhang 0001, Bo Zhao 0001
KDD1
2013 Research-insight: providing insight on research by publication network analysis
abstract
A database contains rich, inter-related, multi-typed data and information, forming one or a set of gigantic, intercon- nected, heterogeneous information networks. Much knowl- edge can be derived from such information networks if we systematically develop an effective and scalable database-oriented information network analysis technology. In this system demo, we take a computer science research publica- tion network as an example, which is an information net- work derived from an integration of DBLP, other web-based information about researchers, and partially available cita- tion data, and construct a Research-Insight system in order to demonstrate the power of database-oriented information network analysis. We show that nontrivial research insight can be obtained from such analysis, including (1) ranking, clustering, classification and similarity search of researchers, terms and venues for research subfields and themes, (2) recommending good researchers and good research papers to read or cite when conducting research on certain topics (3) predicting potential collaborators for certain theme-oriented research, and (4) predicting advisor-advisee rela- tionships and affiliation history based on historical research publications. Although some of these functions have been studied in recent research, effective and scalable realization of such functions in large networks still poses challenging research problems. Moreover, some function are our on- going research tasks. By integrating these functionalities, Research-Insight may not only provide with us insightful rec- ommendations in CS research but also help us gain insight on how to perform effective data mining in large databases.
Fangbo Tao, Xiao Yu 0007, Kin Hou Lei, George Brova, Jiawei Han 0001, Rucha Kanade, Yizhou Sun, Chi Wang 0001, Tim Weninger
SIGMOD Conference1
2011 SkyBoundary: An Improved Approach to Member Promotion in Social Networks
abstract
With the rapid development of Social Network (SN for short), people increasingly pay attention to the importance of the roles which they play in the SNs. As is usually the case, the standard for measuring the importance of the members is multi-objective. The skyline operator is thus introduced to distinguish the important members from the entire community. For decision-making, people are interested in the most potential stars which can be promoted into the skyline with minimum cost, namely the problem of Member Promotion in Social Networks. In this paper, based on the characteristic of the skyline operator and the promotion process, we first of all propose some interesting new concepts such as Promotion Boundary to design a novel promotion boundary-based pruning strategy. After that, we bring forward an effective cost-based pruning strategy on the basis of permutation and combination theories to verify the plans in the ascending order of cost. The Sky Boundary algorithm is therefore proposed to solve the problem effectively by employing the optimization strategies. Extensive experiments on both real and synthetic datasets are conducted to show the application value, effectiveness and efficiency of the Sky Boundary algorithm.
Zhuo Peng, Chaokun Wang, Fangbo Tao
DASC3