Mianwei Zhou

dblp:18/5844 · DBLP profile ↗
← Back
9ranked-venue papers
5as first author
0since 2021 · last 2018
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 8 · 5 first-authorArtificial intelligence and machine learning · 6 · 3 first-authorApplied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
5 papers
Information retrieval · 76% Query processing and optimization · 11% Data models and query languages · 8%
Artificial intelligence
2 papers
Information extraction and text analysis · 100%

Topics — the 17 heaviest of 20, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Information retrieval
ranking
0.422016
Ranking Relevance in Yahoo Search · KDD 2016
Unifying learning to rank and domain adaptation: enabling cross-task document scoring · KDD 2014
Information retrieval › ranking
learning to rank
0.422014
Unifying learning to rank and domain adaptation: enabling cross-task document scoring · KDD 2014
Learning to rank from distant supervision: Exploiting noisy redundancy for relational entity search · ICDE 2013
Natural language and speech › Information extraction and text analysis › sentiment analysis › aspect-based sentiment analysis
aspect-level sentiment classification
0.312018
Target-Sensitive Memory Networks for Aspect Sentiment Classification · ACL (1) 2018
Natural language and speech › Information extraction and text analysis
sentiment analysis
0.312018
Target-Sensitive Memory Networks for Aspect Sentiment Classification · ACL (1) 2018
Information retrieval › ranking › context-aware ranking
location-based ranking
0.212016
Ranking Relevance in Yahoo Search · KDD 2016
Query processing and optimization
query rewriting
0.212016
Ranking Relevance in Yahoo Search · KDD 2016
Information retrieval › ranking › context-aware ranking
recency ranking
0.212016
Ranking Relevance in Yahoo Search · KDD 2016
Information retrieval › ranking › search ranking
relevance ranking
0.212016
Ranking Relevance in Yahoo Search · KDD 2016
Information retrieval
retrieval models
0.212016
Ranking Relevance in Yahoo Search · KDD 2016
Information retrieval
search engines
0.212016
Ranking Relevance in Yahoo Search · KDD 2016
Information retrieval › ranking › transfer ranking
ranking model adaptation
0.212014
Unifying learning to rank and domain adaptation: enabling cross-task document scoring · KDD 2014
Knowledge graphs
distant supervision
0.212013
Learning to rank from distant supervision: Exploiting noisy redundancy for relational entity search · ICDE 2013
Information retrieval › search engines › semantic search
entity retrieval
0.212013
Learning to rank from distant supervision: Exploiting noisy redundancy for relational entity search · ICDE 2013
Information retrieval
document search
0.112010
DoCQS: a prototype system for supporting data-oriented content query · SIGMOD Conference 2010
Data models and query languages
query language
0.112010
DoCQS: a prototype system for supporting data-oriented content query · SIGMOD Conference 2010
Natural language and speech › Information extraction and text analysis
web information extraction
0.012010
Data-oriented content query system: searching for data into text on the web · WSDM 2010
Data models and query languages
relational model
0.012010
DoCQS: a prototype system for supporting data-oriented content query · SIGMOD Conference 2010

Methods — techniques the papers use, named apart from their topics

memory network · 0.3attention mechanism · 0.3semantic matching · 0.2query rewriting · 0.2query processing algorithms · 0.2tree-structured boltzmann machine · 0.2feature decoupling · 0.2distant supervision · 0.2probabilistic graphical model · 0.2pattern-based filtering · 0.2inverted index · 0.2index structures · 0.1index structure · 0.1
YearPublicationVenuePosition
2018 Target-Sensitive Memory Networks for Aspect Sentiment Classification
abstract
Aspect sentiment classification (ASC) is a fundamental task in sentiment analysis.Given an aspect/target and a sentence, the task classifies the sentiment polarity expressed on the target in the sentence.Memory networks (MNs) have been used for this task recently and have achieved state-of-the-art results.In MNs, attention mechanism plays a crucial role in detecting the sentiment context for the given target.However, we found an important problem with the current MNs in performing the ASC task.Simply improving the attention mechanism will not solve it.The problem is referred to as target-sensitive sentiment, which means that the sentiment polarity of the (detected) context is dependent on the given target and it cannot be inferred from the context alone.To tackle this problem, we propose the targetsensitive memory networks (TMNs).Several alternative techniques are designed for the implementation of TMNs and their effectiveness is experimentally evaluated.
Shuai Wang 0020, Sahisnu Mazumder, Bing Liu 0001, Mianwei Zhou, Yi Chang 0001
ACL (1)4
2016 HEER: Heterogeneous graph embedding for emerging relation detection from news
abstract
Real-world knowledge is growing rapidly nowadays. New entities arise with time, resulting in large volumes of relations that do not exist in current knowledge graphs (KGs). These relations containing at least one new entity are called emerging relations. They often appear in news, and hence the latest information about new entities and relations can be learned from news timely. In this paper, we focus on the problem of discovering emerging relations from news. However, there are several challenges for this task: (1) at the beginning, there is little information for emerging relations, causing problems for traditional sentence-based models; (2) no negative relations exist in KGs, creating difficulties in utilizing only positive cases for emerging relation detection from news; and (3) new relations emerge rapidly, making it necessary to keep KGs up to date with the latest emerging relations. In order to address these issues, we start from a global graph perspective and propose a novel Heterogeneous graph Embedding framework for Emerging Relation detection (HEER) that learns a classifier from positive and unlabeled instances by utilizing information from both news and KGs. Furthermore, we implement HEER in an incremental manner to timely update KGs with the latest detected emerging relations. Extensive experiments on real-world news datasets demonstrate the effectiveness of the proposed HEER model.
Chun-Ta Lu, Mianwei Zhou, Sihong Xie, Yi Chang 0001, Philip S. Yu
IEEE BigData3
2016 Ranking Relevance in Yahoo Search
abstract
Search engines play a crucial role in our daily lives. Relevance is the core problem of a commercial search engine. It has attracted thousands of researchers from both academia and industry and has been studied for decades. Relevance in a modern search engine has gone far beyond text matching, and now involves tremendous challenges. The semantic gap between queries and URLs is the main barrier for improving base relevance. Clicks help provide hints to improve relevance, but unfortunately for most tail queries, the click information is too sparse, noisy, or missing entirely. For comprehensive relevance, the recency and location sensitivity of results is also critical. In this paper, we give an overview of the solutions for relevance in the Yahoo search engine. We introduce three key techniques for base relevance -- ranking functions, semantic matching features and query rewriting. We also describe solutions for recency sensitive relevance and location sensitive relevance. This work builds upon 20 years of existing efforts on Yahoo search, summarizes the most recent advances and provides a series of practical relevance solutions. The performance reported is based on Yahoo's commercial search engine, where tens of billions of urls are indexed and served by the ranking system.
Dawei Yin 0001, Yuening Hu, Jiliang Tang, Tim Daly Jr., Mianwei Zhou, Hua Ouyang, Changsung Kang, Hongbo Deng, Chikashi Nobata, Jean-Marc Langlois, Yi Chang 0001
KDD5
2014 Unifying learning to rank and domain adaptation: enabling cross-task document scoring
abstract
For document scoring, although learning to rank and domain adaptation are treated as two different problems in previous works, we discover that they actually share the same challenge of adapting keyword contribution across different queries or domains. In this paper, we propose to study the cross-task document scoring problem, where a task refers to a query to rank or a domain to adapt to, as the first attempt to unify these two problems. Existing solutions for learning to rank and domain adaptation either leave the heavy burden of adapting keyword contribution to feature designers, or are difficult to be generalized. To resolve such limitations, we abstract the keyword scoring principle, pointing out that the contribution of a keyword essentially depends on, first, its importance to a task and, second, its importance to the document. For determining these two aspects of keyword importance, we further propose the concept of feature decoupling, suggesting using two types of easy-to-design features: meta-features and intra-features. Towards learning a scorer based on the decoupled features, we require that our framework fulfill inferred sparsity to eliminate the interference of noisy keywords, and employ distant supervision to tackle the lack of keyword labels. We propose the Tree-structured Boltzmann Machine (T-RBM), a novel two-stage Markov Network, as our solution. Experiments on three different applications confirm the effectiveness of T-RBM, which achieves significant improvement compared with four state-of-the-art baseline methods.
Mianwei Zhou, Kevin Chen-Chuan Chang
KDD1
2013 Entity-centric document filtering: boosting feature mapping through meta-features
abstract
This paper studies the entity-centric document filtering task -- given an entity represented by its identification page (e.g., an Wikpedia page), how to correctly identify its relevant documents. In particular, we are interested in learning an entity-centric document filter based on a small number of training entities, and the filter can predict document relevance for a large set of unseen entities at query time. Towards characterizing the relevance of a document, the problem boils down to learning keyword importance for the query entities. Since the same keyword will have very different importance for different entities, we abstract the entity-centric document filtering problem as a transfer learning problem, and the challenge becomes how to appropriately transfer the keyword importance learned from training entities to query entities. Based on the insight that keywords sharing some similar "properties" should have similar importance for their respective entities, we propose a novel concept of meta-feature to map keywords from different entities. To realize the idea of meta-feature-based feature mapping, we develop and contrast two different models, LinearMapping and BoostMapping. Experiments on three different datasets confirm the effectiveness of our proposed models, which show significant improvement compared with four state-of-the-art baseline methods.
Mianwei Zhou, Kevin Chen-Chuan Chang
CIKM1
2013 Learning to rank from distant supervision: Exploiting noisy redundancy for relational entity search
abstract
In this paper, we study the task of relational entity search which aims at automatically learning an entity ranking function for a desired relation. To rank entities, we exploit the redundancy abound in their snippets; however, such redundancy is noisy as not all the snippets represent information relevant to the desired relation. To explore useful information from such noisy redundancy, we abstract the task as a distantly supervised ranking problem — based on coarse entity-level annotations, deriving a relation-specific ranking function for the purpose of online searching. As the key challenge, without detailed snippet-level annotations, we have to learn an entity ranking function that can effectively filter noise; furthermore, the ranking function should also be online executable. We develop Pattern-based Filter Network (PFNet), a novel probabilistic graphical model, as our solution. To balance the accuracy and efficiency requirements, PFNet selects a limited size of indicative patterns to filter noisy snippets, and inverted indexes are utilized to retrieve required features. Experiments on the large scale CuleWeb09 data set for six different relations confirm the effectiveness of the proposed PFNet model, which outperforms five state-of-the-art relational entity ranking methods.
Mianwei Zhou, Hongning Wang, Kevin Chen-Chuan Chang
ICDE1
2012 Citation Prediction in Heterogeneous Bibliographic Networks
abstract
To reveal information hiding in link space of bibliographical networks, link analysis has been studied from different perspectives in recent years. In this paper, we address a novel problem namely citation prediction, that is: given information about authors, topics, target publication venues as well as time of certain research paper, finding and predicting the citation relationship between a query paper and a set of previous papers. Considering the gigantic size of relevant papers, the loosely connected citation network structure as well as the highly skewed citation relation distribution, citation prediction is more challenging than other link prediction problems which have been studied before. By building a meta-path based prediction model on a topic discriminative search space, we here propose a two-phase citation probability learning approach, in order to predict citation relationship effectively and efficiently. Experiments are performed on real-world dataset with comprehensive measurements, which demonstrate that our framework has substantial advantages over commonly used link prediction approaches in predicting citation relations in bibliographical networks.
Xiao Yu 0007, Quanquan Gu, Mianwei Zhou, Jiawei Han 0001
SDM3
2010 DoCQS: a prototype system for supporting data-oriented content query
abstract
Witnessing the richness of data in document content and many ad-hoc efforts for finding such data, we propose a Data-oriented Content Query System(DoCQS), which is oriented towards fine granularity data of all types by searching directly into document content. DoCQS uses the relational model as the underlying data model, and offers a powerful and flexible Content Query Language(CQL) to adapt to diverse query demands. In this demonstration, we show how to model various search tasks by CQL statements, and how the system architecture efficiently supports the CQL execution. Our online demo of the system is available at http://wisdm.cs.uiuc.edu/demos/docqs/.
Mianwei Zhou, Tao Cheng 0001, Kevin Chen-Chuan Chang
SIGMOD Conference1
2010 Data-oriented content query system: searching for data into text on the web
abstract
As the Web provides rich data embedded in the immense contents inside pages, we witness many ad-hoc efforts for exploiting fine granularity information across Web text, such as Web information extraction, typed-entity search, and question answering. To unify and generalize these efforts, this paper proposes a general search system--Data-oriented Content Query System(DoCQS)--to search directly into document contents for finding relevant values of desired data types. Motivated by the current limitations, we start by distilling the essential capabilities needed by such content querying. The capabilities call for a conceptually relational model, upon which we design a powerful Content Query Language (CQL). For efficient processing, we design novel index structures and query processing algorithms. We evaluate our proposal over two concrete domains of realistic Web corpora, demonstrating that our query language is rather flexible and expressive, and our query processing is efficient with reasonable index overhead.
Mianwei Zhou, Tao Cheng 0001, Kevin Chen-Chuan Chang
WSDM1