VLDB 2026 Research / reviewers in the wild / expert
Minh C. Phan
dblp:204/0140
· DBLP profile ↗
12ranked-venue papers
6as first author
0since 2021 · last 2020
0000-0002-5407-2240ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 9 · 5 first-authorArtificial intelligence and machine learning · 5 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
5 papers |
Knowledge graphs · 41% Data mining · 31% Data integration and cleaning · 17% | |
| Artificial intelligence
5 papers |
Language models and text generation · 26% Question answering and dialogue systems · 24% Deep learning architectures and training · 21% | |
| Human-computer interaction and pervasive computing
1 paper |
Learning and educational technologies · 100% |
Topics — the 19 heaviest of 22, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Knowledge graphs
entity linking |
1.0 | 3 | 2019 | Pair-Linking for Collective Entity Disambiguation: Two Could Be Better Than All · IEEE Trans. Knowl. Data Eng. 2019 Linking Fine-Grained Locations in User Comments · IEEE Trans. Knowl. Data Eng. 2018 CoNEREL: Collective Information Extraction in News Articles · SIGIR 2018 |
Knowledge graphs › entity linking
collective entity linking |
0.7 | 2 | 2019 | Pair-Linking for Collective Entity Disambiguation: Two Could Be Better Than All · IEEE Trans. Knowl. Data Eng. 2019 Linking Fine-Grained Locations in User Comments · IEEE Trans. Knowl. Data Eng. 2018 |
Data integration and cleaning
entity disambiguation |
0.7 | 2 | 2019 | Pair-Linking for Collective Entity Disambiguation: Two Could Be Better Than All · IEEE Trans. Knowl. Data Eng. 2019 Linking Fine-Grained Locations in User Comments · IEEE Trans. Knowl. Data Eng. 2018 |
Natural language and speech › Question answering and dialogue systems
machine reading comprehension |
0.4 | 1 | 2019 | Simple and Effective Curriculum Pointer-Generator Networks for Reading Comprehension over Long Narratives · ACL (1) 2019 |
Natural language and speech › Language models and text generation › language modeling › language model architecture
pointer-generator network |
0.4 | 1 | 2019 | Simple and Effective Curriculum Pointer-Generator Networks for Reading Comprehension over Long Narratives · ACL (1) 2019 |
Natural language and speech › Information extraction and text analysis
entity linking |
0.3 | 1 | 2018 | Linking Fine-Grained Locations in User Comments (Extended Abstract) · ICDE 2018 |
Natural language and speech › Language models and text generation › natural language understanding
neural coherence model |
0.3 | 1 | 2018 | SkipFlow: Incorporating Neural Coherence Features for End-to-End Automatic Text Scoring · AAAI 2018 |
Data mining › structured data mining
graph mining |
0.3 | 1 | 2018 | Linking Fine-Grained Locations in User Comments (Extended Abstract) · ICDE 2018 |
Data mining › text mining
information extraction and text analysis |
0.3 | 1 | 2018 | CoNEREL: Collective Information Extraction in News Articles · SIGIR 2018 |
Data mining › text mining › information extraction
named entity recognition |
0.3 | 1 | 2018 | CoNEREL: Collective Information Extraction in News Articles · SIGIR 2018 |
Data mining
probabilistic graphical models |
0.3 | 1 | 2018 | Linking Fine-Grained Locations in User Comments (Extended Abstract) · ICDE 2018 |
Natural language and speech › Question answering and dialogue systems › community question answering
answer ranking |
0.3 | 1 | 2017 | Learning to Rank Question Answer Pairs with Holographic Dual LSTM Architecture · SIGIR 2017 |
Machine learning › Deep learning architectures and training › recurrent neural network
LSTM |
0.3 | 1 | 2017 | Learning to Rank Question Answer Pairs with Holographic Dual LSTM Architecture · SIGIR 2017 |
Machine learning › Deep learning architectures and training
recurrent neural network |
0.3 | 1 | 2017 | Learning to Rank Question Answer Pairs with Holographic Dual LSTM Architecture · SIGIR 2017 |
Bioinformatics and computational biology
biomedical text mining |
0.1 | 1 | 2019 | Robust Representation Learning of Biomedical Names · ACL (1) 2019 |
Graph algorithms and graph theory › spanning tree
minimum spanning tree |
0.1 | 1 | 2019 | Pair-Linking for Collective Entity Disambiguation: Two Could Be Better Than All · IEEE Trans. Knowl. Data Eng. 2019 |
Computer vision › 3D vision › correspondence estimation
semantic correspondence |
0.1 | 1 | 2017 | Learning to Rank Question Answer Pairs with Holographic Dual LSTM Architecture · SIGIR 2017 |
Web and social media mining › web usage mining
browse log mining |
0.1 | 1 | 2017 | Cross-Device User Linking: URL, Session, Visiting Time, and Device-log Embedding · SIGIR 2017 |
Recommender systems
user modeling |
0.1 | 1 | 2017 | Cross-Device User Linking: URL, Session, Visiting Time, and Device-log Embedding · SIGIR 2017 |
Methods — techniques the papers use, named apart from their topics
graph representation · 1.0synonym-aware representation learning · 0.8minimum spanning tree objective · 0.8iterative pair selection · 0.8contextual and conceptual encoding · 0.8probabilistic model · 0.7deep learning · 0.7LSTM · 0.7pointer-generator network · 0.4curriculum learning · 0.4probabilistic graphical model · 0.3pair-linking · 0.3interactive visualization · 0.3co-reference resolution · 0.3neural tensor layer · 0.3holographic composition · 0.3gradient boosting · 0.3embedding · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2020 | Collective Named Entity Recognition in User Comments via Parameterized Label PropagationabstractNamed entity recognition (NER) in the past has focused on extracting mentions in a local region, within a sentence or short paragraph. When dealing with user‐generated text, the diverse and informal writing style makes traditional approaches much less effective. On the other hand, in many types of text on social media such as user comments, tweets, or question–answer posts, the contextual connections between documents do exist. Examples include posts in a thread discussing the same topic, tweets that share a hashtag about the same entity. Our idea in this work is utilizing the related contexts across documents to perform mention recognition in a collective manner. Intuitively, within a mention coreference graph, the labels of mentions are expected to propagate from more confidence cases to less confidence ones. To this end, we propose a novel semisupervised inference algorithm named parameterized label propagation. In our model, the propagation weights between mentions are learned by an attention‐like mechanism, given their local contexts and the initial labels as input. We study the performance of our approach in the Yahoo! News data set, where comments and articles within a thread share similar context. The results show that our model significantly outperforms all other noncollective NER baselines. Minh C. Phan, Aixin Sun |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2019 | Robust Representation Learning of Biomedical NamesabstractBiomedical concepts are often mentioned in medical documents under different name variations (synonyms).This mismatch between surface forms is problematic, resulting in difficulties pertaining to learning effective representations.Consequently, this has tremendous implications such as rendering downstream applications inefficacious and/or potentially unreliable.This paper proposes a new framework for learning robust representations of biomedical names and terms.The idea behind our approach is to consider and encode contextual meaning, conceptual meaning, and the similarity between synonyms during the representation learning process.Via extensive experiments, we show that our proposed method outperforms other baselines on a battery of retrieval, similarity and relatedness benchmarks.Moreover, our proposed method is also able to compute meaningful representations for unseen names, resulting in high practical utility in real-world applications. Minh C. Phan, Aixin Sun, Yi Tay |
ACL (1) | 1 |
| 2019 | Simple and Effective Curriculum Pointer-Generator Networks for Reading Comprehension over Long NarrativesabstractYi Tay, Shuohang Wang, Anh Tuan Luu, Jie Fu, Minh C. Phan, Xingdi Yuan, Jinfeng Rao, Siu Cheung Hui, Aston Zhang. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. 2019. Yi Tay, Shuohang Wang, Anh Tuan Luu, Jie Fu 0001, Minh C. Phan, Xingdi Yuan, Jinfeng Rao, Siu Cheung Hui, Aston Zhang |
ACL (1) | 5 |
| 2019 | Pair-Linking for Collective Entity Disambiguation: Two Could Be Better Than AllabstractCollective entity disambiguation, or collective entity linking aims to jointly resolve multiple mentions by linking them to their associated entities in a knowledge base. Previous works are primarily based on the underlying assumption that entities within the same document are highly related. However, the extent to which these entities are actually connected in reality is rarely studied and therefore raises interesting research questions. For the first time, this paper shows that the semantic relationships between mentioned entities within a document are in fact less dense than expected. This could be attributed to several reasons such as noise, data sparsity, and knowledge base incompleteness. As a remedy, we introduce MINTREE, a new tree-based objective for the problem of entity disambiguation. The key intuition behind MINTREE is the concept of coherence relaxation which utilizes the weight of a minimum spanning tree to measure the coherence between entities. Based on this new objective, we design Pair-Linking, a novel iterative solution for the MINTREE optimization problem. The idea of Pair-Linking is simple: instead of considering all the given mentions, Pair-Linking iteratively selects a pair with the highest confidence at each step for decision making. Via extensive experiments on eight benchmark datasets, we show that our approach is not only more accurate but also surprisingly faster than many state-of-the-art collective linking algorithms. Minh C. Phan, Aixin Sun, Yi Tay, Jialong Han, Chenliang Li 0005 |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2018 | SkipFlow: Incorporating Neural Coherence Features for End-to-End Automatic Text ScoringabstractDeep learning has demonstrated tremendous potential for Automatic Text Scoring (ATS) tasks. In this paper, we describe a new neural architecture that enhances vanilla neural network models with auxiliary neural coherence features. Our new method proposes a new SkipFlow mechanism that models relationships between snapshots of the hidden representations of a long short-term memory (LSTM) network as it reads. Subsequently, the semantic relationships between multiple snapshots are used as auxiliary features for prediction. This has two main benefits. Firstly, essays are typically long sequences and therefore the memorization capability of the LSTM network may be insufficient. Implicit access to multiple snapshots can alleviate this problem by acting as a protection against vanishing gradients. The parameters of the SkipFlow mechanism also acts as an auxiliary memory. Secondly, modeling relationships between multiple positions allows our model to learn features that represent and approximate textual coherence. In our model, we call this neural coherence features. Overall, we present a unified deep learning architecture that generates neural coherence features as it reads in an end-to-end fashion. Our approach demonstrates state-of-the-art performance on the benchmark ASAP dataset, outperforming not only feature engineering baselines but also other deep learning models. Yi Tay, Minh C. Phan, Anh Tuan Luu, Siu Cheung Hui |
AAAI | 2 |
| 2018 | Linking Fine-Grained Locations in User Comments (Extended Abstract)abstractMany domain-specific websites host a profile page for each entity (e.g., locations on Foursquare, movies on IMDb, and products on Amazon), and users can post comments on it. When commenting on an entity, users often mention other entities for reference or comparison. Compared with web pages and tweets, disambiguating the mentioned entities in user comments has not received much attention. This paper investigates linking fine-grained locations in Foursquare comments. We demonstrate that the focal location, i.e., the location that a comment is posted on, provides rich contexts for linking. To exploit such information, we represent the Foursquare data in a graph, which includes locations, comments, and their relations. A probabilistic model named FocalLink is proposed to estimate the probability that a user mentions a location when commenting on a focal location, by following different kinds of relations. Experimental results show that FocalLink is consistently superior to different baselines. Jialong Han, Aixin Sun, Gao Cong, Wayne Xin Zhao, Zongcheng Ji, Minh C. Phan |
ICDE | 6 |
| 2018 | CoNEREL: Collective Information Extraction in News ArticlesabstractWe present CoNEREL, a system for collective named entity recognition and entity linking focusing on news articles and readers' comments. Different from other systems, CoNEREL processes articles and comments in batch mode, to make the best use of the shared contexts of multiple news stories and their comments. Particularly, a news article provides context for all its comments. To improve named entity recognition, CoNEREL utilizes co-reference of mentions to refine their class labels ( e.g. , person, location). To link the recognized entities to Wikipedia, our system implements Pair-Linking, a state-of-the-art entity linking algorithm. Furthermore, CoNEREL provides an interactive visualization of the Pair-Linking process. From the visualization, one can understand how Pair-Linking achieves decent linking performance through iterative evidence building, while being extremely fast and efficient. The graph formed by the Pair-Linking process naturally becomes a good summary of entity relations, making CoNEREL a useful tool to study the relationships between the entities mentioned in an article, as well as the ones that are discussed in its comments. Minh C. Phan, Aixin Sun |
SIGIR | 1 |
| 2018 | Linking Fine-Grained Locations in User CommentsabstractMany domain-specific websites host a profile page for each entity (e.g., locations on Foursquare, movies on IMDb, and products on Amazon) for users to post comments on. When commenting on an entity, users often mention other entities for reference or comparison. Compared with web pages and tweets, the problem of disambiguating the mentioned entities in user comments has not received much attention. This paper investigates linking fine-grained locations in Foursquare comments. We demonstrate that the focal location, i.e., the location that a comment is posted on, provides rich contexts for the linking task. To exploit such information, we represent the Foursquare data in a graph, which includes locations, comments, and their relations. A probabilistic model named FocalLink is proposed to estimate the probability that a user mentions a location when commenting on a focal location, by following different kinds of relations. Experimental results show that FocalLink is consistently superior under different collective linking settings. Jialong Han, Aixin Sun, Gao Cong, Wayne Xin Zhao, Zongcheng Ji, Minh C. Phan |
IEEE Trans. Knowl. Data Eng. | 6 |
| 2017 | NeuPL: Attention-based Semantic Matching and Pair-Linking for Entity DisambiguationabstractEntity disambiguation, also known as entity linking, is the task of mapping mentions in text to the corresponding entities in a given knowledge base, e.g. Wikipedia. Two key challenges are making use of mention's context to disambiguate (i.e. local objective), and promoting coherence of all the linked entities (i.e. global objective). In this paper, we propose a deep neural network model to effectively measure the semantic matching between mention's context and target entity. We are the first to employ the long short-term memory (LSTM) and attention mechanism for entity disambiguation. We also propose Pair-Linking, a simple but effective and significantly fast linking algorithm. Pair-Linking iteratively identifies and resolves pairs of mentions, starting from the most confident pair. It finishes linking all mentions in a document by scanning the pairs of mentions at most once. Our neural network model combined with Pair-Linking, named NeuPL, outperforms state-of-the-art systems over different types of documents including news, RSS, and tweets. Minh C. Phan, Aixin Sun, Yi Tay, Jialong Han, Chenliang Li 0005 |
CIKM | 1 |
| 2017 | Multi-Task Neural Network for Non-discrete Attribute Prediction in Knowledge GraphsabstractMany popular knowledge graphs such as Freebase, YAGO or DBPedia maintain a list of non-discrete attributes for each entity. Intuitively, these attributes such as height, price or population count are able to richly characterize entities in knowledge graphs. This additional source of information may help to alleviate the inherent sparsity and incompleteness problem that are prevalent in knowledge graphs. Unfortunately, many state-of-the-art relational learning models ignore this information due to the challenging nature of dealing with non-discrete data types in the inherently binary-natured knowledge graphs. In this paper, we propose a novel multi-task neural network approach for both encoding and prediction of non-discrete attribute information in a relational setting. Specifically, we train a neural network for triplet prediction along with a separate network for attribute value regression. Via multi-task learning, we are able to learn representations of entities, relations and attributes that encode information about both tasks. Moreover, such attributes are not only central to many predictive tasks as an information source but also as a prediction target. Therefore, models that are able to encode, incorporate and predict such information in a relational learning context are highly attractive as well. We show that our approach outperforms many state-of-the-art methods for the tasks of relational triplet classification and attribute value prediction. Yi Tay, Anh Tuan Luu, Minh C. Phan, Siu Cheung Hui |
CIKM | 3 |
| 2017 | Cross-Device User Linking: URL, Session, Visiting Time, and Device-log EmbeddingabstractCross-Device User Linking is the task of detecting same users given their browsing logs on different devices (e.g., tablet, mobile phone, PC, etc.). The problem was introduced in CIKM Cup 2016 together with a new dataset provided by Data-Centric Alliance (DCA). In this paper, we present insightful analysis on the dataset and propose a solution to link users based on their visited URLs, visiting time, and profile embeddings. We cast the problem as pairwise classification and use gradient boosting as the leaning-to-rank model. Our model works on a set of features exacted from URLs, titles, time and session data derived from user device-logs. The model outperforms the best solution in the CIKM Cup by a large margin. Minh C. Phan, Aixin Sun, Yi Tay |
SIGIR | 1 |
| 2017 | Learning to Rank Question Answer Pairs with Holographic Dual LSTM ArchitectureabstractWe describe a new deep learning architecture for learning to rank question answer pairs. Our approach extends the long short-term memory (LSTM) network with holographic composition to model the relationship between question and answer representations. As opposed to the neural tensor layer that has been adopted recently, the holographic composition provides the benefits of scalable and rich representational learning approach without incurring huge parameter costs. Overall, we present Holographic Dual LSTM (HD-LSTM), a unified architecture for both deep sentence modeling and semantic matching. Essentially, our model is trained end-to-end whereby the parameters of the LSTM are optimized in a way that best explains the correlation between question and answer representations. In addition, our proposed deep learning architecture requires no extensive feature engineering. Via extensive experiments, we show that HD-LSTM outperforms many other neural architectures on two popular benchmark QA datasets. Empirical studies confirm the effectiveness of holographic composition over the neural tensor layer. Yi Tay, Minh C. Phan, Anh Tuan Luu, Siu Cheung Hui |
SIGIR | 2 |