EDBT 2026 Demo / reviewers in the wild / expert
Xuan-Hieu Phan
dblp:119/2170 · also Xuan Hieu Phan
· DBLP profile ↗
28ranked-venue papers
6as first author
9since 2021 · last 2025
0000-0002-7640-9190ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 23 · 3 first-author · 9 since 2021Databases, data management, data science and information retrieval · 15 · 5 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Uncovering connections: a reference network approach to statute law retrieval
Thi-Hai-Yen Vuong, Hai-Long Nguyen 0001, Tan-Minh Nguyen, Ha-Thanh Nguyen, Minh Le Nguyen 0001, Xuan-Hieu Phan |
Appl. Intell. | 6 |
| 2025 | Mi-CGA: Cross-modal Graph Attention Network for robust emotion recognition in the presence of incomplete modalities
Cam-Van Thi Nguyen, Hai-Dang Kieu, Quang-Thuy Ha, Xuan-Hieu Phan, Duc-Trong Le |
Neurocomputing | 4 |
| 2024 | Aspect-Based Sentiment Analysis of Clothing Reviews in Vietnamese E-commerce
Pham Quoc-Hung, Dinh Van-Dan, Huu-Loi Le, Le Thi-Viet-Huong, Nguyen Thu Ha, Xuan-Hieu Phan, Minh-Tien Nguyen, Pham Ngoc Hung |
PACLIC | 6 |
| 2024 | Learning to generate text with auxiliary tasks
Pham Quoc-Hung, Minh-Tien Nguyen, Shumpei Inoue, Manh Tran-Tien, Xuan-Hieu Phan |
Knowl. Based Syst. | 5 |
| 2024 | Towards Vietnamese Question and Answer Generation: An Empirical StudyabstractQuestion-answer generation (QAG) is a challenging task that generates both questions and answers from a given input paragraph context. The QAG task has recently achieved promising results thanks to the appearance of large pre-trained language models, yet, QAG models are mainly implemented in common languages, e.g., English. There still remains a gap in domain and language adaptation of these QAG models to low-resource languages such as Vietnamese. To address the gap, this article presents a large-scale and systematic study of QAG in Vietnamese. To do that, we first implement several QAG models by using the common fine-tuning techniques based on powerful pre-trained language models. We next introduce a set of instructions designed for the QAG task. These instructions are used to fine-tuned the pre-trained language and large language models. Extensive experimental results of both automatic and human evaluation on five benchmark machine reading comprehension datasets show two important points. First, the instruction-tuning method has the potential to enhance the performance of QAG models. Second, large language models trained in English need more data for fine-tuning to work well on the downstream QAG tasks of low-resource languages. We also provide a prototype system to demonstrate how our QAG models actually work. The code for fine-tuning QAG models and instructions are also made available. Pham Quoc-Hung, Huu-Loi Le, Dang Nhat Minh, T. Tran Khang, Manh Tran-Tien, Viet-Hung Dang, Huy-The Vu, Minh-Tien Nguyen, Xuan-Hieu Phan |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 9 |
| 2023 | GIFT4Rec: An Effective Side Information Fusion Technique Apply to Graph Neural Network for Cold-Start Recommendation
Tran-Ngoc-Linh Nguyen, Chi-Dung Vu, Hoang-Ngan Le, Anh-Dung Hoang, Xuan-Hieu Phan, Quang-Thuy Ha, Hoang-Quynh Le, Mai-Vu Tran |
ACIIDS (1) | 5 |
| 2023 | An Improvement of Diachronic Embedding for Temporal Knowledge Graph Completion
Thuy-Anh Nguyen Thi, Viet-Phuong Ta, Xuan-Hieu Phan, Quang-Thuy Ha |
ACIIDS (2) | 3 |
| 2023 | Passage-based BM25 Hard Negatives: A Simple and Effective Negative Sampling Strategy For Dense Retrieval
Thanh-Do Nguyen, Chi Minh Bui, Thi-Hai-Yen Vuong, Xuan-Hieu Phan |
PACLIC | 4 |
| 2022 | Parameter Distribution Ensemble Learning for Sudden Concept Drift Detection
Khanh-Tung Nguyen, Trung Tran, Xuan-Hieu Phan, Quang-Thuy Ha |
ACIIDS (2) | 4 |
| 2018 | Exploiting User Posts for Web Document SummarizationabstractRelevant user posts such as comments or tweets of a Web document provide additional valuable information to enrich the content of this document. When creating user posts, readers tend to borrow salient words or phrases in sentences. This can be considered as word variation. This article proposes a framework that models the word variation aspect to enhance the quality of Web document summarization. Technically, the framework consists of two steps: scoring and selection. In the first step, the social information of a Web document such as user posts is exploited to model intra-relations and inter-relations in lexical and semantic levels. These relations are denoted by a mutual reinforcement similarity graph used to score each sentence and user post. After scoring, summaries are extracted by using a ranking approach or concept-based method formulated in the form of Integer Linear Programming. To confirm the efficiency of our framework, sentence and story highlight extraction tasks were taken as a case study on three datasets in two languages, English and Vietnamese. Experimental results show that: (i) the framework can improve ROUGE-scores compared to state-of-the-art baselines of social context summarization and (ii) the combination of the two relations benefits the sentence extraction of single Web documents. Minh-Tien Nguyen, Vu D. Tran, Minh Le Nguyen 0001, Xuan-Hieu Phan |
ACM Trans. Knowl. Discov. Data | 4 |
| 2016 | Learning to Filter User Explicit Intents in Online Vietnamese Social Media Texts
Thai-Le Luong, Thi-Hanh Tran, Quoc-Tuan Truong, Thi-Minh-Ngoc Truong, Thi-Thu Phi, Xuan-Hieu Phan |
ACIIDS (2) | 6 |
| 2016 | Identifying User Intents in Vietnamese Spoken Language Commands and Its Application in Smart Mobile Voice Interaction
Thi-Lan Ngo, Van-Hop Nguyen, Thi-Hai-Yen Vuong, Thac-Thong Nguyen, Thi-Thua Nguyen, Son Bao Pham, Xuan-Hieu Phan |
ACIIDS (1) | 7 |
| 2016 | Named Entity Recognition for Vietnamese Spoken Texts and Its Application in Smart Mobile Voice Interaction
Phuong-Nam Tran 0002, Van-Duc Ta, Quoc-Tuan Truong, Quang-Vu Duong, Thac-Thong Nguyen, Xuan-Hieu Phan |
ACIIDS (1) | 6 |
| 2013 | A feature-word-topic model for image annotation and retrievalabstractImage annotation is a process of finding appropriate semantic labels for images in order to obtain a more convenient way for indexing and searching images on the Web. This article proposes a novel method for image annotation based on combining feature-word distributions, which map from visual space to word space, and word-topic distributions, which form a structure to capture label relationships for annotation. We refer to this type of model as Feature-Word-Topic models. The introduction of topics allows us to efficiently take word associations, such as {ocean, fish, coral} or {desert, sand, cactus}, into account for image annotation. Unlike previous topic-based methods, we do not consider topics as joint distributions of words and visual features, but as distributions of words only. Feature-word distributions are utilized to define weights in computation of topic distributions for annotation. By doing so, topic models in text mining can be applied directly in our method. Our Feature-word-topic model, which exploits Gaussian Mixtures for feature-word distributions, and probabilistic Latent Semantic Analysis (pLSA) for word-topic distributions, shows that our method is able to obtain promising results in image annotation and retrieval. Cam-Tu Nguyen, Natsuda Kaothanthong, Takeshi Tokuyama, Xuan-Hieu Phan |
ACM Trans. Web | 4 |
| 2011 | A Hidden Topic-Based Framework toward Building Applications with Short Web DocumentsabstractThis paper introduces a hidden topic-based framework for processing short and sparse documents (e.g., search result snippets, product descriptions, book/movie summaries, and advertising messages) on the Web. The framework focuses on solving two main challenges posed by these kinds of documents: 1) data sparseness and 2) synonyms/homonyms. The former leads to the lack of shared words and contexts among documents while the latter are big linguistic obstacles in natural language processing (NLP) and information retrieval (IR). The underlying idea of the framework is that common hidden topics discovered from large external data sets (universal data sets), when included, can make short documents less sparse and more topic-oriented. Furthermore, hidden topics from universal data sets help handle unseen data better. The proposed framework can also be applied for different natural languages and data domains. We carefully evaluated the framework by carrying out two experiments for two important online applications (Web search result classification and matching/ranking for contextual advertising) with large-scale universal data sets and we achieved significant results. Xuan-Hieu Phan, Cam-Tu Nguyen, Dieu-Thu Le, Minh Le Nguyen 0001, Susumu Horiguchi, Quang-Thuy Ha |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2010 | A feature-word-topic model for image annotationabstractImage annotation is to automatically associate semantic labels with images in order to obtain a more convenient way for indexing and searching images on the Web. This paper proposes a novel method for image annotation based on feature-word and word-topic distributions. The introduction of topics enables us to efficiently take word associations, such as {ocean, fish, coral}, into image annotation. Feature-word distributions are utilized to define weights in computation of topic distributions for annotation. By doing so, topic models in text mining can be applied directly in our method. Experiments show that our method is able to obtain promising improvements over the state-of-the-art method - Supervised Multiclass Labeling (SML) Cam-Tu Nguyen, Natsuda Kaothanthong, Xuan-Hieu Phan, Takeshi Tokuyama |
CIKM | 3 |
| 2009 | Web Search Clustering and Labeling with Hidden TopicsabstractWeb search clustering is a solution to reorganize search results (also called “snippets”) in a more convenient way for browsing. There are three key requirements for such post-retrieval clustering systems: (1) the clustering algorithm should group similar documents together; (2) clusters should be labeled with descriptive phrases; and (3) the clustering system should provide high-quality clustering without downloading the whole Web page. This article introduces a novel framework for clustering Web search results in Vietnamese which targets the three above issues. The main motivation is that by enriching short snippets with hidden topics from huge resources of documents on the Internet, it is able to cluster and label such snippets effectively in a topic-oriented manner without concerning whole Web pages. Our approach is based on recent successful topic analysis models, such as Probabilistic-Latent Semantic Analysis, or Latent Dirichlet Allocation. The underlying idea of the framework is that we collect a very large external data collection called “universal dataset,” and then build a clustering system on both the original snippets and a rich set of hidden topics discovered from the universal data collection. This can be seen as a richer representation of snippets to be clustered. We carry out careful evaluation of our method and show that our method can yield impressive clustering quality. Cam-Tu Nguyen, Xuan-Hieu Phan, Susumu Horiguchi, Thu-Trang Nguyen, Quang-Thuy Ha |
ACM Trans. Asian Lang. Inf. Process. | 2 |
| 2008 | Online Structured Learning for Semantic Parsing with Synchronous and lambda-Synchronous Context Free GrammarsabstractWe formulate semantic parsing as a parsing problem on a synchronous context free grammar (SCFG) which is automatically built on the corpus of natural language sentences and the representation of semantic outputs. We then present an online learning framework for estimating the synchronous SCFG grammar. In addition, our online learning methods for semantic parsing problems are also extended to deal with the case, in which the semantic representation could be represented under lambda-calculus. Experimental results in the domain of semantic parsing show advantages in comparison with previous works. Minh Le Nguyen 0001, Akira Shimazu, Xuan-Hieu Phan, Phuong-Thai Nguyen |
ICTAI (2) | 3 |
| 2008 | Matching and Ranking with Hidden Topics towards Online Contextual AdvertisingabstractIn online contextual advertising, ad messages are displayed related to the content of the target Web page. It leads to the problem in information retrieval community: how to select the most relevant ad messages given the content of a page. To deal with this problem, we propose a framework that takes advantage of large scale external datasets. This framework provides a mechanism to discover the semantic relations between Web pages and ad messages by analyzing topics for them. This helps overcome the problem of mismatch due to unimportant words and the difference in vocabularies between Web pages and ad messages. The framework has been evaluated through a number of experiments. It shows a significant improvement in accuracy over word/lexicon-based matching and ranking methods. Dieu-Thu Le, Cam-Tu Nguyen, Quang-Thuy Ha, Xuan-Hieu Phan, Susumu Horiguchi |
Web Intelligence | 4 |
| 2008 | Learning to classify short and sparse text & web with hidden topics from large-scale data collectionsabstractThis paper presents a general framework for building classifiers that deal with short and sparse text & segments by making the most of hidden topics discovered from large-scale data collections. The main motivation of this work is that many classification tasks working with short segments of text & Web, such as search snippets, forum & chat messages, blog & news feeds, product reviews, and book & movie summaries, fail to achieve high accuracy due to the data sparseness. We, therefore, come up with an idea of gaining external knowledge to make the data more related as well as expand the coverage of classifiers to handle future data better. The underlying idea of the framework is that for each classification task, we collect a large-scale external data collection called universal dataset, and then build a classifier on both a (small) set of labeled training data and a rich set of hidden topics discovered from that data collection. The framework is general enough to be applied to different data domains and genres ranging from search results to medical text. We did a careful evaluation on several hundred megabytes of Wikipedia (30M words) and MEDLINE (18M words) with two tasks: Web search domain disambiguation and disease categorization for medical text, and achieved significant quality enhancement. Xuan-Hieu Phan, Minh Le Nguyen 0001, Susumu Horiguchi |
WWW | 1 |
| 2007 | A Multilingual Dependency Analysis System Using Online Passive-Aggressive Learning
Minh Le Nguyen 0001, Akira Shimazu, Phuong-Thai Nguyen, Xuan-Hieu Phan |
EMNLP-CoNLL | 4 |
| 2006 | Semantic Parsing with Structured SVM Ensemble Classification Models
Minh Le Nguyen 0001, Akira Shimazu, Xuan-Hieu Phan |
ACL | 3 |
| 2006 | Vietnamese Word Segmentation with CRFs and SVMs: An Investigation
Cam-Tu Nguyen, Trung-Kien Nguyen, Xuan-Hieu Phan, Minh Le Nguyen 0001, Quang-Thuy Ha |
PACLIC | 3 |
| 2006 | Improving discriminative sequential learning by discovering important association of statisticsabstractDiscriminative sequential learning models like Conditional Random Fields (CRFs) have achieved significant success in several areas such as natural language processing or information extraction. Their key advantage is the ability to capture various nonindependent and overlapping features of inputs. However, several unexpected pitfalls have a negative influence on the model's performance; these mainly come from a high imbalance among classes, irregular phenomena, and potential ambiguity in the training data. This article presents a data-driven approach that can deal with such difficult data instances by discovering and emphasizing important conjunctions or associations of statistics hidden in the training data. Discovered associations are then incorporated into these models to deal with difficult data instances. Experimental results of phrase-chunking and named entity recognition using CRFs show a significant improvement in accuracy. In addition to the technical perspective, our approach also highlights a potential connection between association mining and statistical learning by offering an alternative strategy to enhance learning performance with interesting and useful patterns discovered from large datasets. Xuan-Hieu Phan, Minh Le Nguyen 0001, Yasushi Inoguchi, Susumu Horiguchi |
ACM Trans. Asian Lang. Inf. Process. | 1 |
| 2005 | Classification with Maximum Entropy Modeling of Predictive Association Rules
Xuan-Hieu Phan, Minh Le Nguyen 0001, Susumu Horiguchi, Yasushi Inoguchi |
ECML | 1 |
| 2005 | Improving discriminative sequential learning with rare--but--important associationsabstractDiscriminative sequential learning models like Conditional Random Fields (CRFs) have achieved significant success in several areas such as natural language processing or information extraction. Their key advantage is the ability to capture various non--independent and overlapping features of inputs. However, several unexpected pitfalls have a negative influence on the model's performance; these mainly come from an imbalance among classes/labels, irregular phenomena, and potential ambiguity in the training data. This paper presents a data--driven approach that can deal with such hard--to--predict data instances by discovering and emphasizing rare--but--important associations of statistics hidden in the training data. Mined associations are then incorporated into these models to deal with difficult examples. Experimental results of English phrase chunking and named entity recognition using CRFs show a significant improvement in accuracy. In addition to the technical perspective, our approach also highlights a potential connection between association mining and statistical learning by offering an alternative strategy to enhance learning performance with interesting and useful patterns discovered from large dataset. Xuan-Hieu Phan, Minh Le Nguyen 0001, Susumu Horiguchi |
KDD | 1 |
| 2005 | A Structured SVM Semantic Parser Augmented by Semantic Tagging with Conditional Random Field
Minh Le Nguyen 0001, Akira Shimazu, Xuan-Hieu Phan |
PACLIC | 3 |
| 2004 | PEWeb: Product Extraction from the Web Based on Entropy EstimationabstractMining product descriptions (PDs) from e-commercial web sites is an important task in information extraction from the Web. In this paper, we propose an efficient technique for this task. The technique first discovers the set of PDs based on the measure of entropy at each internal node in the HTML tag tree. Afterwards, a set of association rules based on heuristic features is employed to filter the output and therefore enhance the precision. The experimental results of PEWeb system show that the proposed method outperforms existing automatic techniques remarkably. Xuan-Hieu Phan, Susumu Horiguchi |
Web Intelligence | 1 |