Thuy Vu

dblp:22/5260 · DBLP profile ↗
← Back
21ranked-venue papers
9as first author
12since 2021 · last 2026
0000-0003-1056-6975ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 18 · 8 first-author · 11 since 2021Databases, data management, data science and information retrieval · 7 · 3 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 CypherSmith: Transforming Text-to-Cypher Generation for LLMs with Synthetic Data
abstract
Knowledge Graph (KG) retrieval is a promising augmentation to address knowledge gaps and hallucinations in LLMs. As KGs in practice are stored in graph databases (e.g., Wikidata, Freebase), accurate retrieval requires translating natural language questions into structured queries (query generation). A key challenge of query generation is Text-to-Cypher, which generates Cypher queries for property graphs (e.g., Neo4j), a paradigm increasingly adopted in industry for their scalable architectures and expressive schemas. However, compared to other query generation tasks such as Text-to-SQL or Text-to-SPARQL, Text-to-Cypher remains underexplored due to scarce public KGs and datasets. Existing datasets are small, domain-limited, and lack diversity, constraining LLM progress. To address this, we introduce CypherSmith, an instruction-tuning dataset over 12\times larger than prior public Text-to-Cypher datasets, spanning diverse domains to better support LLM fine-tuning. Our key distinction lies in fully leveraging open-source LLMs for large-scale synthetic data generation and introducing a novel likelihood-based filtering technique to ensure high-quality Text-to-Cypher data. Extensive experiments demonstrate the effectiveness of CypherSmith, achieving state-of-the-art LLM performance.
Kexuan Sun 0002, Jens-S. Vöckler, Thien Huu Nguyen, Thuy Vu
ACL (1)6
2025 Retrieving Support to Rank Answers in Open-Domain Question Answering
abstract
We introduce a novel Question Answering (QA) architecture that enhances answer selection by retrieving targeted supporting evidence.Unlike traditional methods, which retrieve documents or passages relevant only to a query q, our approach retrieves content relevant to the combined pair (q, a), explicitly emphasizing the supporting relation between the query and a candidate answer a.By prioritizing this relational context, our model effectively identifies paragraphs that directly substantiate the correctness of a with respect to q, leading to more accurate answer verification than standard retrieval systems.Our neural retrieval method also scales efficiently to collections containing hundreds of millions of paragraphs.Moreover, this approach can be used by large language models (LLMs) to retrieve explanatory paragraphs that ground their reasoning, enabling them to tackle more complex QA tasks with greater reliability and interpretability.
Alessandro Moschitti, Thuy Vu
EMNLP3
2024 In Situ Answer Sentence Selection at Web-scale
abstract
Current answer sentence selection (AS2) applied in open-domain question answering (ODQA) selects answers by ranking a large set of candidates, i.e., sentences, extracted from the retrieved text. In this paper, we present Passage-based Extracting Answer Sentence In-place (PEASI), a novel answer selection model optimized for Web-scale setting. This is a Transformer-based network that can jointly (i) rerank passages retrieved for a question and (ii) identify a probable answer from the top passages. We train PEASI with multi-task learning for sharing representations between the passage reranker and answer sentence extractor. We construct a new large-scale QA dataset (WQA) consisting of 800,000+ labeled passages/sentences for 60,000+ questions. The experiment results show that PEASI outperforms AS2 state of the art by 6.51% in accuracy on WQA, from 48.86% to 55.37%.
Zeyu Zhang 0002, Thuy Vu, Alessandro Moschitti
CIKM2
2024 Reinforcement Learning from Answer Reranking Feedback for Retrieval-Augmented Answer Generation
Minh Nguyen 0007, Toàn Quoc Nguyên, Kishan KC, Zeyu Zhang 0002, Thuy Vu
INTERSPEECH5
2023 Question-Answer Sentence Graph for Joint Modeling Answer Selection
abstract
This research studies graph-based approaches for Answer Sentence Selection (AS2), an essential component for retrieval-based Question Answering (QA) systems.During offline learning, our model constructs a small-scale relevant training graph per question in an unsupervised manner, and integrates with Graph Neural Networks.Graph nodes are question sentence to answer sentence pairs.We train and integrate state-of-the-art (SOTA) models for computing scores between question-question, question-answer, and answer-answer pairs, and use thresholding on relevance scores for creating graph edges.Online inference is then performed to solve the AS2 task on unseen queries.Experiments on two well-known academic benchmarks and a real-world dataset show that our approach consistently outperforms SOTA QA baseline models.
Roshni G. Iyer, Thuy Vu, Alessandro Moschitti, Yizhou Sun
EACL2
2023 Question-Context Alignment and Answer-Context Dependencies for Effective Answer Sentence Selection
Minh Nguyen 0007, Kishan KC, Toàn Quoc Nguyên, Thien Huu Nguyen, Ankit Chadha, Thuy Vu
INTERSPEECH6
2023 Efficient Fine-Tuning Large Language Models for Knowledge-Aware Response Planning
Minh Nguyen 0007, Kishan KC, Toàn Quoc Nguyên, Ankit Chadha, Thuy Vu
ECML/PKDD (2)5
2022 WDRASS: A Web-scale Dataset for Document Retrieval and Answer Sentence Selection
abstract
Open-Domain Question Answering (ODQA) systems generate answers from relevant text returned by search engines, e.g., lexical features-based such as BM25, or embeddings-based such as dense passage retrieval (DPR). Few datasets are available for this task: they mainly focus on QA systems based on machine reading (MR) approach, and show problematic evaluation, mostly based on uncontextualized short answer matching. In this paper, we present WDRASS, a dataset for ODQA based on answer sentence selection (AS2) models, which consider sentences as candidate answers for QA systems. WDRASS consists of ∼64k questions and 800k+ labeled passages and sentences extracted from 30M documents. We evaluate the dataset by training models on it and comparing with the same models trained on Google NQ. Our experiments show that WDRASS significantly improves the performance of retrieval and reranking models, thus boosting the accuracy of downstream QA tasks. We believe our dataset can produce significant impact in advancing IR research.
Zeyu Zhang 0002, Thuy Vu, Sunil Gandhi, Ankit Chadha, Alessandro Moschitti
CIKM2
2021 Joint Models for Answer Verification in Question Answering Systems
abstract
Zeyu Zhang, Thuy Vu, Alessandro Moschitti. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Zeyu Zhang 0002, Thuy Vu, Alessandro Moschitti
ACL/IJCNLP (1)2
2021 CDA: a Cost Efficient Content-based Multilingual Web Document Aligner
abstract
We introduce a Content-based Document Alignment approach (CDA), an efficient method to align multilingual web documents based on content in creating parallel training data for machine translation (MT) systems operating at the industrial level.CDA works in two steps: (i) projecting documents of a web domain to a shared multilingual space; then (ii) aligning them based on the similarity of their representations in such space.We leverage lexical translation models to build vector representations using TF×IDF.CDA achieves performance comparable with state-of-the-art systems in the WMT-16 Bilingual Document Alignment Shared Task benchmark while operating in multilingual space.Besides, we created two web-scale datasets to examine the robustness of CDA in an industrial setting involving up to 28 languages and millions of documents.The experiments show that CDA is robust, cost-effective, and is significantly superior in (i) processing large and noisy web data and (ii) scaling to new and low-resourced languages.
Thuy Vu, Alessandro Moschitti
EACL1
2021 Machine Translation Customization via Automatic Training Data Selection from the Web
Thuy Vu, Alessandro Moschitti
ECIR (1)1
2021 AVA: an Automatic eValuation Approach for Question Answering Systems
abstract
We introduce AVA, an automatic evaluation approach for Question Answering, which given a set of questions associated with Gold Standard answers, can estimate system Accuracy. AVA uses Transformer-based language models to encode question, answer, and reference text. This allows for effectively measuring the similarity between the reference and an automatic answer, biased towards the question semantics. To design, train and test AVA, we built multiple large training, development, and test sets on both public and industrial benchmarks. Our innovative solutions achieve up to 74.7% in F1 score in predicting human judgement for single answers. Additionally, AVA can be used to evaluate the overall system Accuracy with an RMSE, ranging from 0.02 to 0.09, depending on the availability of multiple references.
Thuy Vu, Alessandro Moschitti
NAACL-HLT1
2020 TANDA: Transfer and Adapt Pre-Trained Transformer Models for Answer Sentence Selection
abstract
We propose TandA, an effective technique for fine-tuning pre-trained Transformer models for natural language tasks. Specifically, we first transfer a pre-trained model into a model for a general task by fine-tuning it with a large and high-quality dataset. We then perform a second fine-tuning step to adapt the transferred model to the target domain. We demonstrate the benefits of our approach for answer sentence selection, which is a well-known inference task in Question Answering. We built a large scale dataset to enable the transfer step, exploiting the Natural Questions dataset. Our approach establishes the state of the art on two well-known benchmarks, WikiQA and TREC-QA, achieving the impressive MAP scores of 92% and 94.3%, respectively, which largely outperform the the highest scores of 83.4% and 87.5% of previous work. We empirically show that TandA generates more stable and robust models reducing the effort required for selecting optimal hyper-parameters. Additionally, we show that the transfer step of TandA makes the adaptation step more robust to noise. This enables a more effective use of noisy datasets for fine-tuning. Finally, we also confirm the positive impact of TandA in an industrial setting, using domain specific datasets subject to different types of noise.
Siddhant Garg, Thuy Vu, Alessandro Moschitti
AAAI2
2020 Reranking for Efficient Transformer-based Answer Selection
abstract
IR-based Question Answering (QA) systems typically use a sentence selector to extract the answer from retrieved documents. Recent studies have shown that powerful neural models based on the Transformer can provide an accurate solution to Answer Sentence Selection (AS2). Unfortunately, their computation cost prevents their use in real-world applications. In this paper, we show that standard and efficient neural rerankers can be used to reduce the amount of sentence candidates fed to Transformer models without hurting Accuracy, thus improving efficiency up to four times. This is an important finding as the internal representation of shallower neural models is dramatically different from the one used by a Transformer model, e.g., word vs. contextual embeddings.
Yoshitomo Matsubara, Thuy Vu, Alessandro Moschitti
SIGIR2
2017 Extracting Urban Microclimates from Electricity Bills
abstract
Sustainable energy policies are of growing importance in all urban centers.Climate — and climate change — will play increasingly important roles in these policies.Climate zones defined by the California Energy Commissionhave long been influential in energy management.For example, recently a two-zone division of Los Angeles(defined by historical temperature averages) was introduced for electricity rate restructuring.The importance of climate zones has been enormous,and climate change could make them still more important. AI can provide improvements on the ways climate zones are derived and managed.This paper reports on analysis of aggregate household electricity consumption (EC) data from local utilities in Los Angeles,seeking possible improvements in energy management. In this analysis we noticed that EC data permits identificationof interesting geographical zones — regions having EC patterns that are characteristically different from surrounding regions.We believe these zones could be useful in a variety of urban models.
Thuy Vu, Douglas Stott Parker Jr.
AAAI1
2016 $K$-Embeddings: Learning Conceptual Embeddings for Words using Context
abstract
We describe a technique for adding contextual distinctions to word embeddings by extending the usual embedding process -into two phases.The first phase resembles existing methods, but also constructs K classifications of concepts.The second phase uses these classifications in developing refined K embeddings for words, namely word K-embeddings.The technique is iterative, scalable, and can be combined with other methods (including Word2Vec) in achieving still more expressive representations.Experimental results show consistently large performance gains on a Semantic-Syntactic Word Relationship test set for different K settings.For example, an overall gain of 20% is recorded at K = 5.In addition, we demonstrate that an iterative process can further tune the embeddings and gain an extra 1% (K = 10 in 3 iterations) on the same benchmark.The examples also show that polysemous concepts are meaningfully embedded in our K different conceptual embeddings for words.
Thuy Vu, Douglas Stott Parker Jr.
HLT-NAACL1
2015 Node Embeddings in Social Network Analysis
abstract
We introduce a distributed representation of nodes, node embeddings, in social network analysis. We compute embeddings for nodes based on their attributes and links. These embeddings can support many social network applications --- including analyses of community homogeneity, distance, and detection of community connectors (inter-community outliers, people who connect communities) --- thanks to the convenient yet efficient computation provided by node embeddings for structural comparisons. Our experimental results include many interesting insights about the computer science literature network (DBLP). For example, in DBLP prior to 2013 the best way for research in Natural Language & Speech to gain impact toward "best-paper" recognition was to emphasize aspects related to Machine Learning & Pattern Recognition.
Thuy Vu, Douglas Stott Parker Jr.
ASONAM1
2013 Interest mining from user tweets
abstract
We build a system to extract user interests from Twitter messages. Specifically, we extract interest candidates using linguistic patterns and rank them using four different keyphrase ranking techniques: TFIDF, TextRank, LDA-TextRank, and Relevance-Interestingness-Rank (RI-Rank). We also explore the complementary relation between TFIDF and TextRank in ranking interest candidates. Top ranked interests are evaluated with user feedback gathered from an online survey. The results show that TFIDF and TextRank are both suitable for extracting user interests from tweets. Moreover, the combination of TFIDF and TextRank consistently yields the highest user positive feedback.
Thuy Vu, Victor Perez 0002
CIKM1
2009 Feature-Based Method for Document Alignment in Comparable News Corpora
Thuy Vu, AiTi Aw, Min Zhang 0005
EACL1
2008 Term Extraction Through Unithood and Termhood Unification
Thuy Vu, AiTi Aw, Min Zhang 0005
IJCNLP1
2008 Identifying gene-disease associations using centrality on a literature mined gene-interaction network
abstract
MOTIVATION: Understanding the role of genetics in diseases is one of the most important aims of the biological sciences. The completion of the Human Genome Project has led to a rapid increase in the number of publications in this area. However, the coverage of curated databases that provide information manually extracted from the literature is limited. Another challenge is that determining disease-related genes requires laborious experiments. Therefore, predicting good candidate genes before experimental analysis will save time and effort. We introduce an automatic approach based on text mining and network analysis to predict gene-disease associations. We collected an initial set of known disease-related genes and built an interaction network by automatic literature mining based on dependency parsing and support vector machines. Our hypothesis is that the central genes in this disease-specific network are likely to be related to the disease. We used the degree, eigenvector, betweenness and closeness centrality metrics to rank the genes in the network. RESULTS: The proposed approach can be used to extract known and to infer unknown gene-disease associations. We evaluated the approach for prostate cancer. Eigenvector and degree centrality achieved high accuracy. A total of 95% of the top 20 genes ranked by these methods are confirmed to be related to prostate cancer. On the other hand, betweenness and closeness centrality predicted more genes whose relation to the disease is currently unknown and are candidates for experimental study. AVAILABILITY: A web-based system for browsing the disease-specific gene-interaction networks is available at: http://gin.ncibi.org.
Arzucan Özgür, Thuy Vu, Günes Erkan, Dragomir R. Radev
ISMB2