Ngoc Phuoc An Vo

dblp:153/9558 · also Ngoc-Phuoc-An Vo · DBLP profile ↗
← Back
14ranked-venue papers
7as first author
4since 2021 · last 2024
0000-0001-5646-5411ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 6 first-author · 4 since 2021Databases, data management, data science and information retrieval · 6 · 3 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2024 A Hybrid Cognitive Contract Application for Identifying Accounting Risks in Contractual Language
Ngoc Phuoc An Vo, Martin Linhart, Fruzsina Strbik, Istvan Koska, Petros Zerfos, Vadim Sheinin, Jeff Dakin, Milton Laverde
NLDB (2)1
2022 Natural Language Interface for Process Mining Queries in Healthcare
abstract
Recently, the needs of data required for data analysis are becoming more diversified, and research on data extraction and analysis methods has been continuously made in order to effectively respond to various needs. Process mining is a solution that analyzes various system logs built by companies or healthcare institutions so that they can be used for process improvement. From the process model extracted from the system logs, it is possible not only to grasp the exact flow of the current business process, but also to acquire additional information such as repetitive execution of activities in the process where the bottleneck occurs in the business process flow. The manufacturing industry has made great efforts to improve the process management, and as many companies are paying attention to big data these days, various data-related technologies are emerging in the healthcare industry as well to properly provide patients with the care needed. Process mining tools allow users to pull data by programming in a process mining query language using the APIs provided with the process mining tool, or by manually creating reusable analytical documents using user friendly tool. However, these tasks require the users to be familiar with the query language APIs and understand the data model and its relationships with respect to creating analytical documents. This paper proposes a methodology that allows users to easily extract desired data through natural language interface, which relieves nonprofessional users of the burden of programming in a process mining query language. The process mining query engine with natural language interface presented in this paper consists of four major components. Among them, the natural language processing pipeline that not only extracts intermediate representation of entities used when constructing a process mining query language report from natural language queries, but also effectively extracts a query hint from the context of natural language query. The query hint is used to select a process-specific function from the library that fits the context of the user query while transforming a natural language query into a process mining query report. The method proposed in this study has the advantage of being able to roughly grasp the process state for the user just by entering a query in natural language. The proposed system provides users with four query process options. That is, the user 1) retrieves intermediate representation of entities and query hints from the NLP pipeline, 2) retrieves the process mining query language from the query language generator, 3) submits the query language to the process mining engine and execute the query, 4) retrieves description of intermediate representation of entity and query hints in natural language to confirm that the query is processed correctly. The contents proposed in this paper were constructed and executed, and the query reports in process mining query language programmatically generated by the proposed query engine were also executed in a process mining engine and the query results were verified.
Hangu Yeo, Elahe Khorasani, Vadim Sheinin, Irene Manotas, Ngoc Phuoc An Vo, Octavian Popescu, Petros Zerfos
IEEE Big Data5
2022 Addressing Limitations of Encoder-Decoder Based Approach to Text-to-SQL
abstract
Most attempts on Text-to-SQL task using encoder-decoder approach show a big problem of dramatic decline in performance for new databases. For the popular Spider dataset, despite models achieving 70% accuracy on its development or test sets, the same models show a huge decline below 20% accuracy for unseen databases. The root causes for this problem are complex and they cannot be easily fixed by adding more manually created training. In this paper we address the problem and propose a solution that is a hybrid system using automated training-data augmentation technique. Our system consists of a rule-based and a deep learning components that interact to understand crucial information in a given query and produce correct SQL as a result. It achieves double-digit percentage improvement for databases that are not part of the Spider corpus.
Octavian Popescu, Irene Manotas, Ngoc Phuoc An Vo, Hangu Yeo, Elahe Khorashani, Vadim Sheinin
COLING3
2021 Programmatic Database Language Generation for Big Data Applications
abstract
Database management systems offer an efficient way of managing huge amount of data such as financial and healthcare data and he data retrieval from databases requires knowledge of Structured Query Language (SQL). In this paper, an Automatic SQL Generation System is proposed to help users who are inexperienced in querying database with SQL. The proposed SQL generation system reads formatted data items in the query report from the user and converts the data items into SQL statements programmatically with the help of a data model that is pulled from a database. The SQL generation system can handle simple queries composed of a query block with a SELECT statement as well as complex queries composed of multiple query blocks containing multiple SELECT statements. The proposed system is integrated with an NLIDB (Natural Language Interface for Database) system to translate data items (or tokens) extracted from queries in natural languages into SQL query language, and the system is also integrated and adapted with various types of databases and use cases that include financial and healthcare use cases. The experiment results show that the proposed system correctly handles user queries in natural language just like any other neural model based system and more importantly, the proposed SQL generation engine generates SQL queries without syntactic problems with various databases for all queries.
Hangu Yeo, Elahe Khorasani, Vadim Sheinin, Ngoc Phuoc An Vo, Octavian Popescu, Petros Zerfos
IEEE BigData4
2020 Identifying Motion Entities in Natural Language and A Case Study for Named Entity Recognition
abstract
Motion recognition is one of the basic cognitive capabilities of many life forms, however, detecting and understanding motion in text is not a trivial task.In addition, identifying motion entities in natural language is not only challenging but also beneficial for a better natural language understanding.In this paper, we present a Motion Entity Tagging (MET) model to identify entities in motion in a text using the Literal-Motion-in-Text (LiMiT) dataset for training and evaluating the model.Then we propose a new method to split clauses and phrases from complex and long motion sentences to improve the performance of our MET model.We also present results showing that motion features, in particular, entity in motion benefits the Named-Entity Recognition (NER) task.Finally, we present an analysis for the special co-occurrence relation between the person category in NER and animate entities in motion, which significantly improves the classification performance for the person category in NER.
Ngoc Phuoc An Vo, Irene Manotas, Vadim Sheinin, Octavian Popescu
COLING1
2019 Tackling Complex Queries to Relational Databases
Octavian Popescu, Ngoc Phuoc An Vo, Vadim Sheinin, Elahe Khorashani, Hangu Yeo
ACIIDS (1)2
2019 A Natural Language Interface Supporting Complex Logic Questions for Relational Databases
Ngoc Phuoc An Vo, Octavian Popescu, Vadim Sheinin, Elahe Khorasani, Hangu Yeo
NLDB1
2018 A Large Resource of Patterns for Verbal Paraphrases
Octavian Popescu, Ngoc Phuoc An Vo, Vadim Sheinin
LREC2
2018 QUEST: A Natural Language Interface to Relational Databases
Vadim Sheinin, Elahe Khorasani, Hangu Yeo, Ngoc Phuoc An Vo, Octavian Popescu
LREC5
2017 Experimenting word embeddings in assisting legal review
abstract
As advanced technologies, such as data mining become part of the everyday workflow of document reviews in litigations, keyword-search still appears to serve as a cornerstone approach in responsive or privilege review. Keywords are conceptually easy to understand and help culling documents at the early stages of the review. But developing proper keywords to minimize the risk of under/over-inclusiveness can lead to complex strategies. To cope with the burden of designing search terms, we propose to use word embedding techniques in a dynamic manner. This paper describes a system leveraging semantic models in a smart review environment in order to support knowledge workers in eDiscovery.
Ngoc Phuoc An Vo, Caroline Privault, Fabien Guillot
ICAIL1
2016 Multi-layer and Co-learning Systems for Semantic Textual Similarity, Semantic Relatedness and Recognizing Textual Entailment
Ngoc Phuoc An Vo, Octavian Popescu
IC3K1
2016 Corpora for Learning the Mutual Relationship between Semantic Relatedness and Textual Entailment
Ngoc Phuoc An Vo, Octavian Popescu
LREC1
2015 A Preliminary Evaluation of the Impact of Syntactic Structure in Semantic Textual Similarity and Semantic Relatedness Tasks
abstract
The well related tasks of evaluating the Semantic Textual Similarity and Semantic Relatedness have been under a special attention in NLP community. Many different approaches have been proposed, implemented and evaluated at different levels, such as lexical similarity, word/string/POS tags overlapping, semantic modeling (LSA, LDA), etc. However, at the level of syntactic structure, it is not clear how significant it contributes to the overall accuracy. In this paper, we make a preliminary evaluation of the impact of the syntactic structure in the tasks by running and analyzing the results from several experiments regarding to how syntactic structure contributes to solving these tasks.
Ngoc Phuoc An Vo, Octavian Popescu
HLT-NAACL1
2014 Fast and Accurate Misspelling Correction in Large Corpora
abstract
There are several NLP systems whose accuracy depends crucially on finding misspellings fast.However, the classical approach is based on a quadratic time algorithm with 80% coverage.We present a novel algorithm for misspelling detection, which runs in constant time and improves the coverage to more than 96%.We use this algorithm together with a cross document coreference system in order to find proper name misspellings.The experiments confirmed significant improvement over the state of the art.
Octavian Popescu, Ngoc Phuoc An Vo
EMNLP2