VLDB 2026 Research / reviewers in the wild / expert
Tuan-Dung Cao
dblp:49/1711
· DBLP profile ↗
12ranked-venue papers
4as first author
4since 2021 · last 2023
0000-0002-3661-9142ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 6 · 3 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-authorComputer networks · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
1 paper |
Information extraction and text analysis · 100% |
Topics — the 3 heaviest of 3, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Information extraction and text analysis › sentiment analysis
aspect-based sentiment analysis |
0.7 | 1 | 2023 | A Self-enhancement Multitask Framework for Unsupervised Aspect Category Detection · EMNLP 2023 |
Natural language and speech › Information extraction and text analysis › sentiment analysis › aspect-based sentiment analysis
aspect category detection |
0.7 | 1 | 2023 | A Self-enhancement Multitask Framework for Unsupervised Aspect Category Detection · EMNLP 2023 |
Natural language and speech › Information extraction and text analysis › sentiment analysis › aspect-based sentiment analysis
aspect extraction |
0.7 | 1 | 2023 | A Self-enhancement Multitask Framework for Unsupervised Aspect Category Detection · EMNLP 2023 |
Methods — techniques the papers use, named apart from their topics
seed word expansion · 0.7multi-task learning · 0.7data augmentation · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | A Self-enhancement Multitask Framework for Unsupervised Aspect Category DetectionabstractOur work addresses the problem of unsupervised Aspect Category Detection using a small set of seed words.Recent works have focused on learning embedding spaces for seed words and sentences to establish similarities between sentences and aspects.However, aspect representations are limited by the quality of initial seed words, and model performances are compromised by noise.To mitigate this limitation, we propose a simple framework that automatically enhances the quality of initial seed words and selects high-quality sentences for training instead of using the entire dataset.Our main concepts are to add a number of seed words to the initial set and to treat the task of noise resolution as a task of augmenting data for a low-resource task.In addition, we jointly train Aspect Category Detection with Aspect Term Extraction and Aspect Term Polarity to further enhance performance.This approach facilitates shared representation learning, allowing Aspect Category Detection to benefit from the additional guidance offered by other tasks.Extensive experiments demonstrate that our framework surpasses strong baselines on standard datasets. Thi-Nhung Nguyen, Hoang Ngo, Kiem-Hieu Nguyen, Tuan-Dung Cao |
EMNLP | 4 |
| 2023 | An Automatic Method for Building a Taxonomy of Areas of Expertise
Thi Thu Le, Tuan-Dung Cao, Xuan-Lam Xuan, Trung Duc Pham, Toan Luu |
ICAART (3) | 2 |
| 2022 | A Practical Method for Occupational Skills Detection in Vietnamese Job Listings
Viet-Trung Tran, Hai-Nam Cao, Tuan-Dung Cao |
ACIIDS (1) | 3 |
| 2022 | Synonym Prediction for Vietnamese Occupational Skills
Hai-Nam Cao, Duc-Thai Do, Viet-Trung Tran, Tuan-Dung Cao, Young-In Song |
IEA/AIE | 4 |
| 2020 | A Novel Method for Recognizing Vietnamese Voice Commands on Smartphones with Support Vector Machine and Convolutional Neural NetworksabstractThis paper will present a new method of identifying Vietnamese voice commands using Google speech recognition (GSR) service results. The problem is that the percentage of correct identifications of Vietnamese voice commands in the Google system is not high. We propose a supervised machine-learning approach to address cases in which Google incorrectly identifies voice commands. First, we build a voice command dataset that includes hypotheses of GSR for each corresponding voice command. Next, we propose a correction system using support vector machine (SVM) and convolutional neural network (CNN) models. The results show that the correction system reduces errors in recognizing Vietnamese voice commands from 35.06% to 7.08% using the SVM model and 5.15% using the CNN model. Quang H. Nguyen 0001, Tuan-Dung Cao |
Wirel. Commun. Mob. Comput. | 2 |
| 2014 | Toward a platform for building and exploiting semantic annotation of photo taken with smart phone
Tuan-Dung Cao, Thi-Nhu Nguyen |
iiWAS | 1 |
| 2014 | Automatic creation of semantic data about football transfer in sport newsabstractThe semantic annotation of textual content is important for the success of many information integration systems. In this paper, we present a method for generating semantic annotations about transfer in football news. Semantic web technologies are applied to extend the knowledge base of KIM platform so that named entities in this specific domain could be recognized. We study and define language models to recognize the semantic relation for football transfers. At the same time, we propose a pronoun recognition method using extraction rules to improve the relation recognition process. The experiment showed promising results on the data set built from Sky Sports news [27]. The precisions achieved in both cases, with and without integration of the pronoun recognition method, are both over 80%. In particular, the latter helps increase the recall value to around 10%. Minh Nguyen-Quang, Tuan-Dung Cao, Thanh-Tam Nguyen |
iiWAS | 2 |
| 2013 | The VHO Project: A Semantic Solution for Vietnamese History Search System
Dang-Hung Phan, Tuan-Dung Cao |
ACIIDS (1) | 2 |
| 2013 | A Method for the Generation of Semantic Annotation from Sport News Using Ontology Based PatternsabstractIn the framework of this research, we focus on describing the algorithm to collect, process information in natural language to turn it into semantic annotations based on ontology, patterns and the knowledge ground in sport. This paper will concentrate to solve the problem of enhancing the named entities recognition by developing ontology and enriching the knowledge base for the recognition engine. Besides this, it is at the same time solving the problem of semantic recognition in texts using ontology-based patterns, harnessing the power of ontology in identifying synonyms, aliases. Minh Nguyen-Quang, Tuan-Dung Cao, Thanh-Hien Phan, Hoang-Cong Nguyen, Tatsuya Hagino |
KES-AMSTA | 2 |
| 2011 | Integrating open data and generating travel itinerary in semantic-aware tourist information systemabstractThe growth of online data and services on the Web make it become more and more emerging as an indispensable tool of traveling for the tourist industry. It is not denied that various approaches bring benefits for visitors in supporting them of searching tourist attractions, such us interested places for the visit, eating or staying. However, like a coin has two sides, too much information would be the difficulty for people when planning their journeys. Generally, tourists usually have problems when finding a satisfied accommodation without a reference to nearby restaurants, sights or event locations. In addition, travelers suffer from the information overload when they look for information about potential destinations, events and related services. Providing the relevant and up-to-date information for the tourists with different personal interests is still a challenging task for the tourist guide information systems. Tuan-Dung Cao, Minh Nguyen-Quang, Tan-Hung Le |
iiWAS | 1 |
| 2011 | Improving travel information access with semantic search application on mobile environmentabstractIn last decade, we witness a significant increase in performance of mobile devices. In parallel, the growth of online data and services available on the Web make it become more and more emerging as a gold mine for the tourist guide application. With the inherent mobility and the widely supported Internet connectivity, tourist guide applications on smart phone are currently potential substitutes or supplements to travel book. However, it is a nightmare for tourists to select what are really interesting in a sheer volume of information while mobile devices have limited interactivity due to the inconvenient keyboard and small screen size regarding to computer. Therefore, the goal of our research is to develop a smart tool helping tourists find relevant travel information with minimum effort. This tool, implemented on Android platform, is the main component of STAAR (Semantic Tourist informAtion Access and Recommending), a system that supports request from heterogeneous environments. For the purpose, we describe our ontology formalized in RDFS/OWL to describe travel related information and to support integrating data from Linked data source. Through ontology manipulating web services, our Android application generates dynamic ontology-based user interfaces, allowing tourist to express their need in semantic queries with different levels of complexity, and then get access to relevant information. Tuan-Dung Cao, Nguyen-Dat Tuan |
MoMM | 1 |
| 2005 | A semantic web approach for building technology-monitoring systemabstractThe information explosion especially on the World Wide Web has made it become the most immense information source and a gold mine for corporations. In the field of technology monitoring, it is very important to be able to get relevant information from heterogeneous sources on the Web. The arrival of Semantic Web technologies promises intelligent retrieval and access to information through the use of semantic annotations based on relevant ontologies. In a scenario of technology monitoring, the automatic generation of semantic annotations of Web document has a significant role. In this article, after describing a new approach based on Semantic Web technologies for building a technology monitoring system, we present an ontology-based algorithm for automatic search and annotation of Web documents. This algorithm will be encapsulated by agents in a multi agent technology monitoring system. Tuan-Dung Cao, Rose Dieng, Marc Bourdeau, Bruno Fiés |
K-CAP | 1 |