VLDB 2026 Research / reviewers in the wild / expert
Thanh Tho Quan
dblp:52/5843 · also Quan Thanh Tho, Than-Tho Quan, Thanh-Tho Quan, Tho Quan 0001, Tho Quan Thanh, Tho Thanh Quan
· DBLP profile ↗
38ranked-venue papers
7as first author
18since 2021 · last 2026
0000-0003-0467-6254ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 17 · 2 first-author · 12 since 2021Software engineering, systems software and programming languages · 8 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 6 · 3 first-author · 2 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Catching the First Light of Tomorrow: A Hackathon-Based Framework for Introducing High School Students to AI AgentsabstractArtificial Intelligence (AI), particularly in the form of intelligent AI agents, is transforming education, industry, and everyday life. These agents extend the capabilities of Large Language Models (LLMs) by integrating planning, decision-making, tool use, and multi-agent collaboration, enabling systems that can reason, adapt, and act in dynamic environments. As such systems become integral to modern workplaces and everyday problem solving, early exposure equips high school students with systems thinking, practical problem-solving skills, and ethical awareness, while preparing them to create applications that address real-world needs. Yet most high school AI programs focus on basic model usage and overlook the skills required to design and deploy agentic systems. Existing resources are largely aimed at university learners and assume substantial programming expertise, creating a significant accessibility gap. To address this need, we present a structured hackathon-based framework for introducing high school students to the design and application of AI agents. The framework combines expert-led lectures on core topics such as agent architectures, prompting strategies, reasoning methods, and tool-use protocols with a guided hackathon in which students collaboratively develop domain-specific agent-based chatbots. We provide a complete suite of instructional materials, step-by-step tutorials, and starter code to support hands-on learning, enabling participants to build functional agents capable of reasoning and interacting with external tools. Our approach bridges the gap between AI literacy and practical deployment while fostering creativity, collaboration, and responsible innovation, and our findings suggest that early engagement with agent-based AI design equips students with both technical proficiency and the mindset to shape the AI-driven future. Long S. T. Nguyen, Quan Bui, Dung Phan, Dung Le, Khang Vo, Nam Duong, Anh Dinh, Tri Trinh, Chi Phan, Thai Nguyen, Dang Le, Vinh Dang, Thanh Tho Quan |
AAAI | 17 |
| 2026 | SaltGuard: A Reproducible Data-to-Model System for Enhancing Salinity Intrusion Analysis in the Mekong Delta
Thao H. M. Tran, Quynh T. N. Vo, Chi N. L. Phan, Thanh Tho Quan |
IEA/AIE (3) | 4 |
| 2026 | T-HiGra: Temporal Reasoning over Hierarchical Knowledge Graphs for Time-Constrained Open-Domain QA
Bao Le, Thanh Tho Quan |
ISMIS | 4 |
| 2026 | Read as You See: Guiding Unimodal LLMs for Low-Resource Explainable Harmful Meme DetectionabstractDetecting harmful memes is crucial for safeguarding the integrity and harmony of online environments, yet existing detection methods are often resource-intensive, inflexible, and lacking explainability, limiting their applicability in assisting real-world web content moderation. We propose U-CoT+, a resource-efficient framework that prioritizes accessibility, flexibility and transparency in harmful meme detection by fully harnessing the capabilities of lightweight unimodal large language models (LLMs). Instead of directly prompting or fine-tuning large multimodal models (LMMs) as black-box classifiers, we avoid immediate reasoning over complex visual inputs but decouple meme content recognition from meme harmfulness analysis through a high-fidelity meme-to-text pipeline, which collaborates lightweight LMMs and LLMs to convert multimodal memes into natural language descriptions that preserve critical visual information, thus enabling text-only LLMs to "see" memes by "reading". Grounded in textual inputs, we further guide unimodal LLMs' reasoning under zero-shot Chain-of-Thoughts (CoT) prompting with targeted, interpretable, context-aware, and easily obtained human-crafted guidelines, thus providing accountable step-by-step rationales, while enabling flexible and efficient adaptation to diverse sociocultural criteria of harmfulness. Extensive experiments on seven benchmark datasets show that U-CoT+ achieves performance comparable to resource-intensive baselines, highlighting its effectiveness and potential as a scalable, explainable, and low-resource solution to support harmful meme detection. Fengjun Pan, Xiaobao Wu, Thanh Tho Quan, Anh Tuan Luu |
WWW | 3 |
| 2025 | ViPhoVQA: Toward a Phonemic-Based Method for Mitigating Rare and Out-of-Vocab Words in Vietnamese Text-Based Visual Question Answering
Nghia Hieu Nguyen, Duc-Vu Nguyen, Huy Quoc To, Dinh Dien, Thanh Tho Quan, Ngan Luu-Thuy Nguyen |
ICCCI (1) | 5 |
| 2025 | A Nature-Inspired Method to Mine Top-k Multi-Level High-Utility ItemsetsabstractHigh-Utility Itemset Mining (HUIM) is designed to discover sets of itemsets that can bring high profits from the database. However, HUIM encounters several challenges in picking a suitable minimum utility threshold for each database. A class of algorithms that select the top-k itemsets based on their utility has been proposed to address this issue. Although traditional top-k HUI mining algorithms do not require a specific threshold, they tend to be very time-consuming and memory-intensive when dealing with large datasets. To tackle the combinational complexity involved in HUIM algorithms, nature-inspired methods have been suggested and adopted. Nonetheless, these algorithms have traditionally focused on handling conventional, often overlooking critical data structures like product hierarchies. Consequently, they fail to extract crucial insight from this novel database format. Thus, our research introduces a heuristic-based algorithm designed to leverage top-k itemsets from databases enriched with item taxonomy data. We propose a technique involving the early pruning of unpromising items to enhance mining efficiency. Experimental evaluations are conducted on several datasets to assess the method’s performance, both with and without adopting this strategy, demonstrating its effectiveness. N. T. Tung, Trinh D. D. Nguyen, Loan T. T. Nguyen, Thanh Tho Quan, An Mai |
Cybern. Syst. | 4 |
| 2025 | Generative Artificial Intelligence for Software Engineering - A Research AgendaabstractABSTRACT Context Generative artificial intelligence (GenAI) tools have become increasingly prevalent in software development, offering assistance to various managerial and technical project activities. Notable examples of these tools include OpenAI's ChatGPT, GitHub Copilot, and Amazon CodeWhisperer. Objective Although many recent publications have explored and evaluated the application of GenAI, a comprehensive understanding of the current development, applications, limitations, and open challenges remains unclear to many. Particularly, we do not have an overall picture of the current state of GenAI technology in practical software engineering usage scenarios. Method We conducted a literature review and focus groups for a duration of five months to develop a research agenda on GenAI for software engineering. Results We identified 78 open research questions (RQs) in 11 areas of software engineering. Our results show that it is possible to explore the adoption of GenAI in partial automation and support decision‐making in all software development activities. While the current literature is skewed toward software implementation, quality assurance and software maintenance, other areas, such as requirements engineering, software design, and software engineering education, would need further research attention. Common considerations when implementing GenAI include industry‐level assessment, dependability and accuracy, data accessibility, transparency, and sustainability aspects associated with the technology. Conclusions GenAI is bringing significant changes to the field of software engineering. Nevertheless, the state of research on the topic still remains immature. We believe that this research agenda holds significance and practical value for informing both researchers and practitioners about current applications and guiding future research. Anh Nguyen-Duc 0001, Beatriz Cabrero-Daniel, Adam Przybylek, Chetan Arora 0002, Dron Khanna, Tomas Herda, Usman Rafiq, Jorge Melegati, Eduardo Guerra 0001, Kai-Kristian Kemell, Mika Saari, Zheying Zhang, Thanh Tho Quan, Pekka Abrahamsson |
Softw. Pract. Exp. | 14 |
| 2024 | LAMPAT: Low-Rank Adaption for Multilingual Paraphrasing Using Adversarial TrainingabstractParaphrases are texts that convey the same meaning while using different words or sentence structures. It can be used as an automatic data augmentation tool for many Natural Language Processing tasks, especially when dealing with low-resource languages, where data shortage is a significant problem. To generate a paraphrase in multilingual settings, previous studies have leveraged the knowledge from the machine translation field, i.e., forming a paraphrase through zero-shot machine translation in the same language. Despite good performance on human evaluation, those methods still require parallel translation datasets, thus making them inapplicable to languages that do not have parallel corpora. To mitigate that problem, we proposed the first unsupervised multilingual paraphrasing model, LAMPAT (Low-rank Adaptation for Multilingual Paraphrasing using Adversarial Training), by which monolingual dataset is sufficient enough to generate a human-like and diverse sentence. Throughout the experiments, we found out that our method not only works well for English but can generalize on unseen languages as well. Data and code are available at https://github.com/phkhanhtrinh23/LAMPAT. Khoi M. Le, Trinh Pham, Thanh Tho Quan, Anh Tuan Luu |
AAAI | 3 |
| 2024 | Revitalizing Bahnaric Language through Neural Machine Translation: Challenges, Strategies, and Promising OutcomesabstractThe Bahnar, a minority ethnic group in Vietnam with ancient roots, hold a language of deep cultural and historical significance. The government is prioritizing the preservation and dissemination of Bahnar language through online availability and cross-generational communication. Recent AI advances, including Neural Machine Translation (NMT), have transformed translation with improved accuracy and fluency, fostering language revitalization through learning, communication, and documentation. In particular, NMT enhances accessibility for Bahnar language speakers, making information and content more available. However, translating Vietnamese to Bahnar language faces practical hurdles due to resource limitations, particularly in the case of Bahnar language as an extremely low-resource language. These challenges encompass data scarcity, vocabulary constraints, and a lack of fine-tuning data. To address these, we propose transfer learning from selected pre-trained models to optimize translation quality and computational efficiency, capitalizing on linguistic similarities between Vietnamese and Bahnar language. Concurrently, we apply tailored augmentation strategies to adapt machine translation for the Vietnamese-Bahnar language context. Our approach is validated through superior results on bilingual Vietnamese-Bahnar language datasets when compared to baseline models. By tackling translation challenges, we help revitalize Bahnar language, ensuring information flows freely and the language thrives. Hoang Nhat Khang Vo, Duc Dong Le, Tran Minh Dat Phan, Tan Sang Nguyen, Quoc Nguyen Pham, Ngoc Oanh Tran, Tran Minh Hieu Vo, Thanh Tho Quan |
AAAI | 9 |
| 2024 | XGA-Osteo: Towards XAI-Enabled Knee Osteoarthritis Diagnosis with Adversarial Learning
Hieu Phan Trung, Loc Le Tan, Mao Nguyen, Phat Nguyen, Sang Nguyen, Minh-Triet Tran, Thanh Tho Quan |
IJCAI | 7 |
| 2024 | Supervised learning models for social bot detection: Literature review and benchmark
Hoang-Dung Nguyen, Cong-Duy Nguyen, Phong T. To, Danh H. Nguyen, Huy Nguyen-Gia, Long H. Tran, Anh Q. Tran, An Dang-Hieu, Anh Nguyen-Duc 0001, Thanh Tho Quan |
Expert Syst. Appl. | 11 |
| 2023 | Expand BERT Representation with Visual Information via Grounded Language Learning with Multimodal Partial AlignmentabstractLanguage models have been supervised with both language-only objective and visual grounding in existing studies of visual-grounded language learning. However, due to differences in the distribution and scale of visual-grounded datasets and language corpora, the language model tends to mix up the context of the tokens that occurred in the grounded data with those that do not. As a result, during representation learning, there is a mismatch between the visual information and the contextual meaning of the sentence. To overcome this limitation, we propose GroundedBERT - a grounded language learning method that enhances the BERT representation with visually grounded information. GroundedBERT comprises two components: (i) the original BERT which captures the contextual representation of words learned from the language corpora, and (ii) a visual grounding module which captures visual information learned from visual-grounded datasets. Moreover, we employ Optimal Transport (OT), specifically its partial variant, to solve the fractional alignment problem between the two modalities. Our proposed method significantly outperforms the baseline language models on various language tasks of the GLUE and SQuAD datasets. Cong-Duy Nguyen, The-Anh Vu-Le, Thong Nguyen 0003, Thanh Tho Quan, Anh Tuan Luu |
ACM Multimedia | 4 |
| 2023 | Scalable maximal subgraph mining with backbone-preserving graph convolutionsabstractMaximal subgraph mining is increasingly important in various domains, including bioinformatics, genomics, and chemistry, as it helps identify common characteristics among a set of graphs and enables their classification into different categories. Existing approaches for identifying maximal subgraphs typically rely on traversing a graph lattice. However, in practice, these approaches are limited to relatively small subgraphs due to the exponential growth of the search space and the NP-completeness of the underlying subgraph isomorphism test. In this work, we propose SCAMA , an approach that addresses these limitations by adopting a divide-and-conquer strategy for efficient mining of maximal subgraphs. Our approach involves initially partitioning a graph database into equivalence classes using bootstrapped backbones, which are tree-shaped frequent subgraphs. We then introduce a learning process based on a novel graph convolutional network (GCN) to extract maximal backbones for each equivalence class. A critical insight of our approach is that by estimating each maximal backbone directly in the embedding space, we can avoid the exponential traversal of the graph lattice. From the extracted maximal backbones, we construct the maximal frequent subgraphs. Furthermore, we outline how SCAMA can be extended to perform top- k largest frequent subgraph mining and how the discovered patterns facilitate graph classification. Our experimental results demonstrate the effectiveness of SCAMA in identifying almost perfectly maximal frequent subgraphs, while exhibiting approximately 10 times faster performance compared to the best baseline technique. Matthias Weidlich 0001, Thanh Tho Quan, Hongzhi Yin, Karl Aberer, Nguyen Quoc Viet Hung |
Inf. Sci. | 4 |
| 2022 | Optimized fuzzy clustering in wireless sensor networks using improved squirrel search algorithm
Kim Khanh Le-Ngoc, Thanh Tho Quan, Thang H. Bui, Amir Masoud Rahmani, Mehdi Hosseinzadeh 0001 |
Fuzzy Sets Syst. | 2 |
| 2021 | Enriching and Controlling Global Semantics for Text SummarizationabstractRecently, Transformer-based models have been proven effective in the abstractive summarization task by creating fluent and informative summaries.Nevertheless, these models still suffer from the short-range dependency problem, causing them to produce summaries that miss the key points of document.In this paper, we attempt to address this issue by introducing a neural topic model empowered with normalizing flow to capture the global semantics of the document, which are then integrated into the summarization model.In addition, to avoid the overwhelming effect of global semantics on contextualized representation, we introduce a mechanism to control the amount of global semantics supplied to the text generation module.Our method outperforms state-of-the-art summarization models on five common text summarization datasets, namely CNN/DailyMail, XSum, Reddit TIFU, arXiv, and PubMed. Thong Nguyen 0003, Anh Tuan Luu, Truc Lu, Thanh Tho Quan |
EMNLP (1) | 4 |
| 2021 | Nested variational autoencoder for topic modelling on microtexts with word vectorsabstractAbstract Most of the information on the Internet is represented in the form of microtexts, which are short text snippets such as news headlines or tweets. These sources of information are abundant, and mining these data could uncover meaningful insights. Topic modelling is one of the popular methods to extract knowledge from a collection of documents; however, conventional topic models such as latent Dirichlet allocation (LDA) are unable to perform well on short documents, mostly due to the scarcity of word co‐occurrence statistics embedded in the data. The objective of our research is to create a topic model that can achieve great performances on microtexts while requiring a small runtime for scalability to large datasets. To solve the lack of information of microtexts, we allow our method to take advantage of word embeddings for additional knowledge of relationships between words. For speed and scalability, we apply autoencoding variational Bayes, an algorithm that can perform efficient black‐box inference in probabilistic models. The result of our work is a novel topic model called the nested variational autoencoder, which is a distribution that takes into account word vectors and is parameterized by a neural network architecture. For optimization, the model is trained to approximate the posterior distribution of the original LDA model. Experiments show the improvements of our model on microtexts as well as its runtime advantage. Trung Trinh, Thanh Tho Quan, Trung Mai |
Expert Syst. J. Knowl. Eng. | 2 |
| 2021 | Structural representation learning for network alignment with self-supervised anchor links
Minh Tam Pham, Thanh Tam Nguyen, Van Vinh Tong, Nguyen Quoc Viet Hung, Thanh Tho Quan |
Expert Syst. Appl. | 7 |
| 2021 | A comprehensive survey and taxonomy of the SVM-based intrusion detection systems
Mokhtar Mohammadi, Tarik A. Rashid, Sarkhel H. Taher Karim, Adil Hussain Mohammed Aldalwie, Thanh Tho Quan, Moazam Bidaki, Amir Masoud Rahmani, Mehdi Hosseinzadeh 0001 |
J. Netw. Comput. Appl. | 5 |
| 2019 | MarCHGen: A framework for generating a malware concept hierarchyabstractAbstract Automatic classification of virus instances into a concept hierarchy has been attracting much attention from malware research community. However, it is definitely not a trivial work, because malwares usually come in binary forms whose actions are complicated and obfuscated. Therefore, the typical data mining approaches based on feature extraction are not easily applied. In this paper, we tackle this problem by introducing a framework known as MarCHGen (Malware Concept Hierarchy Generation). In this framework, we first apply virus logical concept analysis, which incorporates formal concept analysis with temporal logic to capture malware behaviours and generalize a virus concept lattice accordingly. Second, we propose an on‐the‐fly conceptual clustering technique to generate a malware concept hierarchy. In the MarCHGen framework, the malware concept hierarchy will be monitored by the prelarge data set management technique to avoid reclustering several times unnecessarily. Our approach has been applied in a real data set of virus, and promising experimental results have been acquired. Thien Binh Nguyen, Tran Cong Doi, Thanh Tho Quan |
Expert Syst. J. Knowl. Eng. | 3 |
| 2018 | Auto-detection of sophisticated malware using lazy-binding control flow graph and deep learning
Dung Le Nguyen, Xuan Mao Nguyen, Thanh Tho Quan |
Comput. Secur. | 4 |
| 2017 | Probabilistic modelling for congestion detection on wireless sensor networksabstractRecently, Wireless Sensor Networks (WSNs) attract many researches due to their real applications. WSN is actually a network whose main components are sensors and channels. Based on applications, these components can be worked independently or separately with each others to capture information, process and send it to sink. However, in congestion-based aspect, most researches are assumed that environmental working of components are perfect, i.e. they omit packet-loss aspect due to failed sensors or broken links. This causes a limitation to rationally represent a WSN. Thus, in this proposal, using the reliable probability property, we define a Discrete Time Stochastic Petri Net Model for congestion detection on WSN in order to represent all working scenarios for components on the one hand, and calculate the congestion probability in the network on the other hand. After that, we also present a new algorithm to analyse on that model. Our straight example through this paper emphasises the idea of our model. Khanh Le, Giang V. Trinh, Thang H. Bui, Thanh Tho Quan |
CoDIT | 4 |
| 2016 | Smaller to Sharper: Efficient Web Service Composition and Verification Using On-the-fly Model Checking and Logic-Based Clustering
Khai T. Huynh, Thanh Tho Quan, Thang H. Bui |
ICCSA (4) | 2 |
| 2016 | Multi-threaded On-the-Fly Model Generation of Malware with Hash Compaction
Thanh Tho Quan, Le Duc Anh |
ICFEM | 2 |
| 2015 | Goal-oriented dynamic test generation
TheAnh Do, Siau-Cheng Khoo, Russel Pears, Thanh Tho Quan |
Inf. Softw. Technol. | 5 |
| 2014 | PeCAn: Compositional Verification of Petri Nets Made Easy
Dinh-Thuan Le, Huu-Vu Nguyen, Phuong-Nam Mai, Bao-Trung Pham-Duy, Thanh Tho Quan, Étienne André 0001, Laure Petrucci, Yang Liu 0003 |
ATVA | 6 |
| 2013 | Multi-core Model Checking Algorithms for LTL Verification with Fairness AssumptionsabstractThe main challenge in model checking is the state space explosion. With developments in hardware today, most processors have many cores inside. To leverage on the advances in hardware, we can increase the performance of verifying large models by designing parallel algorithms to run efficiently on multi-core architecture. This work focuses on this problem in the context of Linear Temporal Logic (LTL) model checking, which can be seen as finding accepting cycles in a graph. Recently, there are some parallel algorithms based on Nested Depth First Search (NDFS). In this work, we propose two new parallel algorithms based on strongly connected component (SCC) searching algorithm (i.e., Tarjan's algorithm). By finding all the SCCs in the graph, our approaches can not only check LTL properties, but also handle fairness assumptions all together. The experiments show that our new algorithms are comparable or faster than the state-of-the-art multi-core algorithms. Xuan-Linh Ha, Thanh Tho Quan, Yang Liu 0003, Jun Sun 0001 |
APSEC (1) | 2 |
| 2013 | A Hybrid Approach for Control Flow Graph Construction from Binary CodeabstractBinary code analysis has attracted much attention. The difficulty lies in constructing a Control Flow Graph (CFG), which is dynamically generated and modified, such as mutations. Typical examples are handling dynamic jump instructions, in which destinations may be directly modified by rewriting loaded instructions on memory. In this paper, we describe a PhD project proposal on a hybrid approach that combines static analysis and dynamic testing to construct CFG from binary code. Our aim is to minimize false targets produced when processing indirect jumps during CFG construction. To evaluate the potential of our approach, we preliminarily compare results between our method and Jakstab, a state-of-the-art tool in this field. Thien Binh Nguyen, Thanh Tho Quan, Mizuhito Ogawa |
APSEC (2) | 3 |
| 2012 | SeVe: automatic tool for verification of security protocols
Anh Tuan Luu, Jun Sun 0001, Yang Liu 0003, Jin Song Dong 0001, Xiaohong Li 0001, Thanh Tho Quan |
Frontiers Comput. Sci. China | 6 |
| 2011 | Semantic-lite Retrieval on Imprecise and Incomplete Natural Queries using Conceptual Graphs
Thinh Nhat Phan, Thanh Tho Quan, Thien Cong Pham, Nguyen Tuong Huynh |
ICSOFT (1) | 2 |
| 2010 | Semantic Web Service Composition System Supporting Multiple Service Description Languages
Nhan Cach Dang, Duy Ngan Le, Thanh Tho Quan, Minh Nhut Nguyen |
ACIIDS (1) | 3 |
| 2008 | Fuzzy named entity-based document clusteringabstractTraditional keyword-based document clustering techniques have limitations due to simple treatment of words and hard separation of clusters. In this paper, we introduce named entities as objectives into fuzzy document clustering, which are the key elements defining document semantics and in many cases are of user concerns. First, the traditional keyword-based vector space model is adapted with vectors defined over spaces of entity names, types, name-type pairs, and identifiers, instead of keywords. Then, hierarchical fuzzy document clustering can be performed using a similarity measure of the vectors representing documents. For evaluating fuzzy clustering quality, we propose a fuzzy information variation measure to compare two fuzzy partitions. Experimental results are presented and discussed. Tru Hoang Cao, H. T. Do, Dunt T. Hong, Thanh Tho Quan |
FUZZ-IEEE | 4 |
| 2008 | Ontology-Based Natural Query Retrieval Using Conceptual Graphs
Thanh Tho Quan, Siu Cheung Hui |
PRICAI | 1 |
| 2007 | A citation-based document retrieval system for finding research expertise
Thanh Tho Quan, Siu Cheung Hui |
Inf. Process. Manag. | 1 |
| 2006 | Automatic fuzzy ontology generation for semantic help-desk supportabstractCustomer service support is an important operation for most multinational manufacturing companies. With the advancement of internet technologies, customer services nowadays are supported through web-based systems. More recently, rapid development of the semantic web and semantic web services has prompted us to develop a semantic help-desk for supporting customer services over the semantic web environment, which is presented in this paper. In particular, a fuzzy formal concept analysis (FCA)-based approach is developed for automatic generation of fuzzy machine service ontology that can deal with uncertain information. The proposed automatic fuzzy ontology generation technique consists of the following steps: fuzzy formal concept analysis, fuzzy conceptual clustering, and ontology generation. As such, the supporting machine services provided by the proposed system will potentially improve customer satisfaction in terms of reducing machine down time and increasing productivity. In this paper, an experiment has also been conducted for performance evaluation. The experimental result shows that the proposed approach has attained good performance in terms of both accuracy and efficiency when the queries are associated with appropriate membership values, and a suitable confident threshold is set. Thanh Tho Quan, Siu Cheung Hui |
IEEE Trans. Ind. Informatics | 1 |
| 2006 | Automatic Fuzzy Ontology Generation for Semantic WebabstractOntology is an effective conceptualism commonly used for the semantic Web. Fuzzy logic can be incorporated to ontology to represent uncertainty information. Typically, fuzzy ontology is generated from a predefined concept hierarchy. However, to construct a concept hierarchy for a certain domain can be a difficult and tedious task. To tackle this problem, this paper proposes the FOGA (fuzzy ontology generation framework) for automatic generation of fuzzy ontology on uncertainty information. The FOGA framework comprises the following components: fuzzy formal concept analysis, concept hierarchy generation, and fuzzy ontology generation. We also discuss approximating reasoning for incremental enrichment of the ontology with new upcoming data. Finally, a fuzzy-based technique for integrating other attributes of database to the ontology is proposed. Thanh Tho Quan, Siu Cheung Hui, Tru Hoang Cao |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2004 | Automatic Generation of Ontology for Scholarly Semantic Web
Thanh Tho Quan, Siu Cheung Hui, Tru Hoang Cao |
ISWC | 1 |
| 2003 | A Web Mining Approach for Finding Expertise in Research AreasabstractFinding expertise in a research area helps researchers to know whom are the experts working on the research area. This paper proposes a web mining approach for finding expertise in scientific research areas. In this approach, Indexing Agents search and download scientific publications from web sites that typically include academic web pages, then they extract citations and store them in a Web Citation Database. In addition, researcher information is also saved into the Researcher Database. Data mining techniques are applied to the Web Citation Database on citation keywords and authors to form document clusters and author clusters. The Multi-Clustering technique is proposed to mine the combined information of document clusters and author clusters for information on expertise in specified research areas. Thanh Tho Quan, Siu Cheung Hui |
CW | 1 |
| 2003 | Mining Multiple Clustering Data for Knowledge Discovery
Thanh Tho Quan, Siu Cheung Hui |
Discovery Science | 1 |