VLDB 2026 Research / reviewers in the wild / expert
Thanh-Nam Doan
dblp:160/1537
· DBLP profile ↗
15ranked-venue papers
6as first author
8since 2021 · last 2026
0000-0002-6967-2804ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 9 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 8 · 3 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-authorTheory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MANDO-LLM: Heterogeneous Graph Transformers with Large Language Models for Smart Contract Vulnerability DetectionabstractDetecting vulnerabilities in smart contracts is vital for the security and reliability of decentralized apps. To facilitate vulnerability detection, contract codes, including bug patterns, are represented as heterogeneous graphs with various nodes and edges, like control-flow and function-call graphs. However, existing graph-learning techniques struggle with large, complex graphs. This article presents MANDO-LLM, a novel framework that combines heterogeneous graph transformers (HGTs) with large language models (LLMs) for detecting vulnerabilities in smart contracts represented as heterogeneous contract graphs built upon control-flow and call graphs. MANDO-LLM uses LLMs to capture code features from control-flow and call data, customizes HGTs to learn embeddings with specific node-edge meta relations, and employs classifiers for vulnerability detection in Solidity code at both contract and line levels. Our evaluation shows that MANDO-LLM significantly outperforms existing methods on real-world large-scale imbalanced datasets, with F1-score improvements from 0.59% to 80.72% at the contract level. It is also one of the first effective methods for identifying line-level vulnerabilities, with performance boosts ranging from 3.09% to over 95% across different vulnerability types. MANDO-LLM’s versatility allows easy retraining for various vulnerabilities without needing manually defined patterns. Nhat-Minh Nguyen, Long Le Thanh, Zahra Ahmadi, Thanh-Nam Doan, Daoyuan Wu, Lingxiao Jiang |
ACM Trans. Softw. Eng. Methodol. | 5 |
| 2024 | Rethinking Embedding Vectors for Electric Vehicle Charging Stations: An Empirical StudyabstractElectric vehicle (EV) charging stations are critical in promoting EV adoption and mitigating global warming by reducing reliance on fossil fuels. However, a comprehensive understanding of the latent characteristics of these charging stations remains limited. Unveiling these latent features is essential for enhancing predictive tasks such as utilization prediction, demand forecasting, and strategic infrastructure planning. In this paper, we conduct a comprehensive investigation into methods for extracting embedding vectors of charging stations based on userstation interactions. We explore a spectrum of techniques—from traditional approaches like non-negative matrix factorization to advanced machine learning models such as neural collaborative filtering—to effectively capture these latent features. Through extensive experiments and analyses, we evaluate the quality and effectiveness of the generated embeddings in improving predictive modeling tasks related to charging station usage. Our findings demonstrate that incorporating these embeddings significantly enhances the performance of predictive models, leading to more accurate demand forecasts and better utilization predictions. To the best of our knowledge, this is the first study to delve deeply into extracting and analyzing embedding vectors of charging stations derived from user interaction data. The insights gained from this research provide valuable guidance for optimizing EV charging infrastructure and can inform future developments in the field, ultimately supporting the broader adoption of electric vehicles. Seyedmehdi Khaleghian, Thanh-Nam Doan, Joe Knox, Mina Sartipi |
IEEE Big Data | 2 |
| 2023 | HyperRouter: Towards Efficient Training and Inference of Sparse Mixture of ExpertsabstractTruong Do, Le Khiem, Quang Pham, TrungTin Nguyen, Thanh-Nam Doan, Binh Nguyen, Chenghao Liu, Savitha Ramasamy, Xiaoli Li, Steven Hoi. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Truong Do, Le Khiem, Quang Pham, TrungTin Nguyen, Thanh-Nam Doan, Savitha Ramasamy, Xiaoli Li 0001, Steven C. H. Hoi |
EMNLP | 5 |
| 2023 | MANDO-HGT: Heterogeneous Graph Transformers for Smart Contract Vulnerability DetectionabstractSmart contracts in blockchains have been increasingly used for high-value business applications. It is essential to check smart contracts' reliability before and after deployment. Although various program analysis and deep learning techniques have been proposed to detect vulnerabilities in either Ethereum smart contract source code or bytecode, their detection accuracy and scalability are still limited. This paper presents a novel framework named MANDO-HGT for detecting smart contract vulnerabilities. Given Ethereum smart contracts, either in source code or bytecode form, and vulnerable or clean, MANDO-HGT custom-builds heterogeneous contract graphs (HCGs) to represent control-flow and/or function-call information of the code. It then adapts heterogeneous graph transformers (HGTs) with customized meta relations for graph nodes and edges to learn their embeddings and train classifiers for detecting various vulnerability types in the nodes and graphs of the contracts more accurately. We have collected more than 55K Ethereum smart contracts from various data sources and verified the labels for 423 buggy and 2,742 clean contracts to evaluate MANDO-HGT. Our empirical results show that MANDO-HGT can significantly improve the detection accuracy of other state-of-the-art vulnerability detection techniques that are based on either machine learning or conventional analysis techniques. The accuracy improvements in terms of F1-score range from 0.7% to more than 76% at either the coarse-grained contract level or the fine-grained line level for various vulnerability types in either source code or bytecode. Our method is general and can be retrained easily for different vulnerability types without the need for manually defined vulnerability patterns. Nhat-Minh Nguyen, Chunyao Xie, Zahra Ahmadi, Daniel Kudendo, Thanh-Nam Doan, Lingxiao Jiang |
MSR | 6 |
| 2022 | MANDO: Multi-Level Heterogeneous Graph Embeddings for Fine-Grained Detection of Smart Contract VulnerabilitiesabstractLearning heterogeneous graphs consisting of different types of nodes and edges enhances the results of homogeneous graph techniques. An interesting example of such graphs is control-flow graphs representing possible software code execution flows. As such graphs represent more semantic information of code, developing techniques and tools for such graphs can be highly beneficial for detecting vulnerabilities in software for its reliability. However, existing heterogeneous graph techniques are still insufficient in handling complex graphs where the number of different types of nodes and edges is large and variable. This paper concentrates on the Ethereum smart contracts as a sample of software codes represented by heterogeneous contract graphs built upon both control-flow graphs and call graphs containing different types of nodes and links. We propose MANDO, a new heterogeneous graph representation to learn such heterogeneous contract graphs’ structures. MANDO extracts customized meta-paths, which compose relational connections between different types of nodes and their neighbors. Moreover, it develops a multi-metapath heterogeneous graph attention network to learn multi-level embeddings of different types of nodes and their metapaths in the heterogeneous contract graphs, which can capture the code semantics of smart contracts more accurately and facilitate both fine-grained line-level and coarse-grained contract-level vulnerability detection. Our extensive evaluation of large smart contract datasets shows that MANDO improves the vulnerability detection results of other techniques at the coarse-grained contract level. More importantly, it is the first learning-based approach capable of identifying vulnerabilities at the fine-grained line-level, and significantly improves the traditional code analysis-based vulnerability detection approaches by 11.35% to 70.81% in terms of F1-score. Nhat-Minh Nguyen, Chunyao Xie, Zahra Ahmadi, Daniel Kudendo, Thanh-Nam Doan, Lingxiao Jiang |
DSAA | 6 |
| 2022 | SoChainDB: A Database for Storing and Retrieving Blockchain-Powered Social Network DataabstractSocial networks have become an inseparable part of human activities. Most existing social networks follow a centralized system model, which despite storing valuable information of users, arise many critical concerns such as content ownership and over-commercialization. Recently, decentralized social networks, built primarily on blockchain technology, have been proposed as a substitution to eliminate these concerns. Since decentralized architectures are mature enough to be on par with the centralized ones, decentralized social networks are becoming more and more popular. Decentralized social networks can offer both common options like writing posts and comments and more advanced options such as reward systems and voting mechanisms. They provide rich eco-systems for the influencers to interact with their followers and other users via staking systems based on cryptocurrency tokens. The vast and valuable data of the decentralized social networks open several new directions for the research community to extend human behavior knowledge. However, accessing and collecting data from these social networks is not easy because it requires strong blockchain knowledge, which is not the main focus of computer science and social science researchers. Hence, our work proposes the SoChainDB framework that facilitates obtaining data from these new social networks. To show the capacity and strength of SoChainDB, we crawl and publish Hive data - one of the largest blockchain-based social networks. We conduct extensive analyses to understand the insight of Hive data and discuss some interesting applications, e.g., game, non-fungible tokens market built upon Hive. It is worth mentioning that our framework is well-adaptable to other blockchain social networks with minimal modification. SoChainDB is publicly accessible at http://sochaindb.com and the dataset is available under the CC BY-SA 4.0 license. Dmytro Bozhkov, Zahra Ahmadi, Nhat-Minh Nguyen, Thanh-Nam Doan |
SIGIR | 5 |
| 2022 | MANDO-GURU: vulnerability detection for smart contract source code by heterogeneous graph embeddingsabstractSmart contracts are increasingly used with blockchain systems for high-value applications. It is highly desired to ensure the quality of smart contract source code before they are deployed. This paper proposes a new deep learning-based tool, MANDO-GURU, that aims to accurately detect vulnerabilities in smart contracts at both coarse-grained contract-level and fine-grained line-level. Using a combination of control-flow graphs and call graphs of Solidity code, we design new heterogeneous graph attention neural networks to encode more structural and potentially semantic relations among different types of nodes and edges of such graphs and use the encoded embeddings of the graphs and nodes to detect vulnerabilities. Our validation of real-world smart contract datasets shows that MANDO-GURU can significantly improve many other vulnerability detection techniques by up to 24% in terms of the F1-score at the contract level, depending on vulnerability types. It is the first learning-based tool for Ethereum smart contracts that identify vulnerabilities at the line level and significantly improves the traditional code analysis-based techniques by up to 63.4%. Our tool is publicly available at https://github.com/MANDO-Project/ge-sc-machine. A test version is currently deployed at http://mandoguru.com, and a demo video of our tool is available at http://mandoguru.com/demo-video. Nhat-Minh Nguyen, Hong-Phuc Doan, Zahra Ahmadi, Thanh-Nam Doan, Lingxiao Jiang |
ESEC/SIGSOFT FSE | 5 |
| 2021 | TSLib: A Unified Traffic Signal Control Framework Using Deep Reinforcement Learning and BenchmarkingabstractThe volume and velocity of traffic data have in-creased dramatically due to the wide adoption of new technologies such as cameras, Internet-of-Thing devices, and vehicular net-works. That data can help us to optimize Traffic Signal Controls (TSCs) by using adaptive algorithms. Some direct applications of these algorithms are reducing the CO2 emission, fuel consumption, and traveling time. Recently, Deep Reinforcement Learning (DRL) methods are the de-facto solution due to its ability to handle big data with high performance. However, most open source codes and frameworks for TSCs using DRL algorithms have limited flexibility. That causes a difficulty to reuse the codebases for new contexts. Therefore, it will be difficult to have a benchmark for TSCs using different optimization algorithms. For this reason, our paper introduces TSLib – a Python framework for fast prototyping TSCs. Specifically, TSLib is designed as a modular system with high reusability so that researchers can quickly implement and evaluate new ideas of TSCs. Moreover, our work offers a comprehensive implementation of some well known TSCs algorithm including both traditional and DRL-based methods as well as their performance measurements. Toan Tran 0001, Thanh-Nam Doan, Mina Sartipi |
IEEE BigData | 2 |
| 2020 | Understanding the Effect of COVID-19 on Fuel Consumption of Public Transportation: The Case Study of Chattanooga, TNabstractThe COVID-19 pandemic has caused a drastic change in traffic in the U.S and throughout the world. With the drop in traffic volume, the fuel consumption used for traveling should decline. This study focuses on the changes in fuel consumption of public transportation before and during the pandemic. The fuel consumption volumes of diesel bus fleet in Chattanooga were analyzed to identify the changes between the two periods. Our study provides preparation for future disasters. Le Tuan Phan, Thanh-Nam Doan, Mina Sartipi |
IEEE BigData | 2 |
| 2020 | TransCrossCF: Transition-based Cross-Domain Collaborative FilteringabstractThe success of cross-domain recommender systems in capturing user interests across multiple domains has recently brought much attention to them. These recommender systems aim to improve the quality of suggestions and defy the cold-start problem by transferring information from one (or more) source domain(s) to a target domain. However, most cross-domain recommenders ignore the sequential information in user history. They only rely on an aggregate or snapshot of user feedback in the past. Most importantly, they do not explicitly model how users transition from one domain to another domain as users continue to interact with different item domains. In this paper, we argue that between-domain transitions in user sequences are useful in improving recommendation quality, dealing with the cold-start problem, and revealing interesting aspects of how user interests transform from one domain to another. We propose TransCrossCF, transition-based cross-domain collaborative filtering, that can capture both within and between domain transitions of user feedback sequences while understanding the relationship between different item types in different domains. Specifically, we model each purchase of a user as a transition from his/her previous item to the next one, under the effect of item domains and user preferences. Our intensive experiments demonstrate that TransCrossCF outperforms the state-of-the-art methods in recommendation task on three real-world datasets, both in the cold-start and hot-start scenarios. Moreover, according to our context analysis evaluations, the between-domain relations captured by TransCrossCF are interpretable and intuitive. Thanh-Nam Doan, Shaghayegh Sahebi |
ICMLA | 1 |
| 2019 | Rank-Based Tensor Factorization for Student Performance Prediction
Thanh-Nam Doan, Shaghayegh Sahebi |
EDM | 1 |
| 2019 | Modeling location-based social network data with area attraction and neighborhood competition
Thanh-Nam Doan, Ee-Peng Lim |
Data Min. Knowl. Discov. | 1 |
| 2018 | PACELA: A Neural Framework for User Visitation in Location-based Social NetworksabstractCheck-in prediction using location-based social network data is an important research problem for both academia and industry since an accurate check-in predictive model is useful to many applications, e.g. urban planning, venue recommendation, route suggestion, and context-aware advertising. Intuitively, when considering venues to visit, users may rely on their past observed visit histories as well as some latent attributes associated with the venues. In this paper, we therefore propose a check-in prediction model based on a neural framework called Preference and Context Embeddings with Latent Attributes (PACELA). PACELA learns the embeddings space for the user and venue data as well as the latent attributes of both users and venues. More specifically, we use a probabilistic matrix factorization-based technique to infer the latent attributes specific to users and locations in location-based social networks (LBSNs), considering the user visitation decisions that could be affected by area attraction, neighborhood competition, and social homophily. PACELA also includes a deep learning neural network to combine both embedding and latent features to predict if a user performs check-in on a location. Our experiments on three different real world datasets show that PACELA yields the best check-in prediction accuracy against several baseline methods. Thanh-Nam Doan, Ee-Peng Lim |
UMAP | 1 |
| 2017 | Modeling Check-In Behavior with Geographical Neighborhood Influence of Venues
Thanh-Nam Doan, Ee-Peng Lim |
ADMA | 1 |
| 2016 | Attractiveness versus Competition: Towards an Unified Model for User VisitationabstractModeling user check-in behavior provides useful insights about venues as well as the users visiting them. These insights can be used in urban planning and recommender system applications. Unlike previous works that focus on modeling distance effect on user's choice of check-in venues, this paper studies check-in behaviors affected by two venue-related factors, namely, area attractiveness and neighborhood competitiveness. The former refers to the ability of an area with multiple venues to collectively attract check-ins from users, while the latter represents the ability of a venue to compete with its neighbors in the same area for check-ins. We first embark on a data science study to ascertain the two factors using two Foursquare datasets gathered from users and venues in Singapore and Jakarta, two major cities in Asia. We then propose the VAN model incorporating user-venue distance, area attractiveness and neighborhood competitiveness factors. The results from real datasets show that VAN model outperforms the various baselines in two tasks: home location prediction and check-in prediction. Thanh-Nam Doan, Ee-Peng Lim |
CIKM | 1 |