Hoang Thanh Lam

dblp:66/3866 · also Thanh Lam Hoang · DBLP profile ↗
← Back
23ranked-venue papers
14as first author
10since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 15 · 11 first-author · 4 since 2021Artificial intelligence and machine learning · 14 · 9 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 A study on functionality validation for windows malware mutating using reinforcement learning
Do Thi Thu Hien, Le Viet Tai Man, Le Trong Nhan, Phan Ngoc Yen Nhi, Hoang Thanh Lam, Cam Nguyen Tan, Van-Hau Pham
Inf. Softw. Technol.5
2025 A new framework for evaluating model out-of-distribution generalisation for the biochemical domain
abstract
Quantifying model generalization to out-of-distribution data has been a longstanding challenge in machine learning. Addressing this issue is crucial for leveraging machine learning in scientific discovery, where models must generalize to new molecules or materials. Current methods typically split data into train and test sets using various criteria — temporal, sequence identity, scaffold, or random cross-validation — before evaluating model performance. However, with so many splitting criteria available, existing approaches offer limited guidance on selecting the most appropriate one, and they do not provide mechanisms for incorporating prior knowledge about the target deployment distribution(s). To tackle this problem, we have developed a novel metric, AU-GOOD, which quantifies expected model performance under conditions of increasing dissimilarity between train and test sets, while also accounting for prior knowledge about the target deployment distribution(s), when available. This metric is broadly applicable to biochemical entities, including proteins, small molecules, nucleic acids, or cells; as long as a relevant similarity function is defined for them. Recognizing the wide range of similarity functions used in biochemistry, we propose criteria to guide the selection of the most appropriate metric for partitioning. We also introduce a new partitioning algorithm that generates more challenging test sets, and we propose statistical methods for comparing models based on AU-GOOD. Finally, we demonstrate the insights that can be gained from this framework by applying it to two different use cases: developing predictors for pharmaceutical properties of small molecules, and using protein language models as embeddings to build biophysical property predictors.
Raúl Fernández-Díaz, Hoang Thanh Lam, Vanessa López, Denis C. Shields
ICLR2
2025 Enhancing foundation models for scientific discovery via multimodal knowledge graph representations
abstract
Foundation Models (FMs) hold transformative potential to accelerate scientific discovery, yet reaching their full capacity in complex, highly multimodal domains such as genomics, drug discovery, and materials science requires a deeper consideration of the contextual nature of the scientific knowledge . We revisit the synergy between FMs and Multimodal Knowledge Graph (MKG) representation and learning, exploring their potential to enhance predictive and generative tasks in biomedical contexts like drug discovery. We seek to exploit MKGs to improve generative AI models’ ability to capture intricate domain-specific relations and facilitate multimodal fusion. This integration promises to accelerate discovery workflows by providing more meaningful multimodal knowledge-enhanced representations and contextual evidence. Despite this potential, challenges and opportunities remain, including fusing multiple sequential, structural and knowledge modalities and models leveraging the strengths of each; developing scalable architectures for multi-task multi-dataset learning; creating end-to-end workflows to enhance the trustworthiness of biomedical FMs using knowledge from heterogeneous datasets and scientific literature; the domain data bottleneck and the lack of a unified representation between natural language and chemical representations; and benchmarking, specifically the transfer learning to tasks with limited data (e.g., unseen molecules and proteins, rear diseases). Finally, fostering openness and collaboration is key to accelerate scientific breakthroughs.
Vanessa López, Hoang Thanh Lam, Marcos Martínez Galindo, Raúl Fernández-Díaz, Marco Luca Sbodio, Rodrigo Ordonez-Hurtado, Mykhaylo Zayats, Natasha Mulligan, Joao H. Bettencourt-Silva
J. Web Semant.2
2024 Knowledge Enhanced Representation Learning for Drug Discovery
abstract
Recent research on predicting the binding affinity between drug molecules and proteins use representations learned, through unsupervised learning techniques, from large databases of molecule SMILES and protein sequences. While these representations have significantly enhanced the predictions, they are usually based on a limited set of modalities, and they do not exploit available knowledge about existing relations among molecules and proteins. Our study reveals that enhanced representations, derived from multimodal knowledge graphs describing relations among molecules and proteins, lead to state-of-the-art results in well-established benchmarks (first place in the leaderboard for Therapeutics Data Commons benchmark ``Drug-Target Interaction Domain Generalization Benchmark", with an improvement of 8 points with respect to previous best result). Moreover, our results significantly surpass those achieved in standard benchmarks by using conventional pre-trained representations that rely only on sequence or SMILES data. We release our multimodal knowledge graphs, integrating data from seven public data sources, and which contain over 30 million triples. Pretrained models from our proposed graphs and benchmark task source code are also released.
Hoang Thanh Lam, Marco Luca Sbodio, Marcos Martínez Galindo, Mykhaylo Zayats, Raúl Fernández-Díaz, Víctor Valls, Gabriele Picco, Cesar Berrospi, Vanessa López
AAAI1
2024 TabularFM: An Open Framework For Tabular Foundational Models
abstract
Foundational models (FMs), pretrained on extensive datasets using self-supervised techniques, are capable of learning generalized patterns from large amounts of data. This reduces the need for extensive labeled datasets for each new task, saving both time and resources by leveraging the broad knowledge base established during pretraining. Most research on FMs has primarily focused on unstructured data, such as text and images, or semi-structured data, like time-series. However, there has been limited attention to structured data, such as tabular data, which, despite its prevalence, remains under-studied due to a lack of clean datasets and insufficient research on the transferability of FMs for various tabular data tasks. In response to this gap, we introduce a framework called TabularFM1, which incorporates state-of-the-art methods for developing FMs specifically for tabular data. This includes variations of neural architectures such as GANs, VAEs, and Transformers. We have curated a thousand tabular datasets and released cleaned versions to facilitate the development of tabular FMs. We pretrained FMs on this curated data, benchmarked various learning methods on these datasets, and released the pretrained models along with leaderboards for future comparative studies. Our fully open-sourced system provides a comprehensive analysis of the transferability of tabular FMs.
Quan M. Tran, Suong N. Hoang, Lam M. Nguyen, Dzung T. Phan, Hoang Thanh Lam
IEEE Big Data5
2024 AutoPeptideML: a study on how to build more trustworthy peptide bioactivity predictors
abstract
MOTIVATION: Automated machine learning (AutoML) solutions can bridge the gap between new computational advances and their real-world applications by enabling experimental scientists to build their own custom models. We examine different steps in the development life-cycle of peptide bioactivity binary predictors and identify key steps where automation cannot only result in a more accessible method, but also more robust and interpretable evaluation leading to more trustworthy models. RESULTS: We present a new automated method for drawing negative peptides that achieves better balance between specificity and generalization than current alternatives. We study the effect of homology-based partitioning for generating the training and testing data subsets and demonstrate that model performance is overestimated when no such homology correction is used, which indicates that prior studies may have overestimated their performance when applied to new peptide sequences. We also conduct a systematic analysis of different protein language models as peptide representation methods and find that they can serve as better descriptors than a naive alternative, but that there is no significant difference across models with different sizes or algorithms. Finally, we demonstrate that an ensemble of optimized traditional machine learning algorithms can compete with more complex neural network models, while being more computationally efficient. We integrate these findings into AutoPeptideML, an easy-to-use AutoML tool to allow researchers without a computational background to build new predictive models for peptide bioactivity in a matter of minutes. AVAILABILITY AND IMPLEMENTATION: Source code, documentation, and data are available at https://github.com/IBM/AutoPeptideML and a dedicated web-server at http://peptide.ucd.ie/AutoPeptideML. A static version of the software to ensure the reproduction of the results is available at https://zenodo.org/records/13363975.
Raúl Fernández-Díaz, Rodrigo Cossio-Pérez, Clement Agoni, Hoang Thanh Lam, Vanessa López, Denis C. Shields
Bioinform.4
2023 Attacking c-MARL More Effectively: A Data Driven Approach
abstract
In recent years, a proliferation of methods were developed for cooperative multi-agent reinforcement learning (c-MARL). However, the robustness of c-MARL agents against adversarial attacks has been rarely explored. In this paper, we propose to evaluate the robustness of c-MARL agents via a model-based approach, named c-MBA. Our proposed formulation can craft much stronger adversarial state perturbations of c-MARL agents to lower total team rewards than existing model-free approaches. In addition, we propose the first victim-agent selection strategy and the first data-driven approach to define targeted failure states where each of them allows us to develop even stronger adversarial attack without the expert knowledge to the underlying environment. Our numerical experiments on two representative MARL benchmarks illustrate the advantage of our approach over other baselines: our model-based attack consistently outperforms other baselines in all tested environments.
Nhan H. Pham, Lam M. Nguyen, Jie Chen 0007, Hoang Thanh Lam, Subhro Das, Tsui-Wei Weng
ICDM4
2022 Maximum Bayes Smatch Ensemble Distillation for AMR Parsing
abstract
Young-Suk Lee, Ramón Astudillo, Hoang Thanh Lam, Tahira Naseem, Radu Florian, Salim Roukos. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Young-Suk Lee 0001, Ramón Fernandez Astudillo, Hoang Thanh Lam, Tahira Naseem, Radu Florian, Salim Roukos
NAACL-HLT3
2021 Automated Data Science for Relational Data
abstract
Feature engineering is a crucial but tedious task that requires up to 80% of the total time in data science projects. A significant challenge is when data consists of tables from different data sources, thus data scientists need to wisely aggregate and join tables while performing feature engineering task. In this work, we demonstrate a novel system called OneBM (One Button Machine), that enables data scientists to increase their efficiency with automated feature engineering for relational data. OneBM takes as input a relational dataset with multiple tables and its entity relation diagram (ERD) which can be declared with a novel, easy-to-use drag-and-drop graphical user interface. The system then automatically identifies and executes relevant joins and aggregates in the data, and generates new features with a rich set of transformations for various types of data including but not limited to time-series, sequences, number sets and itemsets, etc. The generated features then can be used by automated model selection and hyper-parameter optimization algorithms to complete a fully end-to-end automated data science (or AutoDS) workflow. A follow-up user evaluation illustrated how data scientists can perform multi-table feature engineering tasks in minutes using our system, compared to repeatedly coding SQL-like queries to transform and aggregate relational data requiring weeks of manual labor for comparable performance. In the live demos we plan to show two use cases with real-world datasets (video demos are available at the links in the footnote): sale prediction1and call center user experience2. Pre-registered partcipants can play with these use-cases and the given datasets via Watson Studio on the cloud.
Hoang Thanh Lam, Beat Buesser, Hong Min, Tran Ngoc Minh, Martin Wistuba, Udayan Khurana, Gregory Bramble, Theodoros Salonidis, Dakuo Wang, Horst Samulowitz
ICDE1
2021 Ensembling Graph Predictions for AMR Parsing
abstract
In many machine learning tasks, models are trained to predict structure data such as graphs. For example, in natural language processing, it is very common to parse texts into dependency trees or abstract meaning representation (AMR) graphs. On the other hand, ensemble methods combine predictions from multiple models to create a new one that is more robust and accurate than individual predictions. In the literature, there are many ensembling techniques proposed for classification or regression problems, however, ensemble graph prediction has not been studied thoroughly. In this work, we formalize this problem as mining the largest graph that is the most supported by a collection of graph predictions. As the problem is NP-Hard, we propose an efficient heuristic algorithm to approximate the optimal solution. To validate our approach, we carried out experiments in AMR parsing problems. The experimental results demonstrate that the proposed approach can combine the strength of state-of-the-art AMR parsers to create new predictions that are more accurate than any individual models in five standard benchmark datasets.
Hoang Thanh Lam, Gabriele Picco, Yufang Hou 0001, Young-Suk Lee 0001, Lam M. Nguyen, Dzung T. Phan, Vanessa López, Ramón Fernandez Astudillo
NeurIPS1
2019 Layered convolutional dictionary learning for sparse coding itemsets
Sameen Mansha, Hoang Thanh Lam, Hongzhi Yin, Faisal Kamiran, Mohsen Ali
World Wide Web2
2016 A concise summary of spatial anomalies and its application in efficient real-time driving behaviour monitoring
abstract
This work is motivated by a smart car application which analyses streams of data generated from cars to enhance transportation safety. We treated the problem as real-time abnormal driving behaviour detection using spatio-temporal data collected from mobile devices including GPS location, speed and steering angle. A concise summary was proposed to summarise spatial patterns from GPS trajectory data for efficient real-time anomaly detection. An approach solving this problem by nearest neighbour search has O(n) space and O(log(n) + k) query time complexity, where k is the neighbourhood size and n is the data size. On the other hand, the concise summary approach requires only O(ε * n) memory space and has O(log(ε * n)) query time complexity, where k is several orders of magnitude smaller than one. Experiments with two large datasets from Porto and Beijing showed that our method used only a few megabytes to summarise datasets with n = 80 million data points and was able to process 30K queries per second which was several orders of magnitude faster than the baseline approach. Besides, in the work, interesting spatio-temporal patterns regarding abnormal driving behaviours from the real-world datasets are also discussed to demonstrate potential application of the work in many industries including insurance, transportation safety enhancement and city transport management.
Hoang Thanh Lam
SIGSPATIAL/GIS1
2016 The SPMF Open-Source Data Mining Library Version 2
Philippe Fournier-Viger, Jerry Chun-Wei Lin, Antonio Gomariz, Ted Gueniche, Azadeh Soltani, Zhi-Hong Deng 0001, Hoang Thanh Lam
ECML/PKDD (3)7
2015 Flexible Sliding Windows for Kernel Regression Based Bus Arrival Time Prediction
Hoang Thanh Lam, Eric Bouillet
ECML/PKDD (3)1
2014 Online event clustering in temporal dimension
abstract
This work is motivated by a real-life application that exploits sensor data available from traffic light control systems currently deployed in many cities around the world. Each sensor consists of an induction loop that generates a stream of events triggered whenever a metallic object e.g. car, bus, or a bicycle, is detected above the sensor. Because of the red phase of traffic lights objects are usually divided into groups that move together. Detecting these groups of objects as long as they pass through the sensor is useful for estimating the status of the toad networks such as car queue length or detecting traffic anomalies. In this work, given a data stream that contains observations of an event, e.g. detection of a moving object, together with the timestamps indicating when the events happen, we study the problem that clusters the events together in real-time based on the proximity of the event's occurrence time. We propose an efficient real-time algorithm that scales up to the large data streams extracted from thousands of sensors in the city of London. Moreover, our algorithm is better than the baseline algorithms in terms of clustering accuracy. We demonstrate motivations of the work by showing a real-life use-case in which clustering results are used for estimating the car queue lengths on the road and detecting traffic anomalies.
Hoang Thanh Lam, Eric Bouillet
SIGSPATIAL/GIS1
2014 Mining Top-K Largest Tiles in a Data Stream
Hoang Thanh Lam, Wenjie Pei, Adriana Prado, Baptiste Jeudy, Élisa Fromont
ECML/PKDD (2)1
2012 Mining Compressing Sequential Patterns
abstract
Compression based pattern mining has been successfully applied to many data mining tasks. We propose an approach based on the minimum description length principle to extract sequential patterns that compress a database of sequences well. We show that mining compressing patterns is NP-Hard and belongs to the class of inapproximable problems. We propose two heuristic algorithms to mining compressing patterns. The first uses a two-phase approach similar to Krimp for itemset data. To overcome performance with the required candidate generation we propose GoKrimp, an effective greedy algorithm that directly mines compressing patterns. We conduct an empirical study on six real-life datasets to compare the proposed algorithms by run time, compressibility, and classification accuracy using the patterns found as features for SVM classifiers.
Hoang Thanh Lam, Fabian Mörchen, Dmitriy Fradkin, Toon Calders
SDM1
2011 Online Discovery of Top-k Similar Motifs in Time Series Data
abstract
A motif is a pair of non-overlapping sequences with very similar shapes in a time series. We study the online top-k most similar motif discovery problem. A special case of this problem corresponding to k = 1 was investigated in the literature by Mueen and Keogh [2]. We generalize the problem to any k and propose space-efficient algorithms for solving it. We show that our algorithms are optimal in term of space. In the particular case when k = 1, our algorithms achieve better performance both in terms of space and time consumption than the algorithm of Mueen and Keogh. We demonstrate our results by both theoretical analysis and extensive experiments with both synthetic and real-life data. We also show possible application of the top-k similar motifs discovery problem.
Hoang Thanh Lam, Toon Calders, Ninh Pham
SDM1
2010 An Incremental Prefix Filtering Approach for the All Pairs Similarity Search Problem
abstract
Given a set of records, a threshold value t and a similarity function, we investigate the problem of finding all pairs of records such that similarity between each pair is above t. We propose several optimizations on the existing approaches to solve the problem. Our algorithm outperforms the state-of-the-art algorithms in the case with large and high-dimensional datasets. The speedup we achieved varied from 30% to 4-x depending on the similarity threshold and the dataset properties.
Hoang Thanh Lam, Dinh Viet Dung, Raffaele Perego 0001, Fabrizio Silvestri
APWeb1
2010 Mining top-k frequent items in a data stream with flexible sliding windows
abstract
We study the problem of finding the k most frequent items in a stream of items for the recently proposed max-frequency measure. Based on the properties of an item, the max-frequency of an item is counted over a sliding window of which the length changes dynamically. Besides being parameterless, this way of measuring the support of items was shown to have the advantage of a faster detection of bursts in a stream, especially if the set of items is heterogeneous. The algorithm that was proposed for maintaining all frequent items, however, scales poorly when the number of items becomes large. Therefore, in this paper we propose, instead of reporting all frequent items, to only mine the top-k most frequent ones. First we prove that in order to solve this problem exactly, we still need a prohibitive amount of memory (at least linear in the number of items). Yet, under some reasonable conditions, we show both theoretically and empirically that a memory-efficient algorithm exists. A prototype of this algorithm is implemented and we present its performance w.r.t. memory-efficiency on real-life data and in controlled experiments with synthetic data.
Hoang Thanh Lam, Toon Calders
KDD1
2010 On Using Query Logs for Static Index Pruning
abstract
Static index pruning techniques aim at removing from the posting lists of an inverted file the references to documents which are likely to be not relevant for answering user queries. The reduction in the size of the index results in a better exploitation of memory hierarchies and faster query processing. On the other hand, pruning may affect the precision of the information retrieval system, since pruned entries are unavailable at query processing time. Static pruning techniques proposed so far exploit query-independent measures to evaluate the importance of a document within a posting list. This paper proposes a general framework that aims at enhancing the precision of any static pruning methods by exploiting usage information extracted from query logs. Experiments conducted on the TREC WT10g Web collection and a large Altavista query log show that integrating usage knowledge into the pruning process is profitable, and increases remarkably performance figures obtained with the state-of-the art Carmel's static pruning method.
Hoang Thanh Lam, Raffaele Perego 0001, Fabrizio Silvestri
Web Intelligence1
2009 Entry Pairing in Inverted File
Hoang Thanh Lam, Raffaele Perego 0001, Quan Thoi Minh Nguyen, Fabrizio Silvestri
WISE1
2007 A heuristic particle swarm optimization
abstract
A heuristic version of the particle swarm optimization (PSO) is introduced in this paper. In this new method called "The heuristic particle swarm optimization(HPSO)", we use heuristics to choose the next particle to update its velocity and position. By using heuristics , the convergence rate to local minimum is faster. To avoid premature convergence of the swarm, the particles are re-initialized with random velocity when moving too close to the global best position. The combination of heuristics and re-initialization mechanism make HPSO outperform the basic PSO and recent versions of PSO.
Hoang Thanh Lam, Nina N. Popova, Quan Thoi Minh Nguyen
GECCO1