EDBT 2026 Demo / reviewers in the wild / expert
Son T. Mai
dblp:124/7084 · also Son Thai Mai, Thai Son Mai
· DBLP profile ↗
34ranked-venue papers
13as first author
19since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 20 · 12 first-author · 5 since 2021Artificial intelligence and machine learning · 16 · 5 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Systems, architecture and hardware · 3 · 3 since 2021Security and privacy · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | HierarNet: Independent Interactive Hierarchical Disease Outbreak ForecastingabstractEarly warning systems for disease outbreaks play a crucial role in public health for management and contingency planning. However, most predictive modeling works focus on flat models that incorporate exogenous inputs (e.g. climate, demographics) to predict future outbreaks at different locations, but do not jointly model multiple spatial aggregation levels. In this paper, we introduce HierarNet, a unique independent-interactive hierarchical forecasting framework that aims to predict disease outbreaks at different levels of spatial resolution, such as provinces, regions, and nations. HierarNet consists of two main phases. In the local phase, we train independent forecasting models for all locations at all levels. In the global phase, all models iteratively interact with others across different levels via their hierarchical relationships under an ensemble fashion to maximize their agreements. This global local hierarchical interactive scheme makes HierarNet a highly effective and flexible method (i.e. it can work with an arbitrary base prediction model and available exogenous data for each location independently). Extensive experiments are conducted on various disease datasets (e.g., Dengue fever, flu, diarrhea, and Bluetongue) in different countries (e.g., France, Vietnam, and USA) to show the performance of HierarNet compared to 19 state-of-the-art (SOTA) methods such as MinT, DYCHEM, WITRAN, SegRNN, TSMixer, PatchTST, or iTransformer. We also illustrate the generability of HierarNet in other domains, e.g., web traffic forecasting. Zichi Zhang, Phi Hung Nguyen, Ngoc Phu Doan, Viet-Hung Tran, Xuan Hoang Nguyen, Hui Wang 0001, Hans Vandierendonck, Son T. Mai |
AAAI | 8 |
| 2026 | ERL-ASC: Entropy-regularized logarithmic adaptive spectral clustering for multivariate time series
Peixin Li, Yaoguo Dang, Son T. Mai |
Neurocomputing | 5 |
| 2026 | Obstructive Sleep Apnea Prediction: A Comprehensive Review and Comparative StudyabstractAbstract Obstructive Sleep Apnea (OSA) is a highly prevalent sleep disorder linked to considerable public health burdens and comorbidities. However, its heterogeneous presentation and the limited accessibility of traditional diagnostic tools such as polysomnography (PSG) lead to widespread underdiagnosis. As a result, artificial intelligence (AI) approaches, including machine learning (ML) and deep learning (DL) models, have attracted attention as an alternative pathway to detection. This paper first provides a comprehensive review of AI-driven OSA diagnosis, covering different diagnosis problems, input-data types, data biases, pre-processing techniques, and model performance. We then leverage the largest clinical dataset used in OSA prediction to date, approximately 110,000 patients with 22,000 having complete entries for all 50 features, to systematically compare the performance of 39 ML/DL models. Our findings highlight the challenging nature of OSA prediction, with accuracies ranging from 29.66% to 46.9% for 4-class prediction and 46.04% to 87.18% for binary tasks. DL models such as DANet and GATE scored highest, whereas ensemble approaches such as LGBM and AdaBoost displayed more consistent performance across folds. However, as severe cases of OSA are easier to predict and over-represented in datasets, accuracy alone is insufficient for model evaluation and we explore a variety of metrics. Finally, imbalance correction and feature selection improved weaker models, but had only marginal effects on the best-performing models. Looking forwards, the development of more sophisticated and tailored DL models and large, high-quality datasets may help to break current performance barriers. We hope that our work can attract more attention to this challenging but interesting research problem. Huynh Thi Khanh Chi, Amonae Dabbs-Brown, Anna Jurek-Loughrey, James Mulhall, Tuan Dung Pham, Ngoc Phu Doan, Viet-Hung Tran, Zichi Zhang, Xuan Hoang Nguyen, Yimeng An, Peixin Li, Phi Hung Nguyen, Thi Linh Hoang, Xinming Shi, Hans Vandierendonck, Sébastien Bailly, Jean Louis Pépin, Son T. Mai |
Mach. Learn. | 18 |
| 2025 | InteDisUX: Intepretation-Guided Discriminative User-Centric Explanation for Time SeriesabstractExplanation for deep learning models on time series classification (TSC) tasks is an important and challenging problem. Most existing approaches use attribution maps to explain outcomes. However, they have limitations in generating explanations that are well-aligned with humans's perceptions. Recently LIME-based approaches provide a more meaningful explanation via segmenting the data. However, these approaches are still suffering from the processes of segment generations and evaluations. In this paper, we propose a novel time series explanation approach called InteDisUX to overcome these problems. Our technique utilizes the segment-level integrated gradient (SIG) for calculating importance scores for an initial set of small and equal segments before iteratively merge two consecutive ones to create better explanations under a unique greedy strategy guided by two new proposed metrics including discrimination and faithfulness gains. By this way, our method does not depend on predefined segments like others while being robusts to instability, poor local fidelity and data imbalance like LIME-based methods. Furthermore, InteDisUX is the first work to use the model's information to improve the set of segments} for time series explanation. Extensive experiments show that our method outperforms LIME-based ones in 12 datasets in terms of faithfulness and 8/12 datasets in terms of robustness. Viet-Hung Tran, Zichi Zhang, Tuan Dung Pham, Ngoc Phu Doan, Anh-Tuan Hoang, Peixin Li, Hans Vandierendonck, Ira Assent, Son T. Mai |
AAAI | 9 |
| 2025 | WaveletMixer: A Multi-Resolution Wavelets Based MLP-Mixer for Multivariate Long-Term Time Series ForecastingabstractTime Series Forecasting (TSF) aims at predicting future values for a time series data and plays a crucial role in many real-world applications, e.g., finance, disease spread, or weather predictions. However, it is also a very challenging task due to complex temporal dependencies in the data, especially for long-term forecasting. In this paper, we introduce WaveletMixer, an iterative multi-levels, multi-resolutions and multi-phases approach to effectively capture long-term dependencies of multivariate time series in both global and local perspectives for improving forecasting performance. WaveletMixer fundamentally differs from existing works in the following key aspects. First, it exploits multi-levels properties of Wavelet transformation to create multiple forecasting models for different frequency domains at various levels of resolutions. Second, the relationships among different frequency domains are exploited to iteratively adjust all prediction models at all levels simultaneously in both local and global perspectives to reduce prediction errors and biases, thus significantly improving the final accuracy. Third, while WaveletMixer is a general framework that can be used to boost the performance of any deep-learning architecture (e.g., MLP, LSTM or Transformer), we additionally introduce TS-Learner, an MLP-based model to further enhance the performance in long-term forecasting. Extensive experiments have been conducted on nine real-world datasets to demonstrate the outstanding performance of WaveletMixer compared to SOTA methods and to reveal its important characteristics. Zichi Zhang, Tuan Dung Pham, Yimeng An, Ngoc Phu Doan, Majed Alsharari, Viet-Hung Tran, Anh-Tuan Hoang, Hans Vandierendonck, Son T. Mai |
AAAI | 9 |
| 2025 | ParaGrapher: A Parallel and Distributed Graph Loading Library for Large-Scale Compressed GraphsabstractWhereas the literature describes an increasing number of graph algorithms, loading graphs remains a time-consuming component of the end-to-end execution time. Graph frameworks often rely on custom graph storage formats, that are not optimized for efficient loading of large-scale graph datasets. Furthermore, graph loading is often not optimized as it is time-consuming to implement. This shows a demand for high-performance libraries capable of efficiently loading graphs to (i) accelerate designing new graph algorithms, (ii) to evaluate the contributions across a wide range of graph datasets, and (iii) to facilitate easy and fast comparisons across different graph frameworks. We present ParaGrapher, a library for loading large-scale compressed graphs in parallel and distributed graph frameworks. ParaGrapher supports (a) loading the graph while the caller is blocked and (b) interleaving graph loading with graph processing. ParaGrapher is designed to support loading graphs in shared-memory, distributed-memory, and out-of-core graph processing. We explain the design of ParaGrapher and present a performance model of graph decompression. Our evaluation shows that ParaGrapher delivers up to 3.2 times speedup in loading and up to 5.2 times speedup in end-to-end execution (i.e., through interleaved loading and execution). Mohsen Koohi Esfahani, Syed Ibtisam Tauhidi, Marco D'Antonio, Son T. Mai, Hans Vandierendonck |
IEEE Big Data | 4 |
| 2025 | DeepPUFSCA: Deep learning for Physical Unclonable Function attack based on Side Channel Analysis supportabstractPhysical Unclonable Function (PUF) poses a vulnerability that it could be imitated by machine learning attacks and side channel attacks, which break its physical uniqueness and unpredictable characteristic. Hence, many works are concerned with enhancing PUF design by introducing more nonlinear modules inside to differentiate approximating PUF behavior from the attacker side. However, the safety of these PUFs are still an open area and need to be verified. In this paper, we propose DeepPUFSCA, which is a deep learning-based model that uniquely combines both challenge and side-channel information features during training to attack PUF. To gather the data, we conduct a design of an arbiter PUF on FPGA and measure its power consumption. Our intensive experiments on this dataset demonstrate that DeepPUFSCA outperforms other machine learning-based methods in terms of attacking accuracy, even the novel ensemble algorithms. Moreover, we also show that combined side channel information boosts the model performance compared to attacking with challenge-response only. Ngoc Phu Doan, Tuan Dung Pham, Zichi Zhang, Viet-Hung Tran, Jack Miskelly, Hans Vandierendonck, Anh-Tuan Hoang, Máire O'Neill, Son T. Mai |
DAC | 9 |
| 2025 | MIX: A Multi-view Time-Frequency Interactive Explanation Framework for Time Series ClassificationabstractDeep learning models for time series classification (TSC) have achieved impressive performance, but explaining their decisions remains a significant challenge. Existing post-hoc explanation methods typically operate solely in the time domain and from a single-view perspective, limiting both faithfulness and robustness. In this work, we propose MIX (Multi-view Time-Frequency Interactive EXplanation Framework), a novel framework that helps to explain deep learning models in a multi-view setting by leveraging multi-resolution, time-frequency views constructed using the Haar Discrete Wavelet Transform (DWT). MIX introduces an interactive cross-view refinement scheme, where explanation's information from one view is propagated across views to enhance overall interpretability. To align with user-preferred perspectives, we propose a greedy selection strategy that traverses the multi-view space to identify the most informative features. Additionally, we present OSIGV, a user-aligned segment-level attribution mechanism based on overlapping windows for each view, and introduce keystone-first IG, a method that refines explanations in each view using additional information from another view. Extensive experiments across multiple TSC benchmarks and model architectures demonstrate that MIX significantly outperforms state-of-the-art (SOTA) methods in terms of explanation faithfulness and robustness. Viet-Hung Tran, Ngoc Phu Doan, Zichi Zhang, Tuan Dung Pham, Phi Hung Nguyen, Xuan Hoang Nguyen, Hans Vandierendonck, Ira Assent, Son T. Mai |
NeurIPS | 9 |
| 2025 | Wasp: Efficient Asynchronous Single-Source Shortest Path on Multicore Systems via Work StealingabstractThe Single-Source Shortest Path (SSSP) problem is a fundamental graph problem with an extensive set of real-world applications. State-of-the-art parallel algorithms for SSSP, such as the Δ -stepping algorithm, create parallelism through priority coarsening. Priority coarsening results in redundant computations that diminish the benefits of parallelization and limit parallel scalability. Marco D'Antonio, Son T. Mai, Philippas Tsigas, Hans Vandierendonck |
SC | 2 |
| 2025 | QL-PGD: An efficient defense against membership inference attack
Tuan Dung Pham, Bao Dung Nguyen, Son T. Mai, Viet Cuong Ta |
J. Inf. Secur. Appl. | 3 |
| 2025 | Efficient Integer-Only-Inference of Gradient Boosting Decision Trees on Low-Power DevicesabstractThere is increasingly interest in developing embedded machine learning hardware as it can offer better performance in terms of privacy, bandwidth efficiency, and scalability. Gradient-boosted decision trees (GBDT) represent a strong candidate as they employ less complex logic, but their efficient implementation in field programmable gate array (FPGA) needs to be explored in detail. In this paper, we propose sophisticated quantisation approaches to balance the dual goals of efficiency and performance. In particular, we introduce quantisation-aware training of GBDT for integer-only and binary arithmetic. Results are presented for implementations on a Zynq UltraScale+ MPSoC FPGA with the best design using only 170 Look-up Tables and 233 flip-flops at a clock speed of 724 MHz. Implementations focused on network intrusion detection and jet substructure classification for large-scale physics experiments are explored. An order of magnitude less FPGA resources are used whilst offering extremely high throughput rate and maintaining accuracy. Code is available athttps://github.com/malsharari/QATGBDT. Majed Alsharari, Son T. Mai, Roger F. Woods, Carlos Reaño |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2024 | Preserving Spatial-Temporal Relationship with Adaptive Node Sampling in Hierarchical Dynamic Graph Transformers
Thi Linh Hoang, Tuan Dung Pham, Son T. Mai, Viet Cuong Ta |
ACML | 3 |
| 2024 | Dynamic weighted ensemble for diarrhoea incidence predictions
Thanh Duy Do, Thuan Dinh Nguyen, Viet Cuong Ta, Duong Tran Anh, Tuyet-Hanh Tran Thi, Diep Phan, Son T. Mai |
Mach. Learn. | 7 |
| 2024 | LCSL: Long-Tailed Classification via Self-LabelingabstractDuring the last decades, deep learning (DL) has been proven to be a very powerful and successful technique in many real-world applications, e.g., video surveillance or object detection. However, when class label distributions are highly skewed, DL classifiers tend to be biased towards majority classes during training phases. This leads to poor generalization of minority classes and consequently reduces the overall accuracy. How to effectively deal with this long-tailed class distribution in DL, i.e., deep long-tailed classification (DLC), remains a challenging problem despite many research efforts. Among various approaches, data augmentation, which aims at generating more samples for reducing label imbalance, is the most common and practical one. However, simply relying on existing class-agnostic augmentation strategies without properly considering the label differences would worsen the problem since more head-class samples can be inevitably augmented than tail-class ones. Moreover, none of the existing works consider the quality and suitability of augmented samples during the training process. Our proposed approach, called Long-tailed Classification via Self-Labeling (LCSL), is specifically designed to address these limitations. LCSL fundamentally differs from existing works by the way it iteratively exploits the preceding network during the training process to re-label the labeled augmented samples and uses the output confidence to decide whether new samples belong to minority classes before adding them to the data. Not only does this help to reduce imbalance ratios among classes, but this also helps to reduce the uncertainty of class prediction problems for minority classes by selecting more confident samples to the data. This incremental learning and generating scheme thus provide a new robust approach for decreasing model over-fitting, thus enhancing the overall accuracy, especially for minority classes. Extensive experiments have demonstrated that LCSL acquires better performance than state-of-the-art long-tailed learning techniques on various standard benchmark datasets. More specifically, our LCSL obtains 85.8%, 54.4%, and 56.2% in terms of accuracy on CIFAR10-LT, CIFAR100-LT, and ImageNet-LT (with moderate to extreme imbalance ratios), respectively. The source code is available athttps://github.com/vdquang1991/lcsl/. Duc-Quang Vu, Trang T. T. Phung, Jia-Ching Wang, Son T. Mai |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2023 | Efficient and Effective Multi-Modal Queries through Heterogeneous Network Embedding (Extended Abstract)abstractRecent information retrieval (IR) systems answer a multi-modal query by considering it as a set of separate uni-modal queries. However, depending on the chosen operationalisation, such an approach is inefficient or ineffective. It either requires multiple passes over the data or leads to inaccuracies since the relations between data modalities are neglected in the relevance assessment. To mitigate these challenges, we present an IR system that has been designed to answer genuine multi-modal queries. It relies on a heterogeneous network embedding, so that features from diverse modalities can be incorporated when representing both, a query and the data over which it shall be evaluated. An experimental evaluation using diverse real-world and synthetic datasets illustrates that our approach returns twice the amount of relevant information compared to baseline techniques, while scaling to large multi-modal databases. Thanh Tam Nguyen, Chi Thang Duong, Hongzhi Yin, Matthias Weidlich 0001, Son T. Mai, Karl Aberer, Nguyen Quoc Viet Hung |
ICDE | 5 |
| 2023 | Detecting rumours with latency guarantees using massive streaming data
Thanh Tam Nguyen, Hongzhi Yin, Matthias Weidlich 0001, Thanh Thi Nguyen 0001, Son T. Mai, Nguyen Quoc Viet Hung |
VLDB J. | 6 |
| 2022 | Dengue Fever: From Extreme Climates to Outbreak PredictionabstractDengue Fever (DF) is an emerging mosquito-borne infectious disease that affects hundred of millions of people each year with considerable morbidity and mortality rates, especially for children. Together with global climate changes, it is continuously increasing in terms of number of cases and new locations. Thus, having effective early warning systems becomes an urgent need to improve disease controls and prevention. In this paper, we introduce a novel framework, called Proximity Time Ensemble, to predict DF outbreaks for multiple areas (provinces) and multiple time steps ahead, and to study the effects of climate data on DF outbreaks. PT-Ensem consists of 6 key components: (1) an event-to-event probabilistic framework to study links among extreme climate events and DF outbreaks; (2) a proximity graph that connects similar provinces; (3) an ensemble prediction technique that combines many different advanced machine learning (ML) methods to predict outbreaks within t time steps in the future using extreme climate events as model inputs; (4) a data aggregate scheme to enrich training data for each province via its neighbors in the proximity graph; (5) a proximity propagation step that propagates predicted results among similar provinces via the proximity graph until maximal agreements are reached among provinces; and (6) a time propagation step to propagate results via different predicted time steps in each province. We use PT-Ensem to predict DF outbreaks for all provinces in Vietnam using data collected from 1997-2016. Experiments show that PT-Ensem acquires significant performance boost compared to many highly-rated ML models like XGBoost, LightGBM and Catboost in the outbreak prediction task. Compared to most recent deep learning approaches like LSTM-ATT, LSTM, CNN and Transformer for predicting DF incidence, PT-Ensem also dominates in both prediction accuracy and computation times. Son T. Mai, Ha T. Phi, Peter Kilpatrick, Hung Q. V. Nguyen, Hans Vandierendonck |
ICDM | 1 |
| 2022 | Incremental Density-Based Clustering on Multicore ProcessorsabstractThe density-based clustering algorithm is a fundamental data clustering technique with many real-world applications. However, when the database is frequently changed, how to effectively update clustering results rather than reclustering from scratch remains a challenging task. In this work, we introduce IncAnyDBC, a unique parallel incremental data clustering approach to deal with this problem. First, IncAnyDBC can process changes in bulks rather than batches like state-of-the-art methods for reducing update overheads. Second, it keeps an underlying cluster structure called the object node graph during the clustering process and uses it as a basis for incrementally updating clusters wrt. inserted or deleted objects in the database by propagating changes around affected nodes only. In additional, IncAnyDBC actively and iteratively examines the graph and chooses only a small set of most meaningful objects to produce exact clustering results of DBSCAN or to approximate results under arbitrary time constraints. This makes it more efficient than other existing methods. Third, by processing objects in blocks, IncAnyDBC can be efficiently parallelized on multicore CPUs, thus creating a work-efficient method. It runs much faster than existing techniques using one thread while still scaling well with multiple threads. Experiments are conducted on various large real datasets for demonstrating the performance of IncAnyDBC. Son T. Mai, Jon Jacobsen, Sihem Amer-Yahia, Ivor T. A. Spence, Nhat-Phuong Tran, Ira Assent, Nguyen Quoc Viet Hung |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | Efficient and Effective Multi-Modal Queries Through Heterogeneous Network EmbeddingabstractThe heterogeneity of today’s Web sources requires information retrieval (IR) systems to handle multi-modal queries. Such queries define a user’s information needs by different data modalities, such as keywords, hashtags, user profiles, and other media. Recent IR systems answer such a multi-modal query by considering it as a set of separate uni-modal queries. However, depending on the chosen operationalisation, such an approach is inefficient or ineffective. It either requires multiple passes over the data or leads to inaccuracies since the relations between data modalities are neglected in the relevance assessment. To mitigate these challenges, we present an IR system that has been designed to answer genuine multi-modal queries. It relies on a heterogeneous network embedding, so that features from diverse modalities can be incorporated when representing both, a query and the data over which it shall be evaluated. By embedding a query and the data in the same vector space, the relations across modalities are made explicit and exploited for more accurate query evaluation. At the same time, multi-modal queries are answered with a single pass over the data. An experimental evaluation using diverse real-world and synthetic datasets illustrates that our approach returns twice the amount of relevant information compared to baseline techniques, while scaling to large multi-modal databases. Chi Thang Duong, Thanh Tam Nguyen, Hongzhi Yin, Matthias Weidlich 0001, Son T. Mai, Karl Aberer, Nguyen Quoc Viet Hung |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2019 | Scalable Interactive Dynamic Graph Clustering on Multicore CPUsabstractThe structural graph clustering algorithm SCAN is a fundamental technique for managing and analyzing graph data. However, its high runtime remains a computational bottleneck, which limits its applicability. In this paper, we propose a novel interactive approach for tackling this problem on multicore CPUs. Our algorithm, called anySCAN, iteratively processes vertices in blocks. The acquired results are merged into an underlying cluster structure consisting of the so-called super-nodes for building clusters. During its runtime, anySCAN can be suspended for examining intermediate results and resumed for finding better results at arbitrary time points, making it an anytime algorithm which is capable of handling very large graphs in an interactive way and under arbitrary time constraints. Moreover, its block processing scheme allows the design of a scalable parallel algorithm on shared memory architectures such as multicore CPUs for speeding up the algorithm further at each iteration. Consequently, anySCAN uniquely is a both interactive and work-efficient parallel algorithm. We further introduce danySCAN an efficient bulk update scheme for anySCAN on dynamic graphs in which the clusters are updated in bulks and in a parallel interactive scheme. Experiments are conducted on very large real graph datasets for demonstrating the performance of anySCAN. They show its ability to acquire very good approximate results early, leading to orders of magnitude speedup compared to SCAN and its variants. Moreover, it scales very well with the number of threads when dealing with both static and dynamic graphs. Son T. Mai, Sihem Amer-Yahia, Ira Assent, Mathias Skovgaard Birk, Martin Storgaard Dieu, Jon Jacobsen, Jesper Kristensen |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2018 | Scalable Active Constrained Clustering for Temporal Data
Son T. Mai, Sihem Amer-Yahia, Ahlame Douzal Chouakria, Ky T. Nguyen, Anh-Duong Nguyen |
DASFAA (1) | 1 |
| 2018 | Scalable Active Temporal Constrained ClusteringabstractInternational audience Son T. Mai, Sihem Amer-Yahia, Ahlame Douzal Chouakria |
EDBT | 1 |
| 2018 | Evolutionary Active Constrained Clustering for Obstructive Sleep Apnea AnalysisabstractWe introduce a novel interactive framework to handle both instance-level and temporal smoothness constraints for clustering large longitudinal data and for tracking the cluster evolutions over time. It consists of a constrained clustering algorithm, called CVQE+ , which optimizes the clustering quality, constraint violation and the historical cost between consecutive data snapshots. At the center of our framework is a simple yet effective active learning technique, named Border , for iteratively selecting the most informative pairs of objects to query users about, and updating the clustering with new constraints. Those constraints are then propagated inside each data snapshot and between snapshots via two schemes, called constraint inheritance and constraint propagation , to further enhance the results. Moreover, a historical constraint is enforced between consecutive snapshots to ensure the consistency of results among them. Experiments show better or comparable clustering results than state-of-the-art techniques as well as high scalability for large datasets. Finally, we apply our algorithm for clustering phenotypes in patients with Obstructive Sleep Apnea as well as for tracking how these clusters evolve over time. Son T. Mai, Sihem Amer-Yahia, Sébastien Bailly, Jean Louis Pépin, Ahlame Douzal Chouakria, Ky T. Nguyen, Anh-Duong Nguyen |
Data Sci. Eng. | 1 |
| 2018 | Anytime parallel density-based clustering
Son T. Mai, Ira Assent, Jon Jacobsen, Martin Storgaard Dieu |
Data Min. Knowl. Discov. | 1 |
| 2018 | Health Monitoring on Social Media over TimeabstractSocial media has become a major source for analyzing all aspects of daily life. Thanks to dedicated latent topic analysis methods such as the Ailment Topic Aspect Model (ATAM), public health can now be observed on Twitter. In this work, we are interested in using social media to monitor people's health overtime. The use of tweets has several benefits including instantaneous data availability at virtually no cost. Early monitoring of health data is complementary to post-factum studies and enables a range of applications such as measuring behavioral risk factors and triggering health campaigns. We formulate two problems: health transition detection and health transition prediction. We first propose the Temporal Ailment Topic Aspect Model (TM-ATAM), a new latent model dedicated to solving the first problem by capturing transitions that involve health-related topics. TM-ATAM is a non-obvious extension to ATAM that was designed to extract health-related topics. It learns health-related topic transitions by minimizing the prediction error on topic distributions between consecutive posts at different time and geographic granularities. To solve the second problem, we develop T-ATAM, a Temporal Ailment Topic Aspect Model where time is treated as a random variable natively inside ATAM. Our experiments on an 8-month corpus of tweets show that TM-ATAM outperforms TM-LDA in estimating health-related transitions from tweets for different geographic populations. We examine the ability of TM-ATAM to detect transitions due to climate conditions in different geographic regions. We then show how T-ATAM can be used to predict the most important transition and additionally compare T-ATAM with CDC (Center for Disease Control) data and Google Flu Trends. Sumit Sidana, Sihem Amer-Yahia, Marianne Clausel, Majdeddine Rebai, Son T. Mai, Massih-Reza Amini |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2017 | Interactive Exploration of Subspace Clusters for High Dimensional Data
Jesper Kristensen, Son T. Mai, Ira Assent, Jon Jacobsen, Bay Vo |
DEXA (1) | 2 |
| 2017 | Scalable and Interactive Graph Clustering Algorithm on Multicore CPUsabstractThe structural graph clustering algorithm SCAN is a fundamental technique for managing and analyzing graph data. However, its high runtime remains a computational bottleneck, which limits its applicability. In this paper, we propose a novel interactive approach for tackling this problem on multicore CPUs. Our algorithm, called anySCAN, iteratively processes vertices in blocks. The acquired results are merged into an underlying cluster structures consisting of the so-called supernodes for building clusters. During its runtime, anySCAN can be suppressed for examining intermediate results and resumed for finding better result at arbitrary time points, making it an anytime algorithm which is capable to deal with very large graphs in an interactive way and under arbitrary time constraints. Moreover, its block processing scheme allows the design of a scalable parallel algorithm on shared memory architectures such as multicore CPUs for further speeding up the algorithm at each iteration. Consequently, anySCAN uniquely is an interactive and parallel algorithm at the same time. Experiments are conducted on very large real graph datasets for demonstrating the performance of anySCAN. It acquires very good approximate results early, leading to orders of magnitude speedup factor compared to SCAN and its variants. Using 16 threads, the acquired speed up factors are up to 13.5 times over its sequential version. Son T. Mai, Martin Storgaard Dieu, Ira Assent, Jon Jacobsen, Jesper Kristensen, Mathias Skovgaard Birk |
ICDE | 1 |
| 2016 | Anytime OPTICS: An Efficient Approach for Hierarchical Density-Based Clustering
Son T. Mai, Ira Assent |
DASFAA (1) | 1 |
| 2016 | AnyDBC: An Efficient Anytime Density-based Clustering Algorithm for Very Large Complex DatasetsabstractThe density-based clustering algorithm DBSCAN is a state-of-the-art data clustering technique with numerous applications in many fields. However, its O(n2) time complexity still remains a severe weakness. In this paper, we propose a novel anytime approach to cope with this problem by reducing both the range query and the label propagation time of DBSCAN. Our algorithm, called AnyDBC, compresses the data into smaller density-connected subsets called primitive clusters and labels objects based on connected components of these primitive clusters for reducing the label propagation time. Moreover, instead of passively performing the range query for all objects like existing techniques, AnyDBC iteratively and actively learns the current cluster structure of the data and selects a few most promising objects for refining clusters at each iteration. Thus, in the end, it performs substantially fewer range queries compared to DBSCAN while still guaranteeing the exact final result of DBSCAN. Experiments show speedup factors of orders of magnitude compared to DBSCAN and its fastest variants on very large real and synthetic complex datasets. Son T. Mai, Ira Assent, Martin Storgaard Dieu |
KDD | 1 |
| 2015 | Anytime density-based clustering of complex data
Son T. Mai, Xiao He 0002, Jing Feng 0003, Claudia Plant, Christian Böhm 0001 |
Knowl. Inf. Syst. | 1 |
| 2014 | Relevant overlapping subspace clusters on categorical dataabstractClustering categorical data poses some unique challenges: Due to missing order and spacing among the categories, selecting a suitable similarity measure is a difficult task. Many existing techniques require the user to specify input parameters which are difficult to estimate. Moreover, many techniques are limited to detect clusters in the full-dimensional data space. Only few methods exist for subspace clustering and they produce highly redundant results. Therefore, we propose ROCAT (Relevant Overlapping Subspace Clusters on Categorical Data), a novel technique based on the idea of data compression. Following the Minimum Description Length principle, ROCAT automatically detects the most relevant subspace clusters without any input parameter. The relevance of each cluster is validated by its contribution to compress the data. Optimizing the trade-off between goodness-of-fit and model complexity, ROCAT automatically determines a meaningful number of clusters to represent the data. ROCAT is especially designed to detect subspace clusters on categorical data which may overlap in objects and/or attributes; i.e. objects can be assigned to different clusters in different subspaces and attributes may contribute to different subspaces containing clusters. ROCAT naturally avoids undesired redundancy in clusters and subspaces by allowing overlap only if it improves the compression rate. Extensive experiments demonstrate the effectiveness and efficiency of our approach. Xiao He 0002, Jing Feng 0003, Bettina Konte, Son T. Mai, Claudia Plant |
KDD | 4 |
| 2013 | Active Density-Based ClusteringabstractThe density-based clustering algorithm DBSCAN is a fundamental technique for data clustering with many attractive properties and applications. However, DBSCAN requires specifying all pair wise (dis)similarities among objects that can be non-trivial to obtain in many applications. To tackle this problem, in this paper, we propose a novel active density-based clustering algorithm, named Act-DBSCAN, which works under a restricted number of used pair wise similarities. Act-DBSCAN exploits the pair wise lower-bounding (LB) similarities to initialize the cluster structure. Then, it adaptively selects the most informative pair wise LB similarities to update with the real ones in order to reconstruct the result until the budget limitation is reached. The goal is to approximate as much as possible the true clustering result with each update. Our Act-DBSCAN framework is built upon a proposed probabilistic model to score the impact of the update of each pair wise LB similarity on the change of the intermediate clustering structure. Deriving from this scoring system and the monotonicity and reduction property of our active clustering process, we propose the two efficient algorithms to iteratively select and update pair wise similarities and cluster structure. Experiments on real datasets show that Act-DBSCAN acquires good clustering results with only a few pair wise similarities, and requires only a small fraction of all pair wise similarities to reach the DBSCAN results. Act-DBSCAN also outperforms other related techniques such as active spectral clustering. Son T. Mai, Xiao He 0002, Nina C. Hubig, Claudia Plant, Christian Böhm 0001 |
ICDM | 1 |
| 2013 | Efficient Anytime Density-based ClusteringabstractMany clustering algorithms suffer from scalability problems on massive datasets and do not support any user interaction during runtime. To tackle these problems, anytime clustering algorithms are proposed. They produce a fast approximate result which is continuously refined during the further run. Also, they can be stopped or suspended anytime and provide an answer. In this paper, we propose a novel anytime clustering algorithm based on the density-based clustering paradigm. Our algorithm called A-DBSCAN is applicable to very high dimensional databases such as time series, trajectory, medical data, etc. The general idea of our algorithm is to use a sequence of lower-bounding functions (LBs) of the true similarity measure to produce multiple approximate results of the true density-based clusters. A-DBSCAN operates in multiple levels w.r.t. the LBs and is mainly based on two algorithmic schemes: (1) an efficient distance upgrade scheme which restricts distance calculations to core-objects at each level of the LBs; (2) a local re-clustering scheme which restricts update operations to the relevant objects only. Extensive experiments demonstrate that A-DBSCAN acquires very good clustering results at very early stages of execution thus saves a large amount of computational time. Even if it runs to the end, A-DBSCAN is still orders of magnitude faster than DBSCAN. Christian Böhm 0001, Jing Feng 0003, Xiao He 0002, Son T. Mai |
SDM | 4 |
| 2012 | A Similarity Model and Segmentation Algorithm for White Matter Fiber TractsabstractRecently, fiber segmentation has become an emerging technique in neuroscience. Grouping fiber tracts into anatomical meaningful bundles allows to study the structure of the brain and to investigate onset and progression of neurodegenerative and mental diseases. In this paper, we propose a novel technique for fiber tracts based on shape similarity and connection similarity. For shape similarity, we propose some new techniques adapted from existing similarity measures for trajectory data. We also propose a new technique called Warped Longest Common Subsequence (WLCS) for which we additionally developed a lower-bounding distance function to speed up the segmentation process. Our segmentation is based on an outlier-robust density-based clustering algorithm. Extensive experiments on diffusion tensor images demonstrate the efficiency and effectiveness of our technique. Son T. Mai, Sebastian Goebl, Claudia Plant |
ICDM | 1 |