Son N. Tran

dblp:139/1001 · DBLP profile ↗
← Back
27ranked-venue papers
6as first author
18since 2021 · last 2026
0000-0002-5912-293XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 19 · 5 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-author · 6 since 2021Databases, data management, data science and information retrieval · 3 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Theory of computation · 1
YearPublicationVenuePosition
2026 The 7th International Workshop on Intelligent Cross-Data Analysis and Retrieval
abstract
With the rapid growth of sensors, communication technologies, and social networks, it is now easy to collect large amounts of data from people and their surrounding environments, such as wearable devices, lifelog cameras, and ambient sensors. These data provide both personal and external views of human activities, but most existing work still focuses on analyzing each type of data separately. As a result, there is a clear gap in understanding how to combine and use cross-data across different sources. This workshop provides a platform for researchers from academia and industry to explore cross-data analysis and retrieval, with a focus on practical challenges such as integrating different data types, handling distributed data, and ensuring data security, aiming to support the development of a smart and sustainable society.
Minh-Son Dao, Duc-Tien Dang-Nguyen, Son N. Tran
ICMR3
2025 ICDAR 25: Intelligent Cross-Data Analysis and Retrieval
abstract
The sixth edition of the Intelligent Cross-Data Analysis and Retrieval (ICDAR) workshop continues to serve as a forum for researchers and practitioners addressing the integration, analysis, and retrieval of heterogeneous data sources. While individual modalities such as wearable sensors, lifelogging cameras, and social media have been well studied, analyzing cross-data that incorporates multiple perspectives remains a crucial yet challenging task for advancing human-centered applications. In 2025, the workshop received 19 submissions, of which 7 were accepted following a careful peer-review process, resulting in an acceptance rate of 37%. The accepted papers covered a wide range of topics, including zero-shot composed image retrieval, vision-language scene understanding, adaptive modality fusion, lightweight fine-tuning with truncated SVD, and real-world federated split learning on mobile devices. By fostering interdisciplinary collaboration across domains such as well-being, disaster mitigation, mobility, food computing, and smart cities, the workshop continues to highlight emerging challenges and solutions for building intelligent, sustainable, and human-centric systems driven by cross-modal and multimodal data analytics.
Takahiro Komamizu, Marc A. Kastner 0001, Minh-Son Dao, Michael Riegler 0001, Duc-Tien Dang-Nguyen, Son N. Tran
ICMR6
2025 Parallel Multi-Scale Deep Supervision Net for Hand Key Point Detection
abstract
Key point detection plays an important role in a wide range of applications. However, predicting key points of small objects such as human hands is a challenging problem. Recent works fuse feature maps of deep Convolutional Neural Networks (CNNs), either via multi-level feature integration or multi-resolution aggregation. Despite achieving some success, the feature fusion approaches increase the complexity and the opacity of CNNs. To address this issue, we propose a novel CNN model named Parallel Multi-Scale Deep Supervision Network (P-MSDSNet) that learns feature maps at different scales in parallel with deep supervisions to produce spatial attention maps for adaptive feature propagation from layer to layer. PMSDSNet has a multi-stage with a parallel structure that fuses multi-scale features from both the same and different depth levels. The deep supervision with spatial attention would enhance relevant features and help improve the transparency of the feature learning at each stage. In the experiment, we show that P-MSDSNet outperforms the state-of-the-art approaches on benchmark datasets while requiring fewer parameters. We also demonstrate the applicability of P-MSDSNet to quantifying finger-tapping hand movements in a neuroscience study.
Renjie Li 0001, Son N. Tran, Saurabh Kumar Garg 0001, Katherine Lawler, Jane E. Alty, Quan Bai 0001
IEEE Trans. Big Data2
2024 Enhance Statistical Features with Changepoint Detection for Driver Behaviour Analysis
Jamal Maktoubian, Son N. Tran, Anna Shillabeer, Muhammad Bilal Amin, Lawrence Sambrooks
PRICAI (4)2
2024 A novel adaptive ensemble learning framework for automated Beggiatoa Spp. coverage estimation
abstract
The presence of Beggiatoa Spp. indicates anoxic conditions or ‘poor condition’ in marine sediments beneath aquaculture pens, resulting from organic enrichment. Currently, the most efficient approach to estimate Beggiatoa Spp. coverage, and thus the extent of the issue, involves video surveys which are scored by human observers for presence of this bacteria. However, this approach is highly time-consuming and relies heavily on the expertise and experience of the individuals involved, thus affecting its accuracy. Machine learning-based computer vision techniques, such as Convolutional Neural Networks (CNNs), offer the potential for automated estimation of Beggiatoa Spp. coverage. However, most existing machine learning methods focus solely on the estimation of the coverage via presence, absence of a single type of Beggiatoa Spp.. These approaches typically rely on binary classification to distinguish the object from the background when estimating coverage. Nevertheless, the inclusion of subordinate categories within high-level classifications poses a great challenge for accurately estimating their coverage rates. In this paper, an adaptive ensemble learning approach was proposed to estimate Beggiatoa Spp. coverage. Unlike other approaches, the proposed approach is capable in adaptively extracting and fusing features from underwater images and accurately estimating the coverages of multiple types of Beggiatoa Spp. through ensemble learning. Experimental results demonstrated that our proposed approach outperforms other approaches in terms of both effectiveness and efficiency in estimating Beggiatoa Spp. coverage.
Yanyu Chen 0001, Yunjue Zhou, Mira Park 0001, Son N. Tran, Scott Hadley, Quan Bai 0001
Expert Syst. Appl.4
2024 Parallel scale de-blur net for sharpening video images for remote clinical assessment of hand movements
abstract
Clinicians and researchers commonly assess hand movements to detect and monitor neurological disorders. With the growing use of deep learning and biomedical informatics, computer vision can be applied to hand movement videos to extract movement features. Such methods promise objective and automated measures of hand movements which can potentially reveal richer details than clinicians in a face-to-face setting. However, extracting valid measures from hand movement video data is a challenging task because motion blur occurs when the hands move quickly. To address this issue, current de-blurring methods have been investigated and a novel ‘Parallel Scale Deblur Net’ (PSDNet) is proposed for hand movement image de-blurring. The results demonstrate that PSDNet achieves better de-blurring performance on both a general blur dataset (available online) and also on our own hand motion dataset.
Renjie Li 0001, Guan Huang 0001, Xinyi Wang 0009, Yanyu Chen 0001, Son N. Tran, Saurabh Kumar Garg 0001, Rebecca J. St George, Katherine Lawler, Jane E. Alty, Quan Bai 0001
Expert Syst. Appl.5
2024 Deep cross-domain transfer for emotion recognition via joint learning
abstract
Abstract Deep learning has been applied to achieve significant progress in emotion recognition from multimedia data. Despite such substantial progress, existing approaches are hindered by insufficient training data, leading to weak generalisation under mismatched conditions. To address these challenges, we propose a learning strategy which jointly transfers emotional knowledge learnt from rich datasets to source-poor datasets. Our method is also able to learn cross-domain features, leading to improved recognition performance. To demonstrate the robustness of the proposed learning strategy, we conducted extensive experiments on several benchmark datasets including eNTERFACE, SAVEE, EMODB, and RAVDESS. Experimental results show that the proposed method surpassed existing transfer learning schemes by a significant margin.
Dung Nguyen 0001, Duc Thanh Nguyen, Sridha Sridharan, Mohamed Almorsy, Simon Denman, Son N. Tran, Clinton Fookes
Multim. Tools Appl.6
2024 Empirical Evaluation of Machine Learning Models for Fuel Consumption, Driver Identification, and Behavior Prediction
abstract
Drivers can be identified through patterns in their routine driving behaviours, as observed by analysing the timing and sequence of various manoeuvres. In contemporary mobility contexts, comprehending and accurately predicting drivers’ behaviours are crucial for informing efficient transportation planning, enhancing traffic safety, reducing emissions, and improving driving efficiency. An increasing number of researchers have explored a variety of machine learning (ML) models to identify, classify, and predict drivers’ behaviours. However, the reliability of these results is often undermined by the complexities associated with the data characteristics, contexts, and the authors’ expertise. Additionally, there is a lack of comprehensive investigation into the effect of driving behaviour on vehicles’ performance, driver identity, and driving activities. This research aims to compare various ML methods to establish a conclusive and generalisable empirical benchmark. The experiments were divided into three phases: estimation of fuel consumption, driver identification, and driver actions’ prediction from drivers’ behaviour during motion. The experiments evaluate prediction accuracy, performance, and computational cost using a different range of temporal and nontemporal ML models and eight datasets from diverse sources, which resulted in 9 tables of outputs. The results have been gauged and scored precisely, and then high-rated and ineffective algorithms were pinpointed for each task. This study is the most in-depth investigation, providing an exhaustive comparison of different ML models for predicting three main criteria of driving behaviour, marking it as the most detailed investigation in this field.
Jamal Maktoubian, Son N. Tran, Anna Shillabeer, Muhammad Bilal Amin, Lawrence Sambrooks, Reza Khoshkangini
IEEE Trans. Intell. Transp. Syst.2
2023 Neurosymbolic Reasoning and Learning with Restricted Boltzmann Machines
abstract
Knowledge representation and reasoning in neural networks has been a long-standing endeavour which has attracted much attention recently. The principled integration of reasoning and learning in neural networks is a main objective of the area of neurosymbolic Artificial Intelligence. In this paper, a neurosymbolic system is introduced that can represent any propositional logic formula. A proof of equivalence is presented showing that energy minimization in restricted Boltzmann machines corresponds to logical reasoning. We demonstrate the application of our approach empirically on logical reasoning and learning from data and knowledge. Experimental results show that reasoning can be performed effectively for a class of logical formulae. Learning from data and knowledge is also evaluated in comparison with learning of logic programs using neural networks. The results show that our approach can improve on state-of-the-art neurosymbolic systems. The theorems and empirical results presented in this paper are expected to reignite the research on the use of neural networks as massively-parallel models for logical reasoning and promote the principled integration of reasoning and learning in deep networks.
Son N. Tran, Artur S. d'Avila Garcez
AAAI1
2023 Multi-output Deep-Supervised Classifier Chains for Plant Pathology
abstract
Plant leaf disease classification is an important task in smart agriculture which plays a critical role in sustainable production. Modern machine learning approaches have shown unprecedented potential in this classification task which offers an array of benefits including time saving and cost reduction. However, most recent approaches directly employ convolutional neural networks where the effect of the relationship between plant species and disease types on prediction performance is not properly studied. In this study, we proposed a new model named Multi-output Deep Supervised Classifier Chains (Mo-DsCC) which weaves the prediction of plant species and disease by chaining the output layers for the two labels. Mo-DsCC consists of three components: A modified VGG-16 network as the backbone, deep supervision training, and a stack of classification chains. To evaluate the advantages of our model, we perform intensive experiments on two benchmark datasets Plant Village and PlantDoc. Comparison to recent approaches, including multi-model, multi-label (Power-set), multi-output and multi-task, demonstrates that Mo-DsCC achieves better accuracy and F1-score. The empirical study in this paper shows that the application of Mo-DsCC could be a useful puzzle for smart agriculture to benefit farms and bring new ideas to industry and academia.
Jianping Yao, Son N. Tran
IJCNN2
2023 Deep Learning for Effective Gender Classification of Tasmania Giant Crabs
abstract
The giant crab fishery in southeast Australia currently suffers from a lack of accurate size and sex data to establish population dynamics essential for the management of the industry. Determining these traits manually on boats or from observers using video would be time-consuming and prone to observer error. This research aims to find an efficient, accurate, and fast way to identify giant crabs' gender to eliminate the manual cost and operation time. This problem can be solved by artificial intelligence technologies, particularly Convolutional Neural Networks (CNN). However, CNNs can detect the crabs but find it challenging to identify their genders from the top view (carapace). Other issues include the lack of training data and the demand for compact system to be deployed on boat easily. With such constraints of effectiveness and efficiency, we address the problem of crab gender classification by proposing a cascading architecture. First, we simplify a light-weight object detection model (MobileNet) for carapace localisation. After that our model extract the carapace area on crab images for classification modelling. According to our experiments' results, with the use of light-weight CNNs for our cascading architecture, we achieved the highest accuracy of up to 96.34% with an efficiency of around 2 frames per second in a Raspberry Pi V4.
Jianping Yao, Son N. Tran, Lianxue Zhang, Jiaxin Ye, Ananda Maiti, Scott Hadley
IJCNN2
2023 Real-time automated detection of older adults' hand gestures in home and clinical settings
Guan Huang 0001, Son N. Tran, Quan Bai 0001, Jane E. Alty
Neural Comput. Appl.2
2022 Wine Characterisation with Spectral Information and Predictive Artificial Intelligence
Jianping Yao, Son N. Tran, Hieu Nguyen 0004, Samantha Sawyer, Rocco Longo
ICONIP (7)2
2022 A survey: From shallow to deep machine learning approaches for blood pressure estimation using biosensors
Sumbal Maqsood, Shuxiang Xu, Son N. Tran, Saurabh Kumar Garg 0001, Matthew Springer, Mohan Karunanithi, Rami Mohawesh
Expert Syst. Appl.3
2022 Deep Auto-Encoders With Sequential Learning for Multimodal Dimensional Emotion Recognition
abstract
Multimodal dimensional emotion recognition has drawn a great attention from the affective computing community and numerous schemes have been extensively investigated, making a significant progress in this area. However, several questions still remain unanswered for most of existing approaches including: (i) how to simultaneously learn compact yet representative features from multimodal data, (ii) how to effectively capture complementary features from multimodal streams, and (iii) how to perform all the tasks in an end-to-end manner. To address these challenges, in this paper, we propose a novel deep neural network architecture consisting of a two-stream auto-encoder and a long short term memory for effectively integrating visual and audio signal streams for emotion recognition. To validate the robustness of our proposed architecture, we carry out extensive experiments on the multimodal emotion in the wild dataset: RECOLA. Experimental results show that the proposed method achieves state-of-the-art recognition performance.
Dung Nguyen 0001, Duc Thanh Nguyen, Thanh Thi Nguyen 0001, Son N. Tran, Thin Nguyen, Sridha Sridharan, Clinton Fookes
IEEE Trans. Multim.5
2021 Compositional Neural Logic Programming
abstract
This paper introduces Compositional Neural Logic Programming (CNLP), a framework that integrates neural networks and logic programming for symbolic and sub-symbolic reasoning. We adopt the idea of compositional neural networks to represent first-order logic predicates and rules. A voting backward-forward chaining algorithm is proposed for inference with both symbolic and sub-symbolic variables in an argument-retrieval style. The framework is highly flexible in that it can be constructed incrementally with new knowledge, and it also supports batch reasoning in certain cases. In the experiments, we demonstrate the advantages of CNLP in discriminative tasks and generative tasks.
Son N. Tran
IJCAI1
2021 Analysis of concept drift in fake reviews detection
Rami Mohawesh, Son N. Tran, Robert Ollington, Shuxiang Xu
Expert Syst. Appl.2
2021 Coconut trees detection and segmentation in aerial imagery using mask region-based convolution neural network
abstract
Abstract Food resources face severe damages under extraordinary situations of catastrophes such as earthquakes, cyclones, and tsunamis. Under such scenarios, speedy assessment of food resources from agricultural land is critical as it supports aid activity in the disaster‐hit areas. In this article, a deep learning approach was presented for the detection and segmentation of coconut trees in aerial imagery provided through the AI competition organised by the World Bank in collaboration with OpenAerialMap and WeRobotics . Masked Region‐based Convolution Neural Network (Mask R‐CNN) approach was used for identification and segmentation of coconut trees. For the segmentation task, Mask R‐CNN model with ResNet50 and ResNet101 based architectures was used. Several experiments with different configuration parameters were performed and the best configuration for the detection of coconut trees with more than 90% confidence factor was reported. For the purpose of evaluation, Microsoft COCO dataset evaluation metric namely mean average precision (mAP) was used.An overall 91% mean average precision for coconut trees’ detection was achieved.
Muhammad Shakaib Iqbal, Hazrat Ali, Son N. Tran, Talha Iqbal
IET Comput. Vis.3
2020 Data Augmentation with Generative Adversarial Networks for Grocery Product Image Recognition
abstract
Image recognition tasks have gained enormous progress with a tremendous amount of training data. However, it isn't easy to obtain such training datasets that contain numerous annotated images in the domain of grocery product recognition. A small number of training data always results in a less than stellar recognition accuracy. Here we attempt to address this challenge by using generative adversarial networks (GAN), which can generate natural images for data augmentation. This paper aims to investigate the feasibility of using GAN to create synthetic training data, and thus to improve grocery product recognition accuracy. In this work, different GAN variants and image rotation are employed to enlarge the fruit datasets. Then, we train the CNN classifier using different data augmentation methods and compare the top-1 accuracy results. Finally, our experiments demonstrate that Auxiliary Classifier GAN (ACGAN) has achieved the best performance, which obtains l.26%~3.44% increase in recognition accuracy. As an additional contribution, the results show that the effectiveness of using generated data is very close to that of using real data, which in our best experimental case, are 93.85% and 94.25%, respectively.
Shuxiang Xu, Son N. Tran, Byeong Ho Kang 0001
ICARCV3
2020 Neuro-Symbolic Probabilistic Argumentation Machines
abstract
Neural-symbolic systems combine the strengths of neural networks and symbolic formalisms. In this paper, we introduce a neural-symbolic system which combines restricted Boltzmann machines and probabilistic semi-abstract argumentation. We propose to train networks on argument labellings explaining the data, so that any sampled data outcome is associated with an argument labelling. Argument labellings are integrated as constraints within restricted Boltzmann machines, so that the neural networks are used to learn probabilistic dependencies amongst argument labels. Given a dataset and an argumentation graph as prior knowledge, for every example/case K in the dataset, we use a so-called K-maxconsistent labelling of the graph, and an explanation of case K refers to a K-maxconsistent labelling of the given argumentation graph. The abilities of the proposed system to predict correct labellings were evaluated and compared with standard machine learning techniques. Experiments revealed that such argumentation Boltzmann machines can outperform other classification models, especially in noisy settings.
Régis Riveret, Son N. Tran, Artur S. d'Avila Garcez
KR2
2020 Mixed-dependency models for multi-resident activity recognition in smart homes
Son N. Tran, Ngo Tung Son, Qing Zhang 0001, Mohan Karunanithi
Multim. Tools Appl.1
2020 Probabilistic approaches for music similarity using restricted Boltzmann machines
Son N. Tran, Ngo Tung Son, Artur S. d'Avila Garcez
Neural Comput. Appl.1
2020 Sequence Classification Restricted Boltzmann Machines With Gated Units
abstract
For the classification of sequential data, dynamic Bayesian networks and recurrent neural networks (RNNs) are the preferred models. While the former can explicitly model the temporal dependences between the variables, and the latter have the capability of learning representations. The recurrent temporal restricted Boltzmann machine (RTRBM) is a model that combines these two features. However, learning and inference in RTRBMs can be difficult because of the exponential nature of its gradient computations when maximizing log likelihoods. In this article, first, we address this intractability by optimizing a conditional rather than a joint probability distribution when performing sequence classification. This results in the "sequence classification restricted Boltzmann machine" (SCRBM). Second, we introduce gated SCRBMs (gSCRBMs), which use an information processing gate, as an integration of SCRBMs with long short-term memory (LSTM) models. In the experiments reported in this article, we evaluate the proposed models on optical character recognition, chunking, and multiresident activity recognition in smart homes. The experimental results show that gSCRBMs achieve the performance comparable to that of the state of the art in all three tasks. gSCRBMs require far fewer parameters in comparison with other recurrent networks with memory gates, in particular, LSTMs and gated recurrent units (GRUs).
Son N. Tran, Artur S. d'Avila Garcez, Tillman Weyde, Jie Yin 0001, Qing Zhang 0001, Mohan Karunanithi
IEEE Trans. Neural Networks Learn. Syst.1
2019 dpUGC: Learn Differentially Private Representation for User Generated Contents (Best Paper Award, Third Place, Shared)
Xuan-Son Vu, Son N. Tran, Lili Jiang 0002
CICLing (1)2
2018 Sentimental Analysis for AIML-Based E-Health Conversational Agents
David Ireland, Hamed Hassanzadeh, Son N. Tran
ICONIP (2)3
2018 Improving Recurrent Neural Networks with Predictive Propagation for Sequence Labelling
Son N. Tran, Qing Zhang 0001, Anthony N. Nguyen, Xuan-Son Vu, Ngo Tung Son
ICONIP (1)1
2018 Human Identification via Unsupervised Feature Learning from UWB Radar Data
Jie Yin 0001, Son N. Tran, Qing Zhang 0001
PAKDD (1)2