Thien Huynh-The

dblp:153/6644 · also Huynh-The Thien · DBLP profile ↗
← Back
50ranked-venue papers
22as first author
29since 2021 · last 2026
0000-0002-9172-2935ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 22 · 7 first-author · 18 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 3 first-author · 5 since 2021Databases, data management, data science and information retrieval · 6 · 4 first-author · 1 since 2021Artificial intelligence and machine learning · 5 · 3 first-author · 3 since 2021Systems, architecture and hardware · 4 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 first-authorHuman-computer interaction and ubiquitous computing · 4 · 2 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Federated learning for big data: A survey on opportunities, applications, and future directions
G. Thippa Reddy, Quoc-Viet Pham, Thien Huynh-The, Hailin Feng, Kai Fang 0001, Sharnil Pandya, Madhusanka Liyanage, Wei Wang 0077, Thanh Thi Nguyen 0001
Eng. Appl. Artif. Intell.3
2026 Temporal-Frequency-Aware Deep Networks for Efficient Waveform Classification in Integrated Radar-Communication (IRC) Systems
abstract
With the growing demand for efficient spectrum utilization in integrated radar-communication (IRC) systems, driven by Internet-of-Things (IoT) and fifth-generation (5G) advancements, robust waveform classification techniques have become increasingly critical. This paper introduces TFINet, a cutting-edge deep learning (DL) architecture designed for waveform classification in spectrally congested environments. TFINet leverages time-frequency representations (TFRs) derived from the Smoothing Pseudo-Wigner-Ville Distribution (SPWVD) to improve feature quality and mitigate cross-term interference, enhancing classification accuracy. The network incorporates two key modules: the Dual-Temporal Frequency Extraction (DTFE) and Time-Frequency Selective Downsampling (TFSD). The DTFE module improves feature extraction by decoupling time and frequency features through dual-branch processing, while the TFSD module intelligently reduces dimensionality, preserving essential features without compromising performance. These innovations enable TFINet to balance computational efficiency and classification accuracy, enhancing its suitability for resource-constrained edge devices. On a diverse synthetic dataset of 12 waveform types, TFINet achieves 91.38% overall classification accuracy with 59K parameters and 0.328 ms inference time. Compared to existing deep models, TFINet demonstrates superior performance in both accuracy and efficiency, validating its suitability for practical IRC systems.
Thien Huynh-The, Thanh-Dat Tran, Nguyen Cong Luong 0001, Dusit Niyato
IEEE Internet Things J.1
2026 Integration of TinyML and LargeML: A Survey of 6G and Beyond
abstract
The evolution from fifth-generation (5G) to sixth-generation (6G) networks is driving an unprecedented demand for advanced machine learning (ML) solutions. Deep learning has already demonstrated significant impact across mobile networking and communication systems, enabling intelligent services such as smart healthcare, smart grids, autonomous vehicles, aerial platforms, digital twins, and the metaverse. At the same time, the rapid proliferation of resource-constrained Internet-of-Things (IoT) devices has accelerated the adoption of tiny machine learning (TinyML) for efficient on-device intelligence, while large machine learning (LargeML) models continue to require substantial computational resources to support large-scale IoT services and ML-generated content. These trends highlight the need for a unified framework that integrates TinyML and LargeML to achieve seamless connectivity, scalable intelligence, and efficient resource management in future 6G systems. This survey provides a comprehensive review of recent advances enabling the integration of TinyML and LargeML in next-generation wireless networks. In particular, we(i)provide an overview of TinyML and LargeML,(ii)analyze the motivations and requirements for unifying these paradigms within the 6G context,(iii)examine efficient bidirectional integration approaches,(iv)review state-of-the-art solutions and their applicability to emerging 6G services, and(v)identify key challenges related to performance optimization, deployment feasibility, resource orchestration, and security. Finally, we outline promising research directions to guide the holistic integration of TinyML and LargeML for intelligent, scalable, and energy-efficient 6G networks and beyond.
Thai-Hoc Vu, Ngo Hoang Tu, Thien Huynh-The, Miroslav Voznak, Kyungchun Lee, Sunghwan Kim 0001, Quoc-Viet Pham
IEEE Internet Things J.3
2026 Analysis and Optimization Framework for STAR-RIS-Aided Short-Packet Systems With Covert Rate-Splitting Signaling
abstract
In this paper, we investigate the performance of simultaneous transmitting and reflecting reconfigurable intelligent surface (STAR-RIS)-enabled short-packet communication (SPC) systems that employ covert rate-splitting (RS) to facilitate applications of Internet-of-Things (IoT). Under the generalized model of the α-η-κ-μ fading, we analyze covertness by first deriving a closed-form expression for the warden’s detection error probability (DEP) and then deducing the optimal detection threshold at which the DEP is minimized. From this optimal DEP, we guide the choice of the feasible power-allocation (PA) region at which the warden’s produced DEP is always beyond the minimal acceptable covertness. On the other hand, we develop mathematical frameworks for evaluating the system’s block-error rate (BLER) and the ergodic rate (ER), as well as providing guidelines on how to access their performance limits at high signal-to-noise ratio, especially the diversity orders and ergodic slopes of the users. Furthermore, we also propose to enhance the system performance by jointly optimizing the PA and RS coefficients in order to: (1) minimize the maximum BLER across private users subject to a minimum covertness requirement and (2) maximize the minimum ER across users subject to covertness and decoding constraints. Numerical results confirm the analytical expressions and show that the proposed optimization efficiently tunes the PA and RS coefficients to achieve the BLER and ER fairness objectives.
Anh-Tu Le, Thai-Hoc Vu, Thien Huynh-The, Miroslav Voznak
IEEE Trans. Commun.3
2026 Transformer Model Embedding Dual Stream for Modulation Classification of Short Signal Samples
abstract
Automatic modulation classification (AMC) is a critical task in modern communication systems, particularly under diverse signal conditions and limited data scenarios. Existing transformer-based AMC models often rely on single-stream architectures and uniform input formats, which limit their effectiveness in capturing rich signal features. To address these limitations, we propose DTNet, a novel transformer-based dual-stream network designed for efficient and accurate modulation classification. DTNet introduces two key innovations: (1) a scale feature and extension (SFE) block that applies a scaling function to transform signals into a structured output map, followed by an extension module that reconstructs the signal into a square matrix and integrates a response map using adaptive filters; and (2) a convolutional stream that extracts discriminative features from multi-scale signal representations. Furthermore, a modified feature embedding to leverage the transformer architecture is introduced to capture global dependencies and contextual information from the input signal, thereby enhancing the modulation classification accuracy. Experimental results show that DTNet achieves superior performance on benchmark datasets, reaching classification accuracies of 93.4% on RML2016.10A and 94.4% on RML2016.10B, outperforming state-of-the-art deep learning methods while maintaining lower computational complexity. The source code is available at https://github.com/daothanh2011/DTNet .
Thien-Thanh Dao, Quoc-Viet Pham, Thien Huynh-The, Won-Joo Hwang
ACM Trans. Intell. Syst. Technol.3
2025 RIS-assisted LoRa networks with diversity: Impact of hardware impairments and phase noise
Thi-Phuong-Anh Hoang, Thien Huynh-The, Tien Hoa Nguyen 0001, Trong Thua Huynh, Nguyen-Son Vo, Tu Lam Thanh
Comput. Commun.2
2025 Latency Minimization for STAR-RIS-Aided Federated Learning Networks With Wireless Power Transfer
abstract
Simultaneous transmitting and reflecting reconfigurable intelligent surfaces (STAR-RISs) introduces revolutionary capabilities by reaching full space coverage for wireless signals, significantly enhancing the efficiency and reliability of Internet of Things (IoT) networks compared to traditional RIS. In this article, we propose a novel framework that leverages STAR-RIS into wirelessly powered federated learning (FL) networks with a multiantenna access point, aiming to minimize system latency. A multivariable nonconvex optimization problem is formulated to optimize phase shift vectors of STAR-RIS, beamforming matrices, time, power, and computation frequency for each user in all phases of FL. Block coordinate descent (BCD) over the combination of an 1-D search algorithm and interior point method is employed to optimize time, power, computation frequency, phase shift vectors of STAR-RIS, and active beamforming matrix in the uplink transmission phase, while semi-definite relaxation via BCD addresses phase shift vectors of STAR-RIS and beamforming matrices optimization in harvesting and downlink transmission phases. On this basis, the optimized downlink transmission time and power are derived. The convergence of the proposed algorithm and the superiority of its performance compared to benchmark schemes are validated through comprehensive simulations. Our findings indicate the potential of FL, multiantenna aggregation server, and STAR-RIS in ushering in a new era of intelligent and efficient IoT networks.
Mohammad Hossein Alishahi, Paul Fortier, Ming Zeng 0002, Thien Huynh-The, Xingwang Li 0001, Quoc-Viet Pham
IEEE Internet Things J.4
2025 XSNet: Lightweight Object Detection Model Using X-Shaped Architecture in Remote Sensing Images
abstract
Remote sensing object detection faces challenges such as small object sizes, complex backgrounds, and computational constraints. To overcome these challenges, we propose XSNet, an efficient deep learning model proficiently designed to enhance feature representation and multi-scale detection. Concretely, XSNet introduces three key innovations: SIner (Swin-Involution Transformer) to improve local self-attention and spatial adaptability; PosWeightRA (Positional Weight Bi-Level Routing Attention) to refine spatial awareness and preserves positional encoding; and an X-shaped multi-scale feature fusion strategy to optimize feature aggregation while reducing computational cost. These components collectively improve detection accuracy, particularly for small and overlapping objects. Through extensive experiments, XSNet achieves impressive mAP0.5and mAP0.95scores of 47.1% and 28.2% on VisDrone2019, and 92.9% and 66.0% on RSOD, respectively. It outperforms state-of-the-art models while maintaining a compact size of 7.11 million parameters and fast inference time of 35.5 milliseconds, making it well-suited for real-time remote sensing in resource-constrained environments.
Dat Minh-Tien Nguyen, Thien Huynh-The
IEEE Geosci. Remote. Sens. Lett.2
2024 WaveNet: Toward Waveform Classification in Integrated Radar-Communication Systems With Improved Accuracy and Reduced Complexity
abstract
The integration of radar and communication systems in 6G networks has led to a significant challenge of spectrum congestion. To address this issue, we propose a deep learning-based method for efficient waveform-based signal classification. Our method is designed to handle large and impaired radar and communication signals, and is crucial for the implementation of resource-limited cognitive radio-enabled Internet-of-Things (CR-IoT) devices. We introduce WaveNet, a cost-efficient deep convolutional neural network that can aptly learn underlying radio features from time-frequency images transformed by a smooth pseudo Wigner-Ville distribution. WaveNet incorporates several innovative modules, including cost-efficient feature awareness, which integrates two well-designed structural blocks: grouped-of-kernel-wise residual connections and dual asymmetric channel attention. These enhancements significantly reduce network size without compromising classification accuracy. Based on various simulations experimented on an impaired signal dataset containing eight radar and communication waveform types, the results demonstrate the effectiveness and robustness of WaveNet, achieving an overall classification accuracy of 92.02%. Compared to the current state-of-the-art deep models, WaveNet has the lowest architectural complexity, with a network size five times smaller, while still outperforming them by approximately 0.5 – 1.69%. Consequently, WaveNet emerges as a valuable solution for waveform classification in integrated radar-communication 6G systems.
Thien Huynh-The, Van-Phuc Hoang, Jae-Woo Kim, Minh-Thanh Le, Ming Zeng 0002
IEEE Internet Things J.1
2024 Enhancing Aerial Semantic Segmentation With Feature Aggregation Network for DeepLabV3+
abstract
As a cutting-edge deep encoder-decoder architecture, DeepLabV3+ has been realized as a cutting-edge solution for image segmentation, especially aerial semantic segmentation in remote sensing applications. This is because of an atrous spatial pyramid pooling (ASPP) block deployed in its encoder with multiple atrous convolutional layers to enrich diversified feature extraction and learning efficiency. However, the DeepLabV3+ encoder-decoder architecture has some limitations, including the lack of information during the upsampling process and some inappropriate customizations that cause incorrect segmentation. To address these shortcomings, we introduce an efficient architecture with a novel feature aggregation network (FAN), which facilitates the extraction of features across multiple scales and stages. Concurrently, we apply some adaptive upgrades to the ASPP block, involving a new set of dilation factors that are adept at accommodating low-resolution inputs. Through simulations, we demonstrate the effectiveness and generalizability of our improved model by evaluating it with different backbones and dilation rates. In addition, compared with recent deep segmentation models, our improved model is superior in terms of mean$F1$score (mF1) by at least 1.32%–6.34% and 0.98%–5.61% on the UAVid and Vaihingen datasets, respectively.
Gia-Vuong Nguyen, Thien Huynh-The
IEEE Geosci. Remote. Sens. Lett.2
2023 Artificial intelligence for the metaverse: A survey
Thien Huynh-The, Quoc-Viet Pham, Xuan-Qui Pham, Thanh Thi Nguyen 0001, Zhu Han 0001, Dong-Seong Kim 0002
Eng. Appl. Artif. Intell.1
2023 Blockchain for the metaverse: A Review
abstract
Since Facebook officially changed its name to Meta in Oct. 2021, the metaverse has become a new norm of social networks and three-dimensional (3D) virtual worlds. The metaverse aims to bring 3D immersive and personalized experiences to users by leveraging many pertinent technologies. Despite great attention and benefits, a natural question in the metaverse is how to secure its users' digital content and data. In this regard, blockchain is a promising solution owing to its distinct features of decentralization, immutability, and transparency. To better understand the role of blockchain in the metaverse, we aim to provide an extensive survey on the applications of blockchain for the metaverse. We first present a preliminary to blockchain and the metaverse and highlight the motivations behind the use of blockchain for the metaverse. Next, we extensively discuss blockchain-based methods for the metaverse from technical perspectives, such as data acquisition, data storage, data sharing, data interoperability, and data privacy preservation. For each perspective, we first discuss the technical challenges of the metaverse and then highlight how blockchain can help. Moreover, we investigate the impact of blockchain on key-enabling technologies in the metaverse, including Internet-of-Things, digital twins, multi-sensory and immersive applications, artificial intelligence, and big data. We also present some major projects to showcase the role of blockchain in metaverse applications and services. Finally, we present some promising directions to drive further research innovations and developments toward the use of blockchain in the metaverse in the future.
Thien Huynh-The, G. Thippa Reddy, Weizheng Wang 0001, Gokul Yenduri, Pasika Ranaweera, Quoc-Viet Pham, Daniel B. da Costa 0001, Madhusanka Liyanage
Future Gener. Comput. Syst.1
2023 Federated Learning for the Healthcare Metaverse: Concepts, Applications, Challenges, and Future Directions
abstract
Recent technological advancements have considerably improved healthcare systems to provide various intelligent services, improving life quality. The Metaverse, often described as the next evolution of the Internet, helps the users interact with each other and the environment, thus offering a seamless connection between the virtual and physical worlds. Additionally, the Metaverse, by integrating emerging technologies, such as artificial intelligence (AI), cloud edge computing, Internet of Things (IoT), blockchain, and semantic communications, can potentially transform many vertical domains in general and the healthcare sector (healthcare Metaverse) in particular. The healthcare Metaverse holds huge potential to revolutionize the development of intelligent healthcare systems, thus presenting new opportunities for significant advancements in healthcare delivery, personalized healthcare experiences, medical education, collaborative research, and so on. However, various challenges are associated with the realization of the healthcare Metaverse, such as privacy, interoperability, data management, and security. Federated learning (FL), a new branch of AI, opens up enormous opportunities to deal with the aforementioned challenges in the healthcare Metaverse by exploiting the data and computing resources available at the distributed devices. This motivated us to present a survey on adopting FL for the healthcare Metaverse. Initially, we present the preliminaries of IoT-based healthcare systems, FL in conventional healthcare, and the healthcare Metaverse. Furthermore, the benefits of the FL in the healthcare Metaverse are discussed. Subsequently, we discuss the several applications of FL-enabled healthcare Metaverse, including medical diagnosis, patient monitoring, medical education, infectious disease, and drug discovery. Finally, we highlight the significant challenges and potential solutions toward realizing FL in the healthcare Metaverse.
Ali Kashif Bashir, Nancy Victor, Sweta Bhattacharya, Thien Huynh-The, Rajeswari Chengoden, Gokul Yenduri, Praveen Kumar Reddy Maddikunta, Quoc-Viet Pham, G. Thippa Reddy, Madhusanka Liyanage
IEEE Internet Things J.4
2023 Energy-Efficient Multiprocessor-Based Computation and Communication Resource Allocation in Two-Tier Federated Learning Networks
abstract
In conventional federated learning (FL), multiple edge devices holding local data jointly train a machine learning model by communicating learning updates with a centralized aggregator without exchanging their data samples. Owing to the communication and computation bottleneck at the centralized aggregator and inaccurate learning model caused by the non-independent and identically distributed (IID) data, we here consider a two-tier FL network, in which Internet of Things (IoT) nodes are the core clients that hold data, the model aggregators at the middle tier are the low altitude aerial platforms (UAVs), and the model aggregator at the top-most layer is the high-altitude aerial platform (UAV with relatively high altitude). Under the assumption that each IoT node has parallel computing ability, we study the energy-efficient computation and communication resource allocation in such a network within some time budget. Upon formulating the problem as an optimization problem, we solve the computation and communication resource allocation problems as the separate subproblems within a time frame, and then propose an iterative algorithm to solve the entire problem jointly. More specifically, we solve both the energy-efficient computation and communication resource allocation subproblems using the dual decomposition technique, and then apply a bisection search-based recursive technique to solve the entire energy efficiency problem jointly. Moreover, we propose offline and online client scheduling schemes that not only select the optimal edge nodes for association but also assign workload to each client based on the data quality and workload constraint. With real data, extensive simulations are conducted to verify the effectiveness of the proposed resource allocation scheme. The results further reveal that the learning performance not only is dependent on the computation and communication energy consumption of the FL process but also the model divergence weight owing to the non-IID data at client IoT nodes.
Rukhsana Ruby, Hailiang Yang, Felipe A. P. de Figueiredo, Thien Huynh-The, Kaishun Wu
IEEE Internet Things J.4
2023 Short-Packet Communications in Multihop Networks With WET: Performance Analysis and Deep Learning-Aided Optimization
abstract
In this paper, we study short-packet communications in multi-hop networks with wireless energy transfer, where relay nodes harvest energy from power beacons to transmit short packets to multiple destinations. It is proposed a novel cooperative beamforming relay selection (CRS) scheme which incorporates partial relay selection and distributed multiuser beamforming to achieve a high-reliable transmission in two consecutive hops. A closed-form expression for the average block error rate (BLER) of the CRS scheme is derived, based on which an asymptotic analysis is also carried out. To achieve optimal channel uses allocation, we formulate a fairness end-to-end throughput maximization problem which is generally NP-hard due to the non-concavity of the objective function and mixed-integer constraints. To solve this challenging problem efficiently, we first relax channel uses to be continuous and transform the relaxed problem into an equivalent non-convex one, but with a more tractable form. We then develop a low-complexity iterative algorithm relying on inner approximation framework to convexify non-convex parts that converges to at least a locally optimal solution. Towards real-time settings, we design an efficient deep convolutional neural network (CNN) with multiscale-accumulation connections to achieve the sub-optimal solution of the relaxed problem via real-time inference processes. Numerical results are presented to verify the analytical derivations and to demonstrate performance improvements of the CRS scheme over the benchmark ones in terms of BLER, reliability, latency, and throughput in various settings. Moreover, the designed CNN provides the lowest root-mean-square error compared to the state-of-the-art deep learning approaches while the CNN-aided optimization framework estimates accurately the optimal channel uses allocation with low execution time.
Van-Dinh Nguyen, Daniel B. da Costa 0001, Thien Huynh-The, Rose Qingyang Hu, Beongku An
IEEE Trans. Wirel. Commun.4
2022 An Efficient Deep CNN Design for EH Short-Packet Communications in Multihop Cognitive IoT Networks
abstract
In this paper, we design an efficient deep convolutional neural network (CNN) to improve and predict the performance of energy harvesting (EH) short-packet communications in multi-hop cognitive Internet-of-Things (IoT) networks. Specifically, we propose a Sum-EH scheme that allows IoT nodes to harvest energy from either a power beacon or primary transmitters to improve not only packet transmissions but also energy harvesting capabilities. We then build a novel deep CNN framework with feature enhancement-collection blocks based on the proposed Sum-EH scheme to simultaneously estimate the block error rate (BLER) and throughput with high accuracy and low execution time. Simulation results show that the proposed CNN framework achieves almost exactly the BLER and throughput of Sum-EH one, while it considerably reduces computational complexity, suggesting a real-time setting for IoT systems under complex scenarios. Moreover, the designed CNN model achieves the root-mean-square-error (RMSE) of 1.33 × 10-2on the considered dataset, which exhibits the lowest RMSE compared to the deep neural network and state-of-the-art machine learning approaches.
Thien Huynh-The, Van-Dinh Nguyen, Daniel B. da Costa 0001, Rose Qingyang Hu, Beongku An
ICC2
2022 UAV-enabled Wireless Powered Communication for Energy-Efficient Federated Learning
abstract
Federated learning (FL) has found numerous applications in wireless and mobile networks thanks to its distinctive features. However, efficient FL networks require to address the energy limitation of FL users. Exploited the flexible deployment and agile mobility of unmanned aerial vehicles (UAVs), this work proposes to dispatch the UAV with edge computing capabilities as an aerial energy source to wirelessly power FL users and as an aerial server for model aggregation. In this regard, we investigate a resource allocation problem that minimizes the energy consumption of FL users and the aerial server. To resolve the nonconvexity of the formulated problem, we propose to decompose the entire set of variables into three blocks, and then develop an iterative algorithm. Simulations results are presented to show that our proposed algorithm significantly outperforms several benchmarks.
Quoc-Viet Pham, Mai Le, Thien Huynh-The, Zhu Han 0001, Won-Joo Hwang
ICC3
2022 2D Skeleton-based Action Recognition Using Action-Snippets and Sequential Deep Learning
abstract
Human action recognition (HAR) is an active and crucial field of computer vision due to its various applications, such as smart surveillance and human-computer interaction. Recently, the human skeleton, which is compact and intuitive for representing actions and body movements, has been widely used in numerous HAR frameworks. Despite the great success of the skeleton-based HAR, several challenges remain, such as intra-class variability and inter-class similarity. In this paper, we address this task by first proposing a discriminative representation of the action-snippet (i.e., the very short sequence) that captures meaningful characteristics of human pose and body transition. We then employ adequate deep sequential neural networks (DSNNs) to thoroughly learn the temporal relation of action-snippets in a whole sequence. In experiments, the results show that the proposed approach achieves high recognition rates on benchmark datasets while maintaining good computational efficiency (i.e., lightweight networks and high recognition speed).
Aizada Askar, Min-Ho Lee, Thien Huynh-The, Nguyen Anh Tu
SMC3
2022 Automatic Modulation Classification with Low-Cost Attention Network for Impaired OFDM Signals
abstract
In this paper, we propose a deep learning (DL)-based method to automatically identify the modulations of orthogonal frequency-division multiplexing (OFDM) signals in wireless communication systems. In particular, a cost-efficient OFDM modulation classification convolutional neural network (COM-ConvNet) is principally designed with grouped convolutional layers to reduce computing complexity significantly. Remarkably, reconstructing the high-dimensional data array of OFDM signals allows our deep network to learn the underlying sample correlations within every symbol and among different symbols sufficiently. We leverage residual connection and attention connection with element-wise addition and element-wise multiplication layers in specific-designed processing blocks to enhance the pattern learning efficiency. For performance evaluation, we test the proposed method on a synthetic six-modulation OFDM signal dataset under impaired channel conditions and conduct diverse simulations, such as ablation study, parameter investigation, and complexity analysis. COM-ConvNet achieves cost efficiency (i.e., small network size and low computational cost) while maintaining an acceptable accuracy when compared with other DL models.
Thien Huynh-The, Quoc-Viet Pham, Daniel B. da Costa 0001, Dong-Seong Kim 0002
WCNC1
2022 Fusion of Federated Learning and Industrial Internet of Things: A survey
M. Parimala Boobalan, R. M. Swarna Priya, Quoc-Viet Pham, Kapal Dev, Sharnil Pandya, Praveen Kumar Reddy Maddikunta, G. Thippa Reddy, Thien Huynh-The
Comput. Networks8
2022 Deep learning for deepfakes creation and detection: A survey
Thanh Thi Nguyen 0001, Nguyen Quoc Viet Hung, Duc Thanh Nguyen, Thien Huynh-The, Saeid Nahavandi, Thanh Tam Nguyen, Quoc-Viet Pham, Cuong M. Nguyen
Comput. Vis. Image Underst.5
2022 Incentive techniques for the Internet of Things: A survey
Praveen Kumar Reddy Maddikunta, Quoc-Viet Pham, Dinh C. Nguyen, Thien Huynh-The, Ons Aouedi, Gokul Yenduri, Sweta Bhattacharya, G. Thippa Reddy
J. Netw. Comput. Appl.4
2022 Underwater Acoustic Target Classification Based on Dense Convolutional Neural Network
abstract
In oceanic remote sensing operations, underwater acoustic target recognition is always a difficult and extremely important task of sonar systems, especially in the condition of complex sound wave propagation characteristics. The expensively learning recognition model for big data analysis is typically an obstacle for most traditional machine learning (ML) algorithms, whereas the convolutional neural network (CNN), a type of deep neural network, can automatically extract features for accurate classification. In this study, we propose an approach using a dense CNN model for underwater target recognition. The network architecture is designed to cleverly reuse all former feature maps to optimize classification rates under various impaired conditions while satisfying low computational cost. In addition, instead of using time–frequency spectrogram images, the proposed scheme allows directly utilizing the original audio signal in the time domain as the network input data. Based on the experimental results evaluated on the real-world data set of passive sonar, our classification model achieves the overall accuracy of 98.85% at 0-dB signal-to-noise ratio (SNR) and outperforms traditional ML techniques, as well as other state-of-the-art CNN models.
Van-Sang Doan, Thien Huynh-The, Dong-Seong Kim 0002
IEEE Geosci. Remote. Sens. Lett.2
2021 Densely-Accumulated Convolutional Network for Accurate LPI Radar Waveform Recognition
abstract
This paper presents a deep learning-based method to automatically recognize low probability of intercept (LPI) radar waveforms against diversified jamming attacks. Concretely, an efficient convolutional neural network (CNN) architecture, namely Densely-Accumulated Network (DANet), is introduced to learn the time-frequency representation transformed by the Wigner-Ville distribution. Such an architecture has several novel densely-accumulated connection modules specified by various symmetric and asymmetric convolutional layers to enrich diversified features at multiple representational maps. Besides, the skip-connection and dense-connection are leveraged to improve feature learning efficiency and prevent the vanishing gradient when the network goes deeper. Some image processing techniques (e.g., global thresholding and digital filtering) are adopted to enhance the quality of time-frequency image. Relying on simulations, we benchmark the proposed method on a synthetic 13-waveform dataset and also investigate the influence of hyper-parameters (such as image size, number of modules, training data size) on the overall recognition performance. Remarkably, with average accuracy of 98.2% at 0 dB signal-to-noise ratio (SNR), DANet outperforms several backbone CNNs and state-of-the-art networks of LPI waveform recognition while keeping a cost-efficient model.
Thien Huynh-The, Quoc-Viet Pham, Van-Sang Doan, Nhan Thanh Nguyen 0001, Daniel B. da Costa 0001, Dong-Seong Kim 0002
GLOBECOM1
2021 A Deep CNN-based Relay Selection in EH Full-Duplex IoT Networks with Short-Packet Communications
abstract
In this paper, we propose an efficient deep convolutional neural network-based relay selection (CNS) scheme to evaluate and improve the end-to-end throughput in energy harvesting full-duplex Internet-of-Things (IoT) networks. In this system, multiple full-duplex relays harvest energy from a power beacon to assist data transmission from a source node to multiple users under short packet communications. We propose a best relay best user (bR-bU) selection scheme to improve the diversity packet transmission. We then develop a deep convolutional neural network framework for relay selection and throughput prediction with high accuracy and low execution time. Simulation results show that the proposed CNS scheme achieves almost exactly the throughput of bR-bU one, while it considerably reduces computational complexity, suggesting a real-time configuration for IoT systems under complex scenarios. Moreover, the designed CNN model achieves the root-mean-square-error (RMSE) of 8.4 × 10−3on the considered dataset, which exhibits the lowest RMSE as compared to the deep neural network and state-of-the-art machine learning approaches.
Thien Huynh-The, Beongku An
ICC2
2021 Physical Activity Recognition With Statistical-Deep Fusion Model Using Multiple Sensory Data for Smart Health
abstract
Nowadays, enhancing the living standard with smart healthcare via the Internet of Things is one of the most critical goals of smart cities, in which artificial intelligence plays as the core technology. Many smart services, deployed according to wearable sensor-based physical activity recognition, have been able to early detect unhealthy daily behaviors and further medical risks. Numerous approaches have studied shallow handcrafted features coupled with traditional machine learning (ML) techniques, which find it difficult to model real-world activities. In this work, by revealing deep features from deep convolutional neural networks (DCNNs) in fusion with conventional handcrafted features, we learn an intermediate fusion framework of human activity recognition (HAR). According to transforming the raw signal value to pixel intensity value, segmentation data acquired from a multisensor system are encoded to an activity image for deep model learning. Formulated by several novel residual triple convolutional blocks, the proposed DCNN allows extracting multiscale spatiotemporal signal-level and sensor-level correlations simultaneously from the activity image. In the fusion model, the hybrid feature merged from the handcrafted and deep features is learned by a multiclass support vector machine (SVM) classifier. Based on several experiments of performance evaluation, our fusion approach for activity recognition has achieved the accuracy over 96.0% on three public benchmark data sets, including Daily and Sport Activities, Daily Life Activities, and RealWorld. Furthermore, the method outperforms several state-of-the-art HAR approaches and demonstrates the superiority of the proposed intermediate fusion model in multisensor systems.
Thien Huynh-The, Cam-Hao Hua, Nguyen Anh Tu, Dong-Seong Kim 0002
IEEE Internet Things J.1
2021 A Deep-Neural-Network-Based Relay Selection Scheme in Wireless-Powered Cognitive IoT Networks
abstract
In this article, we propose an efficient deep-neural-network-based relay selection (DNS) scheme to evaluate and improve the end-to-end throughput in wireless-powered cognitive Internet-of-Things (IoT) networks. In this system, multiple energy harvesting (EH) relays are deployed randomly to assist data transmission from a source node to multiple users under practical nonlinearity of the EH circuits. We first design an incremental relaying protocol, where a selected user will request the help from relays if the direct transmission is not favorable. In such a protocol, we develop a deep neural network framework for relay selection and throughput prediction with high accuracy, less channel feedback amount, and short execution time. Simulation results show that the proposed DNS scheme achieves higher throughput than the conventional relay selection methods, while it considerably reduces computational complexity, suggesting a real-time configuration for IoT systems under complex scenarios. Moreover, the proposed DNS scheme achieves the root-mean-square error (RMSE) of 6.6×10-3on the considered dataset, which exhibits the lowest RMSE as compared to the state-of-the-art machine learning approaches.
Thong Nhat Tran, Kyusung Shim, Thien Huynh-The, Beongku An
IEEE Internet Things J.4
2021 Convolutional Network With Twofold Feature Augmentation for Diabetic Retinopathy Recognition From Multi-Modal Images
abstract
OBJECTIVE: With the scenario of limited labeled dataset, this paper introduces a deep learning-based approach that leverages Diabetic Retinopathy (DR) severity recognition performance using fundus images combined with wide-field swept-source optical coherence tomography angiography (SS-OCTA). METHODS: The proposed architecture comprises a backbone convolutional network associated with a Twofold Feature Augmentation mechanism, namely TFA-Net. The former includes multiple convolution blocks extracting representational features at various scales. The latter is constructed in a two-stage manner, i.e., the utilization of weight-sharing convolution kernels and the deployment of a Reverse Cross-Attention (RCA) stream. RESULTS: The proposed model achieves a Quadratic Weighted Kappa rate of 90.2% on the small-sized internal KHUMC dataset. The robustness of the RCA stream is also evaluated by the single-modal Messidor dataset, of which the obtained mean Accuracy (94.8%) and Area Under Receiver Operating Characteristic (99.4%) outperform those of the state-of-the-arts significantly. CONCLUSION: Utilizing a network strongly regularized at feature space to learn the amalgamation of different modalities is of proven effectiveness. Thanks to the widespread availability of multi-modal retinal imaging for each diabetes patient nowadays, such approach can reduce the heavy reliance on large quantity of labeled visual data. SIGNIFICANCE: Our TFA-Net is able to coordinate hybrid information of fundus photos and wide-field SS-OCTA for exhaustively exploiting DR-oriented biomarkers. Moreover, the embedded feature-wise augmentation scheme can enrich generalization ability efficiently despite learning from small-scale labeled data.
Cam-Hao Hua, Kiyoung Kim, Thien Huynh-The, Jong In You, Seung-Young Yu, Thuong Le-Tien, Sung-Ho Bae, Sungyoung Lee 0001
IEEE J. Biomed. Health Informatics3
2021 Toward efficient and intelligent video analytics with visual privacy protection for large-scale surveillance
Nguyen Anh Tu, Thien Huynh-The, Kok-Seng Wong, M. Fatih Demirci, Young-Koo Lee
J. Supercomput.2
2020 Learning Constellation Map with Deep CNN for Accurate Modulation Recognition
abstract
Modulation classification, recognized as the intermediate step between signal detection and demodulation, is widely deployed in several modern wireless communication systems. Although many approaches have been studied in the last decades for identifying the modulation format of an incoming signal, they often reveal the obstacle of learning radio characteristics for most traditional machine learning algorithms. To overcome this drawback, we propose an accurate modulation classification method by exploiting deep learning for being compatible with constellation diagram. Particularly, a convolutional neural network is developed for proficiently learning the most relevant radio characteristics of gray-scale constellation image. The deep network is specified by multiple processing blocks, where several grouped and asymmetric convolutional layers in each block are organized by a flow-in-flow structure for feature enrichment. These blocks are connected via skip-connection to prevent the vanishing gradient problem while effectively preserving the information identity throughout the network. Regarding several intensive simulations on the constellation image dataset of eight digital modulations, the proposed deep network achieves the remarkable classification accuracy of approximately 87% at 0 dB signal-to-noise ratio (SNR) under a multipath Rayleigh fading channel and further outperforms some state-of-the-art deep models of constellation-based modulation classification.
Van-Sang Doan, Thien Huynh-The, Cam-Hao Hua, Quoc-Viet Pham, Dong-Seong Kim 0002
GLOBECOM2
2020 Chain-Net: Learning Deep Model for Modulation Classification Under Synthetic Channel Impairment
abstract
Modulation classification, an intermediate process between signal detection and demodulation in a physical layer, is now attracting more interest to the cognitive radio field, wherein the performance is powered by artificial intelligence algorithms. However, most existing conventional approaches pose the obstacle of effectively learning weakly discriminative modulation patterns. This paper proposes a robust modulation classification method by taking advantage of deep learning to capture the meaningful information of modulation signal at multi-scale feature representations. To this end, a novel architecture of convolutional neural network, namely Chain-Net, is developed with various asymmetric kernels organized in two processing flows and associated via depth-wise concatenation and element-wise addition for optimizing feature utilization. The network is evaluated on a big dataset of 14 challenging modulation formats, including analog and high-order digital techniques. The simulation results demonstrate that Chain-Net robustly classifies the modulation of radio signals suffering from a synthetic channel deterioration and further performs better than other deep networks.
Thien Huynh-The, Van-Sang Doan, Cam-Hao Hua, Quoc-Viet Pham, Dong-Seong Kim 0002
GLOBECOM1
2020 Learning Geometric Features with Dual-stream CNN for 3D Action Recognition
abstract
Recently, regarding several beneficial properties of depth camera, numerous 3D action recognition frameworks have studied high-level features by exploiting deep learning techniques, but nevertheless they cannot seize the meaningful characteristics of static human pose and dynamic action motion of a whole sequence. This paper introduces a deep network configured by two parallel streams of convolutional stacks for fully learning the deep intra-frame joint associations and inter-frame joint correlations, wherein the structure of each stream is learned from Inception-v3. In experiments, besides the compatibility verification with various backbone networks, the proposed approach achieves the state-of-theart performance in battle with several deep learning-based methods on the updated NTU RGB+D 120 dataset..
Thien Huynh-The, Cam-Hao Hua, Nguyen Anh Tu, Dong-Seong Kim 0002
ICASSP1
2020 Exploiting a low-cost CNN with skip connection for robust automatic modulation classification
abstract
Recently, deep learning (DL) is an innovative machine learning (ML) technique that has gained the outstanding achievements in computer vision and natural language processing. This work takes advantage of DL for effectively handling automatic modulation classification (AMC), which is the fundamental function of numerous cognitive radio-based and spectrum sensing-based applications in many modern communication systems. Concretely, a novel deep convolutional neural network (DCNN) is proposed for learning a classification model from a massive amount of modulated signals, in which the network architecture has several convolutional blocks specialized to simultaneously capture the temporal intra-signal correlations and the spatial inter-signal relations. To this end, each block comprises various convolutional layers of asymmetric convolution kernels, whose outputs are gathered via a concatenation layer. For the enrichment of multi-scale deep feature and the prevention of gradient vanishing problem, these blocks are associated by skip connections to take into account the useful residual information. In experiments, the proposed CNN-based AMC method achieves the overall 24-modulation classification rate of 88.22% at 10dB SNR on the well-known DeepSig dataset.
Thien Huynh-The, Cam-Hao Hua, Jae-Woo Kim, Seung-Hwan Kim 0003, Dong-Seong Kim 0002
WCNC1
2020 Learning 3D spatiotemporal gait feature by convolutional network for person identification
Thien Huynh-The, Cam-Hao Hua, Nguyen Anh Tu, Dong-Seong Kim 0002
Neurocomputing1
2020 Sum-Rate Maximization for UAV-Assisted Visible Light Communications Using NOMA: Swarm Intelligence Meets Machine Learning
abstract
As the integration of unmanned aerial vehicles (UAVs) into visible light communications (VLCs) can offer many benefits for massive-connectivity applications and services in 5G and beyond, this article considers a UAV-assisted VLC using nonorthogonal multiple-access. More specifically, we formulate a joint problem of power allocation and UAV's placement to maximize the sum rate of all users, subject to constraints on power allocation, quality of service of users, and UAV's position. Since the problem is nonconvex and NP-hard in general, it is difficult to be solved optimally. Moreover, the problem is not easy to be solved by conventional approaches, e.g., coordinate descent algorithms, due to channel modeling in VLC. Therefore, we propose using the Harris hawks optimization (HHO) algorithm to solve the formulated problem and obtain an efficient solution. We then use the HHO algorithm together with artificial neural networks to propose a design that can be used in real-time applications and avoid falling into the “local minima” trap in conventional trainers. Numerical results are provided to verify the effectiveness of the proposed algorithm and further demonstrate that the proposed algorithm/HHO trainer is superior to several alternative schemes and existing metaheuristic algorithms.
Quoc-Viet Pham, Thien Huynh-The, Mamoun Alazab, Jun Zhao 0007, Won-Joo Hwang
IEEE Internet Things J.2
2020 Cross-Attentional Bracket-shaped Convolutional Network for semantic image segmentation
Cam-Hao Hua, Thien Huynh-The, Sung-Ho Bae, Sungyoung Lee 0002
Inf. Sci.2
2020 Image representation of pose-transition feature for 3D skeleton-based action recognition
Thien Huynh-The, Cam-Hao Hua, Trung-Thanh Ngo, Dong-Seong Kim 0002
Inf. Sci.1
2020 Encoding Pose Features to Images With Data Augmentation for 3-D Action Recognition
abstract
Recently, numerous methods have been introduced for three-dimensional (3-D) action recognition using handcrafted feature descriptors coupled traditional classifiers. However, they cannot learn high-level features of a whole skeleton sequence exhaustively. In this paper, a novel encoding technique - namely, pose feature to image (PoF2I), is introduced to transform the pose features of joint-joint distance and orientation to color pixels. By concatenating the features of all skeleton frames in a sequence, a color image is generated to depict spatial joint correlations and temporal pose dynamics of an action appearance. The strategy of end-to-end fine-tuning a pretrained deep convolutional neural network, which completely capture multiple high-level features at multiscale action representation, is implemented for learning recognition models. We further propose an efficient data augmentation mechanism for informative enrichment and overfitting prevention. The experimental results on six challenging 3-D action recognition datasets demonstrate that the proposed method outperforms state-of-the-art approaches.
Thien Huynh-The, Cam-Hao Hua, Dong-Seong Kim 0002
IEEE Trans. Ind. Informatics1
2019 Data Augmentation For CNN-Based 3D Action Recognition on Small-Scale Datasets
abstract
Video-based human action recognition recently plays a vital role in many industrial applications thanks to the popularity of depth sensors. A large number of conventional approaches, which have combined handcrafted features and traditional classifiers, cannot deal with various challenges in the field such as the complexity of human actions in the realistic environment. In order to improve recognition performance by exploiting more high-level discriminative features, an efficient skeleton-based action recognition method using deep convolutional neural networks (CNNs) is studied with an image encoder to transform skeleton coordinate data to image-formed data. Since deep learning techniques are fundamentally designed for efficiently working with large datasets, the network overfitting usually occurs if training CNNs on small-scale datasets. To address this issue, a novel data augmentation technique is proposed for both the informative enrichment and overfitting prevention, wherein a skeleton sequence is depicted by manifold action images based on randomly adding some skeleton frames during the data transformation and preparation for the training set. Experimental results on several small-scale challenging datasets demonstrate that the proposed method outperforms state-of-the-art approaches in terms of action recognition accuracy.
Thien Huynh-The, Dong-Seong Kim 0002
INDIN1
2019 ML-HDP: A Hierarchical Bayesian Nonparametric Model for Recognizing Human Actions in Video
abstract
Action recognition from videos is an important area of computer vision research due to its various applications, ranging from visual surveillance to human-computer interaction. To address action recognition problems, this paper presents a framework that jointly models multiple complex actions and motion units at different hierarchical levels. We achieve this by proposing a generative topic model, namely, multi-label hierarchical Dirichlet process (ML-HDP). The ML-HDP model formulates the co-occurrence relationship of actions and motion units, and enables highly accurate recognition. In particular, our topic model possesses the three-level representation in action understanding, where low-level local features are connected to high-level actions via mid-level atomic actions. This allows the recognition model to work discriminatively. In our ML-HDP, atomic actions are treated as latent topics and automatically discovered from data. In addition, we incorporate the notion of class labels into our model in a semi-supervised fashion to effectively learn and infer multi-labeled videos. Using discovered topics and inferred labels, which are jointly assigned to local features, we present the straightforward methods to perform three recognition tasks including action classification, joint classification and segmentation of continuous actions, and spatiotemporal action localization. In experiments, we explore the use of three different features and demonstrate the effectiveness of our proposed approach for these tasks on four public datasets: KTH, MSR-II, Hollywood2, and UCF101.
Nguyen Anh Tu, Thien Huynh-The, Kifayat-Ullah Khan, Young-Koo Lee
IEEE Trans. Circuits Syst. Video Technol.2
2018 Convolutional Networks with Bracket-Style Decoder for Semantic Scene Segmentation
abstract
To build up a state-of-the-art semantic scene segmentation model, a balanced combination between coarsely and finely contextual details is required for eliminating class-wise ambiguities and reaching high accuracy of pixel-wise labeling, respectively. Accordingly, with deep learning integration, prior works have achieved impressive performance in general, but found difficulties in correctly labeling medium to small objects. For the purpose of overcoming such issue, this paper proposes a deep convolutional network with bracket-style decoder, namely B-Net, to leverage the utilization of features learned at middle layers in the backbone networks (encoder) for constructing a final prediction map of densely enhanced semantic information. In particular, every feature map of interest combines with its adjacent version of higher spatial resolution through lateral connection modules to produce finer outputs that repeat such routine round-by-round until retrieving the finest-resolution map for dense prediction. Consequently, benchmarking results on CamVid dataset showed the effectiveness of the proposed method with mean class-wise accuracy, pixel-wise accuracy, and mean union intersection of 76.2%, 87.1%, and 66.4%, respectively.
Cam-Hao Hua, Thien Huynh-The, Sungyoung Lee 0002
SMC2
2018 Selective bit embedding scheme for robust blind color image watermarking
Thien Huynh-The, Cam-Hao Hua, Nguyen Anh Tu, Tae Ho Hur, Jae Hun Bang, Dohyeong Kim, Muhammad Bilal Amin, Byeong Ho Kang 0001, Hyonwoo Seung, Sungyoung Lee 0001
Inf. Sci.1
2018 Hierarchical topic modeling with pose-transition feature for action recognition using 3D skeleton data
Thien Huynh-The, Cam-Hao Hua, Nguyen Anh Tu, Tae Ho Hur, Jae Hun Bang, Dohyeong Kim, Muhammad Bilal Amin, Byeong Ho Kang 0001, Hyonwoo Seung, Soo-Yong Shin, Eun-Soo Kim, Sungyoung Lee 0001
Inf. Sci.1
2018 Early fault detection in IaaS cloud computing based on fuzzy logic and prediction technique
Dinh-Mao Bui, Thien Huynh-The, Sungyoung Lee 0001
J. Supercomput.2
2017 ADM-HIPaR: An efficient background subtraction approach
abstract
This paper presents a novel background subtraction method that is flexible for various background scenarios. The method includes automated-directional masking (ADM) algorithm for adaptive background modeling and historical intensity pattern reference (HIPaR) algorithm for foreground segmentation. By selecting an appropriate mask in a set based on directional feature, ADM updates background smoothly and precisely following a boundary-based strategy with an intensity correction rule. In order to segment foreground, HIPaR refers intensity patterns of previous backgrounds and input frames and then compares their mean difference with a checking threshold to make foreground decision. Experimental results prove that our proposed ADM-HIPaR outperforms other state-of-the-art methods in terms of foreground detection accuracy.
Thien Huynh-The, Sungyoung Lee 0002, Cam-Hao Hua
AVSS1
2017 Color image watermarking using selective MSB-LSB embedding and 2D Otsu thresholding
abstract
This paper proposes a novel digital image water-marking method, namely SMLE, that allows to intelligently embed a gray-scale watermark image into a color host image in the wavelet domain. By decomposing a gray-scale image to binary images in digits ordering from Least Significant Bit (LSB) to Most Significant Bit (MSB), binary bits are efficiently embedded to optimal wavelet coefficient blocks using a quantization technique which encodes wavelet coefficient differences to either of two pre-identified thresholds for corresponding 0-bits or 1-bits. To boost visual quality, an embedding rule is improved by equal spreading coefficient adjustment on two middle-frequency sub-bands instead of only one as existing approaches. Additionally, 2D Otsu algorithm, more proficient than 1D Otsu algorithm for binary classification under attacking scenarios, is modified to flexibly calculate an optimal threshold for high-rate watermark extraction. According to experimental results, our proposed SMLE watermarking model produces remarkable impercepti-bility as well as high robustness against common digital image transformations and mostly does better than other existing methods at same payload rate.
Thien Huynh-The, Sungyoung Lee 0002
SMC1
2017 NIC: A Robust Background Extraction Algorithm for Foreground Detection in Dynamic Scenes
abstract
This paper presents a robust foreground detection method capable of adapting to different motion speeds in scenes. A key contribution of this paper is the background estimation using a proposed novel algorithm, neighbor-based intensity correction (NIC), that identifies and modifies the motion pixels from the difference of the background and the current frame. Concretely, the first frame is considered as an initial background that is updated with the pixel intensity from each new frame based on the examination of neighborhood pixels. These pixels are formed into windows generated from the background and the current frame to identify whether a pixel belongs to the background or the current frame. The intensity modification procedure is based on the comparison of the standard deviation values calculated from two pixel windows. The robustness of the current background is further measured using pixel steadiness as an additional condition for the updating process. Finally, the foreground is detected by the background subtraction scheme with an optimal threshold calculated by the Otsu method. This method is benchmarked on several well-known data sets in the object detection and tracking domain, such as CAVIAR 2004, AVSS 2007, PETS 2009, PETS 2014, and CDNET 2014. We also compare the accuracy of the proposed method with other state-of-the-art methods via standard quantitative metrics under different parameter configurations. In the experiments, NIC approach outperforms several advanced methods on depressing the detected foreground confusions due to light artifact, illumination change, and camera jitter in dynamic scenes.
Thien Huynh-The, Oresti Baños, Sungyoung Lee 0001, Byeong Ho Kang 0001, Eun-Soo Kim, Thuong Le-Tien
IEEE Trans. Circuits Syst. Video Technol.1
2016 Describing body-pose feature - poselet - activity relationship using Pachinko Allocation Model
abstract
Understanding video-based activities have remained the challenge regardless of efforts from the image processing and artificial intelligence community. However, the rapid developing of computer vision in 3D area has brought an opportunity for the human pose estimation and so far for the activity recognition. In this research, the authors suggest an impressive approach for understanding daily life activities in the indoor using the skeleton information collected from the Microsoft Kinect device. The approach comprises two significant components as the contribution: the pose-based feature extraction under the spatio-temporal relation and the topic model based learning. For extracting feature, the distance between two articulated points and the angle between horizontal axis and joint vector are measured and normalized on each detected body. A codebook is then constructed using the K-means algorithm to encode the merged set of distance and angle. For modeling activities from sparse features, a hierarchical model developed on the Pachinko Allocation Model is proposed to describe the flexible relationship between features - poselets - activities in the temporal dimension. Finally, the activities are classified by using three different state-of-the-art machine learning techniques: Support Vector Machine, K-Nearest Neighbor, and Random Forest. In the experiment, the proposed approach is benchmarked and compared with existing methods in the overall classification accuracy.
Thien Huynh-The, Ba-Vui Le, Sungyoung Lee 0001
SMC1
2016 Improving digital image watermarking by means of optimal channel selection
Thien Huynh-The, Oresti Baños, Sungyoung Lee 0001, Yongik Yoon 0001, Thuong Le-Tien
Expert Syst. Appl.1
2016 Interactive activity recognition using pose-based spatio-temporal relation features and four-level Pachinko Allocation Model
Thien Huynh-The, Ba-Vui Le, Sungyoung Lee 0001, Yongik Yoon 0001
Inf. Sci.1