Jaehyo Jung

dblp:172/5757 · DBLP profile ↗
← Back
17ranked-venue papers
1as first author
17since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 9 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 7 since 2021Computer networks · 2 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 DepSTDSAM: Spatio-temporal dual-sparse attention model for depression prediction based on facial behavioral structure representations
Gengjia Zhang, Jisun Hong, Daegil Choi, Jaehyo Jung, Meina Li
Neurocomputing6
2026 A U-Net Based LSGAN Framework for 12-Lead ECG Synthesis in Wearable IoMT Systems
abstract
Wearable devices, such as smartwatches, offer a convenient way to monitor cardiac conditions through single- or few-lead electrocardiogram (ECG) recordings. However, these devices fail to provide the comprehensive diagnostic information available from the standard 12-lead ECG used in clinical settings. To address this technological gap, this paper proposes a novel framework for generating a complete 12-lead ECG from only two leads (Lead I and II). We designed a deep learning model based on a Least Squares Generative Adversarial Network (LSGAN) that employs a U-Net architecture as its generator. The U-Net effectively extracts multi-scale features and morphological details of the ECG signal, while the LSGAN framework ensures a stable training process and generates high-fidelity waveforms. Furthermore, an L1 loss was incorporated to enforce pixel-wise accuracy in the generated signals. Experimental results on the large-scale PTB-XL dataset, using a patient-wise split for evaluation, show that our proposed model achieves a high average Pearson Correlation Coefficient (PCC) of 0.905 across all ten generated leads, outperforming state-of-the-art methods. These findings confirm that our model can reliably generate clinically significant 12-lead ECGs, presenting a practical solution for enhancing remote cardiac monitoring in the Internet of Medical Things (IoMT) environment.
Min-Gu Kim, Jaehyo Jung
IEEE Internet Things J.2
2025 LEFORMER: Liquid Enhanced Multimodal Learning for Depression Severity Estimation
abstract
According to the World Health Organization's (WHO) 2023 statistics, approximately 5% of people worldwide experience depression. Early diagnosis is crucial. However, misdiagnosis and delayed diagnosis are common because of professional subjectivity and reliance on patient responses. To address this issue, audio- and text-based methods for depression prediction have been a focus of recent research. However, previous methods are limited in generalizability, adaptability to new data, and prediction accuracy because they cannot fully reflect an individual's speech habits and symptom levels. To overcome this problem, this study proposes a liquid feed-forward neural network-enhanced multimodal former (LEFORMER) that incorporates an individual's symptom scores, along with learnable and dynamic parameters, into the transformer block. The LEFORMER consists of two main blocks: the symptom prediction block, which predicts patients' symptoms and incorporates third-party assessments, and the audio-text interaction block, which captures depression-related speech patterns while accounting for individual speech habits. In the depression score prediction experiment based on DAIC-WOZ, the LEFORMER achieved an MAE of 2.87 and an RMSE of 4.12, demonstrating an improvement of 0.4 in MAE and 0.71 in RMSE compared to previous studies.
Jisun Hong, Daegil Choi, Jaehyo Jung
CBMS4
2025 LTD-Conformer: Speech Depression Detection with Speaking and Listening Perspectives
abstract
Depression is a pervasive mental health problem worldwide and requires quick and accurate diagnosis. Recently, machine learning and deep learning techniques have been actively applied to depression diagnosis research, especially as audio signals are attracting attention as non-invasive and costeffective modality. This study proposes the Long-Term DilatedConformer (LTD-Conformer), an extension of the existing Conformer model designed to utilize audio signals for more accurate depression detection. The LTD-Conformer employs dilated depthwise convolution to achieve a wide receptive field and integrates a Long-Term Module to capture sequential information in audio features. This model comprehensively captures and analyzes the local, global, and sequential patterns in audio signals. In addition, we combined listening features (Mel-spectrogram) and speaking features (HuBERT) to effectively analyze both perspectives of audio signal. The experiment was conducted using DAIC-WOZ dataset, and the LTD-Conformer model achieved an accuracy of$\mathbf{8 7. 0 4 \%}$and an F1-score of 0.87, demonstrating a 4 % improvement in accuracy and 0.04 increase in the F1 score compared to the existing Conformer model. This study presents the possibility that the audio signal-based depression LTD-Conformer model can be effectively applied to depression diagnosis and will develop into a strong audio-based depression diagnosis model in the future.
Jisun Hong, Daegil Choi, Jaehyo Jung
CBMS4
2025 STSFF-Net: A two-stream network for depression detection from facial expressions
Daegil Choi, Jisun Hong, Gengjia Zhang, Jaehyo Jung
Neurocomputing6
2024 Depression Diagnosis Algorithm Based on R2U-Net Using Facial Images
abstract
Depression is a chronic and common disease that not only causes social dysfunction depending on the severity but also accompanies social problems such as suicide. Therefore, it is very important to judge the severity of depression rather than to diagnose it. We propose an algorithm that combines Convolutional Neural Networks (CNN) and R2U-Net to analyze spatial features of facial images for depression diagnosis and to verify their performance using the AVEC 2014 database. The proposed algorithm leverages Recurrent Residual Convolution Units (RRCUs) to improve feature extraction with multiple iterations to help the model improve spatial information across multiple layers. Furthermore, CNN improves the model’s ability to analyze spatial features of facial images to extract useful information for depression diagnosis. The proposed model shows improved depression score prediction performance compared to other models including U-Net, R2U-Net, and ResNet50.
Daegil Choi, Jisun Hong, Jaehyo Jung
BIBM4
2024 Mental Stress Detection Using PPG Signals Based on Transformer-LSTM Model
abstract
Long-term stress not only affects physical health but also leads to the deterioration of mental state. To prevent such adverse effects, it is very important to detect and manage stress promptly. Therefore, it is particularly necessary to develop effective stress detection methods. This study utilizes the photoplethysmogram (PPG) signal in the WESAD database to develop a deep learning model for detecting psychological stress. The PPG signal is preprocessed by data filtering, down sampling, and sliding windows. The local features of the PPG signal are extracted using a convolutional neural network (CNN). Then, the Transformer model is used to analyze and capture the correlation between features. The Transformer is based on the attention mechanism, which enables the model to focus on important local features. Lastly, the long short-term memory (LSTM) is utilized to process the temporal relationship between features and detect stress status. Evaluations on binary, ternary, and quaternary stress classification tasks show significant performance: 96.07%, 92.27%, and 87.04% accuracy are achieved, respectively. Comparison with other recent research findings demonstrates the superiority of the proposed method. The proposed method does not require manual feature extraction engineering, simplifies the data processing process, and improves the accuracy of stress detection, indicating its potential in real-life and practical applications.
Ziyu Shi, Gengjia Zhang, Jaehyo Jung, Meina Li
BIBM4
2024 Building an Explainable Deep Residual Network for Depression Detection from ECG Signals
abstract
This study focuses on developing and validating interpretable deep learning models for depression detection using the proprietary Database. The preprocessing of the proprietary database involves noise reduction, normalization, and segment extraction. The paper proposes a Deep ResNet Network that utilizes overlapping ECG signal segments to capture underlying patterns related to depression. To enhance the understanding of important areas of the input ECG signal in the model, the SHAP (SHapley Additive exPlanations) framework was introduced to increase the interpretability of the results. Independent testing confirmed the robustness and reliability of the Deep ResNet Network, which achieved an accuracy of 92.63%, establishing it as a promising and interpretable tool for future depression diagnosis.
Gengjia Zhang, Ziyu Shi, Jaehyo Jung, Meina Li
BIBM4
2024 Depression Classification Algorithm Based on Voice Signals Using MFCC and CNN Autoencoders
abstract
Depression is a very common mental illness. In severe cases, it is a scary disease that can lead to suicide. Consequently, early diagnosis is essential because it can improve with appropriate treatment if discovered early. Recently, research on voice-based automatic depression detection systems has been actively conducted. Most existing studies to date have diagnosed depression by analyzing the characteristics of voice signals fragmented. Data is not used efficiently because the signal is analyzed only from a specific and fragmentary aspect. To solve this problem, we propose a method to extract features from both the time series and spatial aspects of the signal. First, MFCC (Mel-Frequency Cepstral Coefficient) is used to extract important features through time series analysis of speech signal. Then these extracted features are inputted into a CNN AE (Convolutional Auto-Encoder) to capture complex patterns among the initial features. This methodology can extract features that contain important information from both time series and spatial aspects. Consequently, it makes it possible to learn features and subtle signals highly associated with depression from speech data. The proposed method analyzes voice signals in both time series and spatial aspects, allowing the model to utilize more diverse information. As a result, depression can be diagnosed more accurately. To evaluate the performance of the proposed method, experiments are performed on the Distress Analysis Interview Corpus-Wizard of Oz (DAIC-WOZ) dataset and compare with frequently used feature extraction methods in voice signal analysis, such as MFCC and CNN. As a result, the proposed method showed an accuracy of 91%, an improvement of about 15% compared to previous studies.
Jisun Hong, Daegil Choi, Jaehyo Jung
ICMLA4
2024 Mental Stress Classification by Attention-Based CNN-LSTM Algorithm of Electrocardiogram Signal
abstract
Stress is a state of tension felt when exposed to a difficult situation, and excessive stress can lead to chronic diseases, method to diagnose it early is needed. Electrocardiogram (ECG) signals reflect human physiological phenomena and can be easily obtained in a noninvasive manner, which can efficiently diagnose stress. Recent studies using ECG to diagnose stress tend to use only a single or continuous cycle of ECG signals. However, if only a single cycle is used, there is a problem that the characteristics of the continuous cycle cannot be analyzed, and vice versa, the same problem arises. To solve this problem, this study proposes an Attention-based CNN-LSTM model that uses a single cycle and continuous cycle of ECG together to diagnose stress. Using a single cycle and a continuous cycle together improves the stress classification performance because it learns the long and short-term patterns of the ECG. In addition, the model in this study uses a parallel structure Convolutional neural network (CNN) to extract and combine local features of single cycles and continuous cycles, and then highlights temporal patterns and important details through Long short-term memory (LSTM) and attention mechanisms to accurately identify physiological changes in complex ECG signals. Experiments on three multi-classes using the DREAMER database have achieved an average accuracy of 97% and an average f1 score of 0.969 and shown outstanding stress analysis efficiency of the proposed model. This approach shows higher performance and more accurate stress diagnosis when using both cycles together than when using a single cycle or a continuous cycle alone.
Jisun Hong, Daegil Choi, Jaehyo Jung
ICMLA4
2024 Development of continuous cuffless blood pressure prediction platform using enhanced 1-D SENet-LSTM
Gengjia Zhang, Daegil Choi, Jaehyo Jung
Expert Syst. Appl.3
2024 Development of Optimized User-Recognition Technology Using Multilayered XAI-Based ECG Signals
abstract
Owing to the swift advancements in information and technology and wearable device technology, research involving bio-signals is being conducted for convenient personal identification authentication. Therefore, deep learning using electrocardiogram (ECG) signals is being developed as a next-generation user-recognition technology for application in real-life environments based on high security and accuracy. This study proposes a deep learning model that combines multilayered explainable artificial intelligence for user recognition in a real-life environment. The ECG-signal database utilizes the MIT-BIH normal sinus rhythm database (NSRDB) and a self-acquired database. A comparison between a single deep learning model and the proposed multilayer deep learning model using these signals confirmed that the proposed model achieved higher accuracy, with minimum accuracy differences of 15% and 0.2%. In addition, the recognition accuracies of the single deep learning model and multilayered deep learning model were compared five times to confirm the difference in output results caused by various factors in the single deep learning model. The user-recognition accuracies of a single deep learning model for the NSRDB exhibited a difference of 3.4%, whereas those of the multilayered model showed a difference of 0.9%. The user-recognition accuracy difference when applying the self-acquired database to the single deep learning model was 1.3%, but that of the multilayered deep learning model was 0.3%. Therefore, the multilayered deep learning model exhibited a stable user-recognition accuracy compared to that of the single structure, and the optimized area that affects the output result was confirmed.
Jaehyo Jung
IEEE Internet Things J.1
2024 Reconstruction of arterial blood pressure waveforms based on squeeze-and-excitation network models using electrocardiography and photoplethysmography signals
Gengjia Zhang, Daegil Choi, Jaehyo Jung
Knowl. Based Syst.3
2023 Depression Diagnosis Algorithm Based on 2-stream CNN Using Facial Image
abstract
In this study, a 2-stream CNN model was proposed to improve depression prediction accuracy by analyzing facial images. For image pre-processing, the AVEC2014 database was used to cut the participant's eyebrows, eyes, nasolabial folds, mouth, and tail to create images and use them as input data. It also computes and visualizes the optical flow between successive frames. It is expected that various depression patterns can be analyzed with high accuracy by predicting depression scores using a proposed model combining ResNet and Squeeze and Excitation Networks (SENet).
Daegil Choi, Gengjia Zhang, Da Eun Kim, Jaehyo Jung
ICIS4
2023 Reconstruction of ABP Waveform from the ECG or PPG Signals Using Enhanced 1-D UNetwork
abstract
ABP waveforms show how well a patient's blood responds to changes in arterial flow and can be used to analyze a patient's cardiovascular health, including heart rate analysis, stress analysis, and cardiovascular disease diagnosis. However, it is more difficult to obtain the ABP waveform than the ECG and PPG signals, and many studies are being conducted on reconstruction of the ABP waveform by recombination of bio-signals. Reconstruction of the ABP waveform has emerged as a major problem in signal distortion due to waveform destruction due to normalization/normalization of the input ABP waveform. In this study, we propose a method that can simultaneously output blood pressure estimation and ABP waveforms by inputting the raw ABP signal into the model as it is and using the U-net network coupled multi-task learning architecture model. Three signals (ECG, PPG, ECG-PPG) were trained as feature vectors through the U-net model and the results were compared. The results demonstrate the potential for replacing ECG signals with PPG signals and overcoming the current constraints of optimizing noninvasive reconstruction of ABP waveforms via U-net networks. In addition, we demonstrate that the U-net network combined multi-task learning architecture model can simultaneously perform blood pressure estimation and output of ABP waveforms without damaging the raw signal.
Gengjia Zhang, Daegil Choi, Da Eun Kim, Jaehyo Jung
ICIS4
2023 Depression Diagnosis Algorithm Based on R2U-Net Using Facial Images
abstract
Depression is a chronic and common disease that not only causes social dysfunction depending on the severity but also accompanies social problems such as suicide. Therefore, it is very important to judge the severity of depression rather than to diagnose it. Physiological studies show that depressed patients and normal people have different facial expressions. Currently, many studies are being conducted to analyze patterns of depression using physiological data. In this work, we propose a method to analyze important depression patterns by combining convolutional neural networks (CNN) and R2U-Net. For models that ensemble CNN and R2U-Net, feedback connections can be used to store information over time through the Recurrent Residual Convolution Neural Network (RCNN), allowing the network to take advantage of more and more neighbor information over time. In addition, CNN blocks and Global Average Pooling (GAP) can be used to extract useful features by including full image information in a small feature map. Through this process, complex patterns of depression can be analyzed. In the future, the video processing method will be improved to increase the performance of the model. In addition, an efficient model will be developed by evenly increasing the number of uneven data according to the severity of depression and reducing the comolexity of the model structure.
Daegil Choi, Gengjia Zhang, Jaehyo Jung
ICMLA3
2021 Estimation of a physical activity energy expenditure with a patch-type sensor module using artificial neural network
abstract
Summary Chronic diseases such as coronary artery diseases and diabetes are caused by lack of physical activities and are leading causes of high death and morbidity rates. In particular, the imbalance of consumption energy and intake energy has increased adult diseases such as obesity with high mortality. Until recently, direct calorimetry by production calorie and indirect calorimetry by energy expenditure have been regarded as the best methods for estimating physical activity and energy expenditure. These calorimetry methods are associated with limited practicality such as data acquisition in a limited time, high cost, and wearing an inconvenient mask for oxygen uptake measurement. In this study, we propose the most accurate method using a wireless patch‐type sensor to predict the energy expenditure of physical activities. Through the optimization of the prediction of energy expenditure of physical activities using the neural network algorithm, we achieved RMSE of 0.1893 and R2 of 0.91 for the energy expenditures of aerobic and anaerobic exercises. These results indicate that the proposed system is useful and reliable for monitoring user's energy expenditure when using attached patch‐type sensors workouts.
Kyeung Ho Kang, Siho Shin, Jaehyo Jung, Youn Tae Kim
Concurr. Comput. Pract. Exp.3