Daegil Choi

dblp:354/9659 · DBLP profile ↗
← Back
12ranked-venue papers
4as first author
12since 2021 · last 2026
0009-0003-5856-5282ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 2 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021
YearPublicationVenuePosition
2026 DepSTDSAM: Spatio-temporal dual-sparse attention model for depression prediction based on facial behavioral structure representations
Gengjia Zhang, Jisun Hong, Daegil Choi, Jaehyo Jung, Meina Li
Neurocomputing5
2025 LEFORMER: Liquid Enhanced Multimodal Learning for Depression Severity Estimation
abstract
According to the World Health Organization's (WHO) 2023 statistics, approximately 5% of people worldwide experience depression. Early diagnosis is crucial. However, misdiagnosis and delayed diagnosis are common because of professional subjectivity and reliance on patient responses. To address this issue, audio- and text-based methods for depression prediction have been a focus of recent research. However, previous methods are limited in generalizability, adaptability to new data, and prediction accuracy because they cannot fully reflect an individual's speech habits and symptom levels. To overcome this problem, this study proposes a liquid feed-forward neural network-enhanced multimodal former (LEFORMER) that incorporates an individual's symptom scores, along with learnable and dynamic parameters, into the transformer block. The LEFORMER consists of two main blocks: the symptom prediction block, which predicts patients' symptoms and incorporates third-party assessments, and the audio-text interaction block, which captures depression-related speech patterns while accounting for individual speech habits. In the depression score prediction experiment based on DAIC-WOZ, the LEFORMER achieved an MAE of 2.87 and an RMSE of 4.12, demonstrating an improvement of 0.4 in MAE and 0.71 in RMSE compared to previous studies.
Jisun Hong, Daegil Choi, Jaehyo Jung
CBMS3
2025 LTD-Conformer: Speech Depression Detection with Speaking and Listening Perspectives
abstract
Depression is a pervasive mental health problem worldwide and requires quick and accurate diagnosis. Recently, machine learning and deep learning techniques have been actively applied to depression diagnosis research, especially as audio signals are attracting attention as non-invasive and costeffective modality. This study proposes the Long-Term DilatedConformer (LTD-Conformer), an extension of the existing Conformer model designed to utilize audio signals for more accurate depression detection. The LTD-Conformer employs dilated depthwise convolution to achieve a wide receptive field and integrates a Long-Term Module to capture sequential information in audio features. This model comprehensively captures and analyzes the local, global, and sequential patterns in audio signals. In addition, we combined listening features (Mel-spectrogram) and speaking features (HuBERT) to effectively analyze both perspectives of audio signal. The experiment was conducted using DAIC-WOZ dataset, and the LTD-Conformer model achieved an accuracy of$\mathbf{8 7. 0 4 \%}$and an F1-score of 0.87, demonstrating a 4 % improvement in accuracy and 0.04 increase in the F1 score compared to the existing Conformer model. This study presents the possibility that the audio signal-based depression LTD-Conformer model can be effectively applied to depression diagnosis and will develop into a strong audio-based depression diagnosis model in the future.
Jisun Hong, Daegil Choi, Jaehyo Jung
CBMS3
2025 STSFF-Net: A two-stream network for depression detection from facial expressions
Daegil Choi, Jisun Hong, Gengjia Zhang, Jaehyo Jung
Neurocomputing1
2024 Depression Diagnosis Algorithm Based on R2U-Net Using Facial Images
abstract
Depression is a chronic and common disease that not only causes social dysfunction depending on the severity but also accompanies social problems such as suicide. Therefore, it is very important to judge the severity of depression rather than to diagnose it. We propose an algorithm that combines Convolutional Neural Networks (CNN) and R2U-Net to analyze spatial features of facial images for depression diagnosis and to verify their performance using the AVEC 2014 database. The proposed algorithm leverages Recurrent Residual Convolution Units (RRCUs) to improve feature extraction with multiple iterations to help the model improve spatial information across multiple layers. Furthermore, CNN improves the model’s ability to analyze spatial features of facial images to extract useful information for depression diagnosis. The proposed model shows improved depression score prediction performance compared to other models including U-Net, R2U-Net, and ResNet50.
Daegil Choi, Jisun Hong, Jaehyo Jung
BIBM1
2024 Depression Classification Algorithm Based on Voice Signals Using MFCC and CNN Autoencoders
abstract
Depression is a very common mental illness. In severe cases, it is a scary disease that can lead to suicide. Consequently, early diagnosis is essential because it can improve with appropriate treatment if discovered early. Recently, research on voice-based automatic depression detection systems has been actively conducted. Most existing studies to date have diagnosed depression by analyzing the characteristics of voice signals fragmented. Data is not used efficiently because the signal is analyzed only from a specific and fragmentary aspect. To solve this problem, we propose a method to extract features from both the time series and spatial aspects of the signal. First, MFCC (Mel-Frequency Cepstral Coefficient) is used to extract important features through time series analysis of speech signal. Then these extracted features are inputted into a CNN AE (Convolutional Auto-Encoder) to capture complex patterns among the initial features. This methodology can extract features that contain important information from both time series and spatial aspects. Consequently, it makes it possible to learn features and subtle signals highly associated with depression from speech data. The proposed method analyzes voice signals in both time series and spatial aspects, allowing the model to utilize more diverse information. As a result, depression can be diagnosed more accurately. To evaluate the performance of the proposed method, experiments are performed on the Distress Analysis Interview Corpus-Wizard of Oz (DAIC-WOZ) dataset and compare with frequently used feature extraction methods in voice signal analysis, such as MFCC and CNN. As a result, the proposed method showed an accuracy of 91%, an improvement of about 15% compared to previous studies.
Jisun Hong, Daegil Choi, Jaehyo Jung
ICMLA3
2024 Mental Stress Classification by Attention-Based CNN-LSTM Algorithm of Electrocardiogram Signal
abstract
Stress is a state of tension felt when exposed to a difficult situation, and excessive stress can lead to chronic diseases, method to diagnose it early is needed. Electrocardiogram (ECG) signals reflect human physiological phenomena and can be easily obtained in a noninvasive manner, which can efficiently diagnose stress. Recent studies using ECG to diagnose stress tend to use only a single or continuous cycle of ECG signals. However, if only a single cycle is used, there is a problem that the characteristics of the continuous cycle cannot be analyzed, and vice versa, the same problem arises. To solve this problem, this study proposes an Attention-based CNN-LSTM model that uses a single cycle and continuous cycle of ECG together to diagnose stress. Using a single cycle and a continuous cycle together improves the stress classification performance because it learns the long and short-term patterns of the ECG. In addition, the model in this study uses a parallel structure Convolutional neural network (CNN) to extract and combine local features of single cycles and continuous cycles, and then highlights temporal patterns and important details through Long short-term memory (LSTM) and attention mechanisms to accurately identify physiological changes in complex ECG signals. Experiments on three multi-classes using the DREAMER database have achieved an average accuracy of 97% and an average f1 score of 0.969 and shown outstanding stress analysis efficiency of the proposed model. This approach shows higher performance and more accurate stress diagnosis when using both cycles together than when using a single cycle or a continuous cycle alone.
Jisun Hong, Daegil Choi, Jaehyo Jung
ICMLA3
2024 Development of continuous cuffless blood pressure prediction platform using enhanced 1-D SENet-LSTM
Gengjia Zhang, Daegil Choi, Jaehyo Jung
Expert Syst. Appl.2
2024 Reconstruction of arterial blood pressure waveforms based on squeeze-and-excitation network models using electrocardiography and photoplethysmography signals
Gengjia Zhang, Daegil Choi, Jaehyo Jung
Knowl. Based Syst.2
2023 Depression Diagnosis Algorithm Based on 2-stream CNN Using Facial Image
abstract
In this study, a 2-stream CNN model was proposed to improve depression prediction accuracy by analyzing facial images. For image pre-processing, the AVEC2014 database was used to cut the participant's eyebrows, eyes, nasolabial folds, mouth, and tail to create images and use them as input data. It also computes and visualizes the optical flow between successive frames. It is expected that various depression patterns can be analyzed with high accuracy by predicting depression scores using a proposed model combining ResNet and Squeeze and Excitation Networks (SENet).
Daegil Choi, Gengjia Zhang, Da Eun Kim, Jaehyo Jung
ICIS1
2023 Reconstruction of ABP Waveform from the ECG or PPG Signals Using Enhanced 1-D UNetwork
abstract
ABP waveforms show how well a patient's blood responds to changes in arterial flow and can be used to analyze a patient's cardiovascular health, including heart rate analysis, stress analysis, and cardiovascular disease diagnosis. However, it is more difficult to obtain the ABP waveform than the ECG and PPG signals, and many studies are being conducted on reconstruction of the ABP waveform by recombination of bio-signals. Reconstruction of the ABP waveform has emerged as a major problem in signal distortion due to waveform destruction due to normalization/normalization of the input ABP waveform. In this study, we propose a method that can simultaneously output blood pressure estimation and ABP waveforms by inputting the raw ABP signal into the model as it is and using the U-net network coupled multi-task learning architecture model. Three signals (ECG, PPG, ECG-PPG) were trained as feature vectors through the U-net model and the results were compared. The results demonstrate the potential for replacing ECG signals with PPG signals and overcoming the current constraints of optimizing noninvasive reconstruction of ABP waveforms via U-net networks. In addition, we demonstrate that the U-net network combined multi-task learning architecture model can simultaneously perform blood pressure estimation and output of ABP waveforms without damaging the raw signal.
Gengjia Zhang, Daegil Choi, Da Eun Kim, Jaehyo Jung
ICIS2
2023 Depression Diagnosis Algorithm Based on R2U-Net Using Facial Images
abstract
Depression is a chronic and common disease that not only causes social dysfunction depending on the severity but also accompanies social problems such as suicide. Therefore, it is very important to judge the severity of depression rather than to diagnose it. Physiological studies show that depressed patients and normal people have different facial expressions. Currently, many studies are being conducted to analyze patterns of depression using physiological data. In this work, we propose a method to analyze important depression patterns by combining convolutional neural networks (CNN) and R2U-Net. For models that ensemble CNN and R2U-Net, feedback connections can be used to store information over time through the Recurrent Residual Convolution Neural Network (RCNN), allowing the network to take advantage of more and more neighbor information over time. In addition, CNN blocks and Global Average Pooling (GAP) can be used to extract useful features by including full image information in a small feature map. Through this process, complex patterns of depression can be analyzed. In the future, the video processing method will be improved to increase the performance of the model. In addition, an efficient model will be developed by evenly increasing the number of uneven data according to the severity of depression and reducing the comolexity of the model structure.
Daegil Choi, Gengjia Zhang, Jaehyo Jung
ICMLA1