VLDB 2026 Research / reviewers in the wild / expert
Seung-Hee Yang
dblp:143/3966 · also Seung Hee Yang
· DBLP profile ↗
10ranked-venue papers
1as first author
8since 2021 · last 2024
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 1 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Towards Intelligent Speech Assistants in Operating Rooms: A Multimodal Model for Surgical Workflow Analysis
Kubilay Can Demir, Belén Lojo Rodríguez, Tobias Weise, Andreas K. Maier, Seung-Hee Yang |
INTERSPEECH | 5 |
| 2024 | Speaker- and Text-Independent Estimation of Articulatory Movements and Phoneme Alignments from Speech
Tobias Weise, Philipp Klumpp, Kubilay Can Demir, Paula Andrea Pérez-Toro, Maria Schuster, Elmar Nöth, Björn Heismann, Andreas K. Maier, Seung-Hee Yang |
INTERSPEECH | 9 |
| 2024 | Indoor Synthetic Data Generation: A Systematic ReviewabstractDeep learning-based object recognition, 6D pose estimation, and semantic scene understanding require a large amount of training data to achieve generalization. Time-consuming annotation processes, privacy, and security aspects lead to a scarcity of real-world datasets. To overcome this lack of data, synthetic data generation has been proposed, including multiple facets in the area of domain randomization to extend the data distribution. The objective of this review is to identify methods applied for synthetic data generation aiming to improve 6D pose estimation, object recognition, and semantic scene understanding in indoor scenarios. We further review methods used to extend the data distribution and discuss best practices to bridge the gap between synthetic and real-world data. We adhered to the guidelines of the systematic PRISMA technique. Three databases, IEEE Xplore, Springer Link, and ACM, and an additional manual search were conducted. In total, we identified 241 studies and included 34 in our systematic review. In summary, synthetic data generation has been performed using crop-out methods, graphic APIs, 3D modeling or authoring tools, or game engine-based methods. To extend the data distribution, varying scene parameters, i.e., lighting conditions or textures and the use of distracting objects in the scene are promising. Hannah Schieber, Kubilay Can Demir, Constantin Kleinbeck, Seung-Hee Yang, Daniel Roth 0001 |
Comput. Vis. Image Underst. | 4 |
| 2023 | Federated Learning for Secure Development of AI Models for Parkinson's Disease Detection Using Speech from Different LanguagesabstractParkinson's disease (PD) is a neurological disorder impacting a person's speech. Among automatic PD assessment methods, deep learning models have gained particular interest. Recently, the community has explored cross-pathology and cross-language models which can improve diagnostic accuracy even further. However, strict patient data privacy regulations largely prevent institutions from sharing patient speech data with each other. In this paper, we employ federated learning (FL) for PD detection using speech signals from 3 real-world language corpora of German, Spanish, and Czech, each from a separate institution. Our results indicate that the FL model outperforms all the local models in terms of diagnostic accuracy, while not performing very differently from the model based on centrally combined training sets, with the advantage of not requiring any data sharing among collaborators. This will simplify inter-institutional collaborations, resulting in enhancement of patient outcomes. Soroosh Tayebi Arasteh, Cristian D. Ríos-Urrego, Elmar Nöth, Andreas K. Maier, Seung-Hee Yang, Jan Rusz, Juan Rafael Orozco-Arroyave |
INTERSPEECH | 5 |
| 2023 | PoCaPNet: A Novel Approach for Surgical Phase Recognition Using Speech and X-Ray ImagesabstractSurgical phase recognition is a challenging and necessary task for the development of context-aware intelligent systems that can support medical personnel for better patient care and effective operating room management.In this paper, we present a surgical phase recognition framework that employs a Multi-Stage Temporal Convolution Network using speech and X-Ray images for the first time.We evaluate our proposed approach using our dataset that comprises 31 port-catheter placement operations and report 82.56 % frame-wise accuracy with eight surgical phases.Additionally, we investigate the design choices in the temporal model and solutions for the class-imbalance problem.Our experiments demonstrate that speech and X-Ray data can be effectively utilized for surgical phase recognition, providing a foundation for the development of speech assistants in operating rooms of the future. Kubilay Can Demir, Tobias Weise, Matthias S. May, Axel Schmid, Andreas K. Maier, Seung-Hee Yang |
INTERSPEECH | 6 |
| 2023 | Deep Learning in Surgical Workflow Analysis: A Review of Phase and Step RecognitionabstractOBJECTIVE: In the last two decades, there has been a growing interest in exploring surgical procedures with statistical models to analyze operations at different semantic levels. This information is necessary for developing context-aware intelligent systems, which can assist the physicians during operations, evaluate procedures afterward or help the management team to effectively utilize the operating room. The objective is to extract reliable patterns from surgical data for the robust estimation of surgical activities performed during operations. The purpose of this article is to review the state-of-the-art deep learning methods that have been published after 2018 for analyzing surgical workflows, with a focus on phase and step recognition. METHODS: Three databases, IEEE Xplore, Scopus, and PubMed were searched, and additional studies are added through a manual search. After the database search, 343 studies were screened and a total of 44 studies are selected for this review. CONCLUSION: The use of temporal information is essential for identifying the next surgical action. Contemporary methods used mainly RNNs, hierarchical CNNs, and Transformers to preserve long-distance temporal relations. The lack of large publicly available datasets for various procedures is a great challenge for the development of new and robust models. As supervised learning strategies are used to show proof-of-concept, self-supervised, semi-supervised, or active learning methods are used to mitigate dependency on annotated data. SIGNIFICANCE: The present study provides a comprehensive review of recent methods in surgical workflow analysis, summarizes commonly used architectures, datasets, and discusses challenges. Kubilay Can Demir, Hannah Schieber, Tobias Weise, Daniel Roth 0001, Matthias S. May, Andreas K. Maier, Seung-Hee Yang |
IEEE J. Biomed. Health Informatics | 7 |
| 2022 | Cross-lingual Self-Supervised Speech Representations for Improved Dysarthric Speech RecognitionabstractState-of-the-art automatic speech recognition (ASR) systems perform well on healthy speech.However, the performance on impaired speech still remains an issue.The current study explores the usefulness of using Wav2Vec self-supervised speech representations as features for training an ASR system for dysarthric speech.Dysarthric speech recognition is particularly difficult as several aspects of speech such as articulation, prosody and phonation can be impaired.Specifically, we train an acoustic model with features extracted from Wav2Vec, Hubert, and the cross-lingual XLSR model.Results suggest that speech representations pretrained on large unlabelled data can improve word error rate (WER) performance.In particular, features from the multilingual model led to lower WERs than filterbanks (Fbank) or models trained on a single language.Improvements were observed in English speakers with cerebral palsy caused dysarthria (UASpeech corpus), Spanish speakers with Parkinsonian dysarthria (PC-GITA corpus) and Italian speakers with paralysis-based dysarthria (EasyCall corpus).Compared to using Fbank features, XLSR-based features reduced WERs by 6.8%, 22.0%, and 7.0% for the UASpeech, PC-GITA, and EasyCall corpus, respectively. Abner Hernandez, Paula Andrea Pérez-Toro, Elmar Nöth, Juan Rafael Orozco-Arroyave, Andreas K. Maier, Seung-Hee Yang |
INTERSPEECH | 6 |
| 2022 | Disentangled Latent Speech Representation for Automatic Pathological Intelligibility AssessmentabstractSpeech intelligibility assessment plays an important role in the therapy of patients suffering from pathological speech disorders.Automatic and objective measures are desirable to assist therapists in their traditionally subjective and labor-intensive assessments.In this work, we investigate a novel approach for obtaining such a measure using the divergence in disentangled latent speech representations of a parallel utterance pair, obtained from a healthy reference and a pathological speaker.Experiments on an English database of Cerebral Palsy patients, using all available utterances per speaker, show high and significant correlation values (R = -0.9)with subjective intelligibility measures, while having only minimal deviation (±0.01) across four different reference speaker pairs.We also demonstrate the robustness of the proposed method (R = -0.89deviating ±0.02 over 1000 iterations) by considering a significantly smaller amount of utterances per speaker.Our results are among the first to show that disentangled speech representations can be used for automatic pathological speech intelligibility assessment, resulting in a reference speaker pair invariant method, applicable in scenarios with only few utterances available. Tobias Weise, Philipp Klumpp, Andreas K. Maier, Elmar Nöth, Björn Heismann, Maria Schuster, Seung-Hee Yang |
INTERSPEECH | 7 |
| 2019 | Self-Imitating Feedback Generation Using GAN for Computer-Assisted Pronunciation TrainingabstractSelf-imitating feedback is an effective and learner-friendly method for non-native learners in Computer-Assisted Pronunciation Training. Acoustic characteristics in native utterances are extracted and transplanted onto learner's own speech input, and given back to the learner as a corrective feedback. Previous works focused on speech conversion using prosodic transplantation techniques based on PSOLA algorithm. Motivated by the visual differences found in spectrograms of native and non-native speeches, we investigated applying GAN to generate self-imitating feedback by utilizing generator's ability through adversarial training. Because this mapping is highly under-constrained, we also adopt cycle consistency loss to encourage the output to preserve the global structure, which is shared by native and non-native utterances. Trained on 97,200 spectrogram images of short utterances produced by native and non-native speakers of Korean, the generator is able to successfully transform the non-native spectrogram input to a spectrogram with properties of self-imitating feedback. Furthermore, the transformed spectrogram shows segmental corrections that cannot be obtained by prosodic transplantation. Perceptual test comparing the self-imitating and correcting abilities of our method with the baseline PSOLA method shows that the generative approach with cycle consistency loss is promising. Seung-Hee Yang, Minhwa Chung |
INTERSPEECH | 1 |
| 2019 | WithDorm: Dormitory Solution for Linking RoommatesabstractExperiences in universities are important for emotional maturation and offer an opportunity to develop individual characteristics and skills needed for social life. There are diverse issues affecting the quality of dormitory life and roommate relationships, which can influence one's psychosocial development. In this paper, we propose WithDorm, a mobile application to help communication with roommates and tighten their connections, and thereby assisting the users' emotional health and psychosocial development. We analyzed dormitory roommate issues from a human-centered perspective and narrowed down to three design implications after dormitory life modeling. Furthermore, we implemented the design implications in a prototype and performed a usability test to evaluate and improve the design. The final design, WithDorm, is aware of dormitory-specific concerns, collects and adapts to users' lifestyles, and initiates humanhuman interaction among roommates. Minji Kwak, Seung-Hee Yang, Jaeseo Lim, Byoung-Tak Zhang |
MobileHCI | 3 |