Suyeon Lee

dblp:10/8703 · DBLP profile ↗
← Back
11ranked-venue papers
4as first author
9since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Security and privacy · 2 · 1 first-authorSoftware engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Griffin: Coherency-Aware Task Scheduling and Memory Allocation for CXL Interconnects
abstract
CXL is an emerging interconnect that has the potential to efficiently realize memory disaggregation. This is because CXL enables the expansion of memory beyond individual hosts, and supports coherent memory sharing among multiple hosts. However, CXL introduces several performance overheads due to the cache coherency protocol for memory sharing, as well as placement constraints for shared data, which, if ignored, can lead to correctness issues. This paper presents the first analysis of the impact of CXL memory sharing and shows that the overheads of hardware-based coherency in CXL interconnects are substantial. We then propose Griffin, a new coherency-aware task and memory allocator for CXL disaggregated memory systems. Griffin introduces new abstractions and algorithms that allow it to prioritize which data is allocated remotely and to which memory node, to efficiently reduce the coherence overheads associated with both the amount of shared data and the load on CXL coherence resources. Our simulation results show that Griffin reduces the total memory time by up to 4.29 × compared to a standard baseline and 1.71 × compared to an advanced baseline.
Suyeon Lee, Khaled Diab 0001, Diman Zad Tootaghaj, Lianjie Cao, Puneet Sharma 0001, Ada Gavrilovska
ICS1
2026 AXLE: Coordinated Offloading with Asynchronous Back-Streaming in Computational Memory Systems
Suyeon Lee, Kangkyu Park, Kwangsik Shin, Ada Gavrilovska
ISCA1
2025 Grudon: A System for Deploying Graph Workloads on Disaggregated Architectures with Near-Data Processing
abstract
Emerging memory disaggregation and near-data processing (NDP) technologies have shown promise in addressing performance, scaling and efficiency challenges for data-intensive workloads such as graph analytics. However, they expose new tradeoffs related to data movement, runtime, and cost. This paper introduces Grudon, a novel system for deploying graph workloads on Disaggregated NDP (DiNDP) platforms. Grudon combines the benefits of memory disaggregation and NDP to tackle bandwidth limitations, improve resource flexibility, and reduce energy demands to improve overall workload performance. To achieve this, Grudon provides support for dynamic configuration of graph workload deployments on DiNDP platforms via new control plane functionality that enables lightweight evaluation of the tradeoff space and dynamic reconfiguration of operations executed on the different DiNDP elements. Evaluation on an emulated DiNDP testbed demonstrates that Grudon achieves a geomean speedup of 2.12× while reducing energy consumption by 3.3× compared to state-of-the-art distributed runtimes. These results show the potential to unlock the full capabilities of DiNDP platforms for efficient distributed graph analytics.
Vishal Rao, Nikhil Ram Shashidhar, Suyeon Lee, Ada Gavrilovska
HPDC3
2025 KMI: A Dataset of Korean Motivational Interviewing Dialogues for Psychotherapy
abstract
Hyunjong Kim, Suyeon Lee, Yeongjae Cho, Eunseo Ryu, Yohan Jo, Suran Seong, Sungzoon Cho. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Hyunjong Kim, Suyeon Lee, Yeongjae Cho, Eunseo Ryu, Yohan Jo, Suran Seong, Sungzoon Cho
NAACL (Long Papers)2
2025 Model-Guided Dual-Role Alignment for High-Fidelity Open-Domain Video-to-Audio Generation
abstract
We present MGAudio, a novel flow-based framework for open-domain video-to-audio generation, which introduces model-guided dual-role alignment as a central design principle. Unlike prior approaches that rely on classifier-based or classifier-free guidance, MGAudio enables the generative model to guide itself through a dedicated training objective designed for video-conditioned audio generation. The framework integrates three main components: (1) a scalable flow-based Transformer denoiser, (2) a dual-role alignment mechanism where the audio-visual encoder serves both as a conditioning module and as a feature aligner to improve generation quality, and (3) a model-guided objective that enhances cross-modal coherence and audio realism. MGAudio achieves state-of-the-art performance on VGGSound, reducing FAD to 0.40, substantially surpassing the best classifier-free guidance baselines, and consistently outperforms existing methods across FD, IS, and alignment metrics. It also generalizes well to the challenging UnAV-100 benchmark. These results highlight model-guided dual-role alignment as a powerful and scalable paradigm for conditional video-to-audio generation. Code is available at: https://github.com/pantheon5100/mgaudio
Kang Zhang 0008, Trung X. Pham, Suyeon Lee, Axi Niu, Arda Senocak, Joon Son Chung
NeurIPS3
2024 TalkNCE: Improving Active Speaker Detection with Talk-Aware Contrastive Learning
abstract
The goal of this work is Active Speaker Detection (ASD), a task to determine whether a person is speaking or not in a series of video frames. Previous works have dealt with the task by exploring network architectures while learning effective representations has been less explored. In this work, we propose TalkNCE, a novel talk-aware contrastive loss. The loss is only applied to part of the full segments where a person on the screen is actually speaking. This encourages the model to learn effective representations through the natural correspondence of speech and facial movements. Our loss can be jointly optimized with the existing objectives for training ASD models without the need for additional supervision or training data. The experiments demonstrate that our loss can be easily integrated into the existing ASD frameworks, improving their performance. Our method achieves state-of-the-art performances on AVA-ActiveSpeaker and ASW datasets.
Chaeyoung Jung, Suyeon Lee, Kihyun Nam, Kyeongha Rho, You Jin Kim, Youngjoon Jang 0001, Joon Son Chung
ICASSP2
2024 Seeing Through The Conversation: Audio-Visual Speech Separation Based on Diffusion Model
abstract
The objective of this work is to extract the target speaker’s voice from a mixture of voices using visual cues. Existing works on audio-visual speech separation have demonstrated their performance with promising intelligibility, but maintaining naturalness remains challenging. To address this issue, we propose AVDiffuSS, an audio-visual speech separation model based on a diffusion mechanism known for its capability to generate natural samples. We also propose a cross-attention-based feature fusion mechanism for an effective fusion of the two modalities for diffusion. This mechanism is specifically tailored for the speech domain to integrate the phonetic information from audio-visual correspondence in speech generation. In this way, the fusion process maintains the high temporal resolution of the features, without excessive computational requirements. We demonstrate that the proposed framework achieves state-of-the-art results on two benchmarks, including VoxCeleb2 and LRS3, producing speech with notably better naturalness. Project page with demo: https://mm.kaist.ac.kr/projects/avdiffuss/
Suyeon Lee, Chaeyoung Jung, Youngjoon Jang 0001, Jaehun Kim, Joon Son Chung
ICASSP1
2024 An Autonomous Parallelization of Transformer Model Inference on Heterogeneous Edge Devices
abstract
The utilization of advancing transformer-based deep neural network (DNN) models in edge environments holds the promise of improving productivity for intelligent tasks. However, deploying these models on edge devices with limited resources encounters significant performance challenges. Previous solutions have attempted to distribute computation tasks across devices and perform parallel inferences but often fall short of meeting service-level objectives (SLO). This limitation arises from their inability to effectively harness parallelization in transformer-based models and consider the resource diversity of edge devices. In this paper, we propose Hepti, a practical framework designed to facilitate parallel inference of transformer-based DNN models on heterogeneous edge environments. Hepti is armed with: 1) an understanding of transformer model architecture to enable effective parallel inference and 2) dynamic workload optimization to adapt to changing network and device resource capabilities. Our evaluations confirmed that the Hepti autonomously assesses the resource diversity of edge devices and network status. Furthermore, Hepti achieves a maximum performance improvement of 49.1% and 37.1% compared to the local inference approach and state-of-the-art model parallelisms on the BERT-Large model.
Juhyeon Lee, Insung Bahk, Hoseung Kim, Sinjin Jeong, Suyeon Lee, Donghyun Min
ICS5
2024 FlowAVSE: Efficient Audio-Visual Speech Enhancement with Conditional Flow Matching
Chaeyoung Jung, Suyeon Lee, Joon Son Chung
INTERSPEECH2
2015 Screening smartphone applications using malware family signatures
Jehyun Lee, Suyeon Lee, Heejo Lee
Comput. Secur.2
2013 Screening Smartphone Applications Using Behavioral Signatures
Suyeon Lee, Jehyun Lee, Heejo Lee
SEC1