Soonyoung Lee

dblp:18/6230 · DBLP profile ↗
← Back
13ranked-venue papers
1as first author
10since 2021 · last 2026
0009-0005-0764-337XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 1 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 1 first-author · 8 since 2021Systems, architecture and hardware · 2
YearPublicationVenuePosition
2026 RAPID: A Rapid Prototyping Platform for Industrial Automation
abstract
Industrial automation in smart logistics and factories requires simulation platforms that support rapid environment building before costly physical deployment. Yet existing tools often require substantial expertise, complex setup, and long configuration times, hindering agile prototyping. We present RAPID, a simulation platform with two components: layout design, which enables intuitive visual configuration of factory layouts, and behavior simulation and validation, which allows users to attach behavior models and evaluate system performance. RAPID lowers the entry barrier to industrial simulation, letting users apply existing behavior models or trained reinforcement learning (RL) agents to new layouts with minimal effort. This approach lets practitioners prototype facilities in minutes rather than weeks and gives researchers a standardized environment for benchmarking multi-agent RL and coordination algorithms. By combining rapid design with simulation-based validation, RAPID accelerates automation development from concept to implementation.
Sunghoon Hong, Whiyoung Jung, Deunsol Yoon, Woohyung Lim, Soonyoung Lee, Kanghoon Lee
AAAI6
2026 RL-Studio: A System for Multi-Phase Reinforcement Learning Experimentation
abstract
Reinforcement learning (RL) has evolved beyond monolithic training, yet existing frameworks remain limited to single algorithms or simple offline-to-online transitions. We present multi-phase RL, a framework that orchestrates multiple learning phases for continual policy improvement. It enables efficient fine-tuning of pretrained policies with new data and smooth adaptation from simulation to real-world environments. To support this paradigm, we introduce RL-Studio, a platform that addresses key implementation barriers, including neural architecture mismatches, parameter transfer complexities, and experiment management overhead. It provides phase orchestration, transition-point monitoring, and full experiment lineage tracking. We demonstrate the effectiveness of multi-phase RL through representative scenarios and highlight RL-Studio’s capabilities.
Whiyoung Jung, Sunghoon Hong, Deunsol Yoon, Jeonghye Kim, Yongjae Shin, Suhyun Jung, Hyundam Yoo, Chanwoo Moon, Woohyung Lim, Soonyoung Lee, Kanghoon Lee
AAAI11
2026 AEGIS: Toward Expert-in-the-loop Industrial Anomaly Detection
abstract
Anomaly detection platforms in real-world environments require continuous interaction between automated systems and domain experts, as anomalies evolve dynamically and their definitions vary across contexts. Therefore, an effective platform must collaborate with experts and incorporate their feedback to update the system. This paper introduces AEGIS, an anomaly detection platform that aims to support interaction between domain experts and data-driven agents through three core capabilities: (1) data-driven insights through real-time monitoring, explanations, and distribution shift detection, which invoke customized tools to generate appropriate responses, (2) an expert feedback interface for labeling and direct updates via chat-based interaction, and (3) autonomous model construction that leverages expert-labeled data with LLM-driven hyperparameter optimization. Through this design, AEGIS fosters continuous interaction in which the platform provides insights while experts guide model improvement, ensuring user intent is reflected and robustness is maintained under evolving data distributions.
Ye Seul Sim, Suhee Yoon, Sanghyu Yoon, Seungdong Yoa, Soonyoung Lee, Woohyung Lim
AAAI6
2026 OrcheCause Agent: From Textual Knowledge to End-to-End Causal Inference
abstract
Causal agents have emerged as promising tools for automating causal analysis based on user queries. However, existing causal agent systems are often limited to a single causal task, limiting their ability to handle complex queries. In addition, they accept only numerical data as input, preventing the integration of domain knowledge expressed in natural language. To overcome these limitations, we propose the OrcheCause agent, a causal agent leveraging textual knowledge for end-to-end causal inference. Specifically, OrcheCause is designed to orchestrate a sequence of interrelated causal tasks in response to user queries. Furthermore, OrcheCause supports diverse data types—numerical as well as textual data—by extracting cause-effect pairs from the relevant sources and incorporating them into causal discovery (CD), thereby improving the performance of CD. OrcheCause also introduces a metric-based hyperparameter optimization framework for CD when ground-truth graphs are not available.
Jinseok Yang, Juhyun Lyu, Soonyoung Lee, Woohyung Lim
AAAI4
2025 MASH-VLM: Mitigating Action-Scene Hallucination in Video-LLMs through Disentangled Spatial-Temporal Representations
abstract
In this work, we tackle action-scene hallucination in Video Large Language Models (Video-LLMs), where models incorrectly predict actions based on the scene context or scenes based on observed actions. We observe that existing Video-LLMs often suffer from action-scene hallucination due to two main factors. First, existing Video-LLMs intermingle spatial and temporal features by applying an attention operation across all tokens. Second, they use standard Rotary Position Embedding (RoPE), which causes the text tokens to overemphasize certain types of tokens depending on their sequential orders. To address these issues, we introduce MASH-VLM, Mitigating Action-Scene Hallucination in Video-LLMs through disentangled spatial-temporal representations. Our approach includes two key innovations: (1) DST-attention, a novel attention mechanism that disentangles spatial and temporal tokens within the LLM by using masked attention to restrict direct interactions be tween spatial and temporal tokens; (2) Harmonic-RoPE, which extends the dimensionality of the positional IDs, allowing spatial and temporal tokens to maintain balanced positions relative to the text tokens. To evaluate the action-scene hallucination in Video-LLMs, we introduce the UN-SCENE benchmark with 1,320 videos and 4,078 QA pairs. MASH-VLM achieves state-of-the-art performance on the UNSCENE benchmark, as well as on existing video understanding benchmarks.
Kyungho Bae, Jinhyung Kim, Sihaeng Lee, Soonyoung Lee, Jinwoo Choi 0001
CVPR4
2025 ReSpec: Relevance and Specificity Grounded Online Filtering for Learning on Video-Text Data Streams
abstract
The rapid growth of video-text data presents challenges in storage and computation during training. Online learning, which processes streaming data in real-time, offers a promising solution to these issues while also allowing swift adaptations in scenarios demanding real-time responsiveness. One strategy to enhance the efficiency and effectiveness of learning involves identifying and prioritizing data that enhances performance on target downstream tasks. We propose Relevance and Specificity-based online filtering framework (ReSpec) that selects data based on four criteria: (i) modality alignment for clean data, (ii) task relevance for target focused data, (iii) specificity for informative and detailed data, and (iv) efficiency for low-latency processing. Relevance is determined by the probabilistic alignment of incoming data with downstream tasks, while specificity employs the distance to a root embedding representing the least specific data as an efficient proxy for informativeness. By establishing reference points from target task data, ReSpec filters incoming data in real-time, eliminating the need for extensive storage and compute. Evaluating on large-scale datasets WebVid2M and VideoCC3M, ReSpec attains state-of-the-art performance on five zero-shot video retrieval tasks, using as little as 5% of the data while incurring minimal compute. The source code is available at https://github.com/cdjkim/ReSpec.
Chris Dongjoo Kim, Jihwan Moon 0002, Sangwoo Moon 0001, Heeseung Yun, Sihaeng Lee, Aniruddha Kembhavi, Soonyoung Lee, Gunhee Kim, Sangho Lee 0008
CVPR7
2025 Single-stage table structure recognition approach via efficient sequence modeling
Yeonsik Jo, Soonyoung Lee, Nam Ik Cho
Pattern Recognit. Lett.3
2024 Bi-directional Contextual Attention for 3D Dense Captioning
Minjung Kim 0001, Hyung Suk Lim, Soonyoung Lee, Bumsoo Kim 0005, Gunhee Kim
ECCV (18)3
2022 L-Verse: Bidirectional Generation Between Image and Text
abstract
Far beyond learning long-range interactions of natural language, transformers are becoming the de-facto standard for many vision tasks with their power and scalability. Especially with cross-modal tasks between image and text, vector quantized variational autoencoders (VQ-VAEs) are widely used to make a raw RGB image into a sequence of feature vectors. To better leverage the correlation between image and text, we propose L-Verse, a novel architecture consisting of feature-augmented variational autoencoder (AugVAE) and bidirectional auto-regressive transformer (BiART) for image-to-text and text-to-image generation. Our AugVAE shows the state-of-the-art reconstruction performance on ImageNetlK validation set, along with the robustness to unseen images in the wild. Unlike other models, BiART can distinguish between image (or text) as a conditional reference and a generation target. L-Verse can be directly used for image-to-text or text-to-image generation without any finetuning or extra object detection framework. In quantitative and qualitative experiments, L-Verse shows impressive results against previous methods in both image-to-text and text-to-image generation on MS-COCO Captions. We furthermore assess the scalability of L-Verse architecture on Conceptual Captions and present the initial result of bidirectional vision-language representation learning on general domain.
Gwangmo Song, Sihaeng Lee, Yewon Seo, Soonyoung Lee, Honglak Lee, Kyunghoon Bae
CVPR6
2022 CEDe: A collection of expert-curated datasets with atom-level entity annotations for Optical Chemical Structure Recognition
abstract
Optical Chemical Structure Recognition (OCSR) deals with the translation from chemical images to molecular structures, this being the main way chemical compounds are depicted in scientific documents. Traditionally, rule-based methods have followed a framework based on the detection of chemical entities, such as atoms and bonds, followed by a compound structure reconstruction step. Recently, neural architectures analog to image captioning have been explored to solve this task, yet they still show to be data inefficient, using millions of examples just to show performances comparable with traditional methods. Looking to motivate and benchmark new approaches based on atomic-level entities detection and graph reconstruction, we present CEDe, a unique collection of chemical entity bounding boxes manually curated by experts for scientific literature datasets. These annotations combine to more than 700,000 chemical entity bounding boxes with the necessary information for structure reconstruction. Also, a large synthetic dataset containing one million molecular images and annotations is released in order to explore transfer-learning techniques that could help these architectures perform better under low-data regimes. Benchmarks show that detection-reconstruction based models can achieve performances on par with or better than image captioning-like models, even with 100x fewer training examples.
Rodrigo Hormazabal, Changyoung Park, Soonyoung Lee, Sehui Han, Yeonsik Jo, Jaewan Lee, Ahra Jo, Jaegul Choo, Moontae Lee, Honglak Lee
NeurIPS3
2014 An Efficient Multiple Cell Upsets Tolerant Content-Addressable Memory
abstract
Multiple cell upsets (MCUs) become more and more problematic as the size of technology reaches or goes below 65 nm. The percentage of MCUs is reported significantly larger than that of single cell upsets (SCUs) in 20 nm technology. In SRAM and DRAM, MCUs are tackled by incorporating single-error correcting double-error detecting (SEC-DED) code and interleaved data columns. However, in content-addressable memory (CAM), column interleaving is not practically possible. A novel error correction code (ECC) scheme is proposed in this paper that will cater for ever-increasing MCUs. This work demonstrated that m parity bits are sufficient to cater for up to m-bit MCUs, with an understanding of the physical grouping of MCUs. The results showed that the proposed scheme requires 85% fewer parity bits compared to traditional Hamming distance based schemes.
Syed Mohsin Abbas, Soonyoung Lee, Sanghyeon Baeg, Sungju Park
IEEE Trans. Computers2
2012 Soft Error Issues with Scaling Technologies
abstract
As transistor geometry shrinks, the erroneous and spurious charge from a particle strike tends to be shared by multiple nodes and causes multiple nodes upset. Such SEU mechanism invalidates the hardening principle of protecting a single node in relatively larger technologies. SEU needs to be accordingly understood and evaluated. In an 28-nm design example, SEU can happen in 6-day interval if no mitigation technique is used.
Sanghyeon Baeg, Jongsun Bae, Soonyoung Lee, Chul Seung Lim, Sang Hoon Jeon, Hyeonwoo Nam
Asian Test Symposium3
2008 New smith predictor control using disturbance observer for steam superheater and steam pressure of the boiler
abstract
The steam superheater and the steam pressure systems are time delay systems that have poles near the origin in the left half plane or a pure integrator. Smith predictor can't be applied to these systems any more, because of the steady state error for the step disturbance. In this paper, a new Smith predictor controller is proposed for the steam superheater and the steam pressure. A disturbance observer to estimate an input disturbance and a new controller to eliminate an effect of a disturbance are proposed. The simulation results verifies the efficiency of the proposed system.
Soonyoung Lee, Yonghwan Shin, Eunsung Jang, Hwi-Beom Shin
ICARCV1