EDBT 2026 Demo / reviewers in the wild / expert
Hyunwoo Yu
dblp:183/7958
· DBLP profile ↗
8ranked-venue papers
3as first author
6since 2021 · last 2024
0009-0009-4426-8272ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Security and privacy · 2 · 1 first-author · 1 since 2021Computer networks · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Segmentation and scene understanding · 41% Deep learning architectures and training · 30% Efficient and distributed learning · 14% | |
| Network and information security
3 papers |
Privacy and data protection · 74% Digital forensics and information hiding · 21% Systems and software security · 5% |
Topics — the 16 heaviest of 17, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Segmentation and scene understanding
semantic segmentation |
1.4 | 2 | 2024 | Embedding-Free Transformer with Inference Spatial Reduction for Efficient Semantic Segmentation · ECCV (42) 2024 FeedFormer: Revisiting Transformer Decoder for Efficient Semantic Segmentation · AAAI 2023 |
Computer vision › Vision and language
cross-modal alignment |
0.8 | 1 | 2024 | Cross-Aware Early Fusion With Stage-Divided Vision and Language Transformer Encoders for Referring Image Segmentation · IEEE Trans. Multim. 2024 |
Machine learning › Efficient and distributed learning › inference efficiency
efficient transformer inference |
0.8 | 1 | 2024 | Embedding-Free Transformer with Inference Spatial Reduction for Efficient Semantic Segmentation · ECCV (42) 2024 |
Computer vision › Segmentation and scene understanding
referring image segmentation |
0.8 | 1 | 2024 | Cross-Aware Early Fusion With Stage-Divided Vision and Language Transformer Encoders for Referring Image Segmentation · IEEE Trans. Multim. 2024 |
Machine learning › Deep learning architectures and training
transformer |
0.8 | 1 | 2024 | Cross-Aware Early Fusion With Stage-Divided Vision and Language Transformer Encoders for Referring Image Segmentation · IEEE Trans. Multim. 2024 |
Machine learning › Deep learning architectures and training › transformer
transformer decoder |
0.7 | 1 | 2023 | FeedFormer: Revisiting Transformer Decoder for Efficient Semantic Segmentation · AAAI 2023 |
Privacy and data protection › data confidentiality › content privacy › multimedia privacy
video privacy |
0.6 | 2 | 2018 | Pinto: Enabling Video Privacy for Commodity IoT Cameras · CCS 2018 ViewMap: Sharing Private In-Vehicle Dashcam Videos · NSDI 2017 |
Privacy and data protection › data confidentiality › content privacy
audio privacy |
0.6 | 1 | 2022 | Overo: Sharing Private Audio Recordings · CCS 2022 |
Privacy and data protection
privacy-preserving data sharing |
0.6 | 1 | 2022 | Overo: Sharing Private Audio Recordings · CCS 2022 |
Digital forensics and information hiding
watermarking |
0.6 | 1 | 2022 | Overo: Sharing Private Audio Recordings · CCS 2022 |
Privacy and data protection
image privacy |
0.3 | 1 | 2018 | Pinto: Enabling Video Privacy for Commodity IoT Cameras · CCS 2018 |
Privacy and data protection › privacy-preserving computation › privacy-preserving multimedia processing
privacy-preserving video |
0.3 | 1 | 2018 | Pinto: Enabling Video Privacy for Commodity IoT Cameras · CCS 2018 |
Machine learning › Deep learning architectures and training › transformer
vision transformer |
0.2 | 1 | 2023 | FeedFormer: Revisiting Transformer Decoder for Efficient Semantic Segmentation · AAAI 2023 |
Systems and software security
data integrity |
0.2 | 1 | 2022 | Overo: Sharing Private Audio Recordings · CCS 2022 |
Digital forensics and information hiding › content authentication
video authentication |
0.1 | 1 | 2018 | Pinto: Enabling Video Privacy for Commodity IoT Cameras · CCS 2018 |
Vehicular, aerial and satellite networks
vehicular networks |
0.1 | 1 | 2017 | ViewMap: Sharing Private In-Vehicle Dashcam Videos · NSDI 2017 |
Methods — techniques the papers use, named apart from their topics
transformer · 1.4stage-divided encoders · 0.8spatial reduction attention · 0.8feature-based alignment · 0.8attention mechanism · 0.7voice conversion · 0.6redaction · 0.6realtime signature · 0.3post-processing blurring · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Embedding-Free Transformer with Inference Spatial Reduction for Efficient Semantic Segmentation
Hyunwoo Yu, Yubin Cho, Beoungwoo Kang, Seunghun Moon, Kyeongbo Kong, Suk-Ju Kang |
ECCV (42) | 1 |
| 2024 | MetaSeg: MetaFormer-based Global Contexts-aware Network for Efficient Semantic SegmentationabstractBeyond the Transformer, it is important to explore how to exploit the capacity of the MetaFormer, an architecture that is fundamental to the performance improvements of the Transformer. Previous studies have exploited it only for the backbone network. Unlike previous studies, we explore the capacity of the Metaformer architecture more extensively in the semantic segmentation task. We propose a powerful semantic segmentation network, MetaSeg, which leverages the Metaformer architecture from the backbone to the decoder. Our MetaSeg shows that the MetaFormer architecture plays a significant role in capturing the useful contexts for the decoder as well as for the backbone. In addition, recent segmentation methods have shown that using a CNN-based backbone for extracting the spatial information and a decoder for extracting the global information is more effective than using a transformer-based backbone with a CNN-based decoder. This motivates us to adopt the CNN-based backbone using the MetaFormer block and design our MetaFormer-based decoder, which consists of a novel self-attention module to capture the global contexts. To consider both the global contexts extraction and the computational efficiency of the self-attention for semantic segmentation, we propose a Channel Reduction Attention (CRA) module that reduces the channel dimension of the query and key into the one dimension. In this way, our proposed MetaSeg outperforms the previous state-of-the-art methods with more efficient computational costs on popular semantic segmentation and a medical image segmentation benchmark, including ADE20K, Cityscapes, COCO-stuff, and Synapse. Beoungwoo Kang, Seunghun Moon, Yubin Cho, Hyunwoo Yu, Suk-Ju Kang |
WACV | 4 |
| 2024 | Cross-Aware Early Fusion With Stage-Divided Vision and Language Transformer Encoders for Referring Image SegmentationabstractReferring segmentation aims to segment a target object related to a natural language expression. Key challenges of this task are understanding the meaning of complex and ambiguous language expressions and determining the relevant regions in the image with multiple objects by referring to the expression. Recent models have focused on the early fusion with the language features at the intermediate stage of the vision encoder, but these approaches have a limitation that the language features cannot refer to the visual information. To address this issue, this paper proposes a novel architecture, Cross-aware early fusion with stage-divided Vision and Language Transformer encoders (CrossVLT), which allows both language and vision encoders to perform the early fusion for improving the ability of the cross-modal context modeling. Unlike previous methods, our method enables the vision and language features to refer to each other's information at each stage to mutually enhance the robustness of both encoders. Furthermore, unlike the conventional scheme that relies solely on the high-level features for the cross-modal alignment, we introduce a feature-based alignment scheme that enables the low-level to high-level features of the vision and language encoders to engage in the cross-modal alignment. By aligning the intermediate cross-modal features in all encoder stages, this scheme leads to effective cross-modal fusion. In this way, the proposed approach is simple but effective for referring image segmentation, and it outperforms the previous state-of-the-art methods on three public benchmarks. Yubin Cho, Hyunwoo Yu, Suk-Ju Kang |
IEEE Trans. Multim. | 2 |
| 2023 | FeedFormer: Revisiting Transformer Decoder for Efficient Semantic SegmentationabstractWith the success of Vision Transformer (ViT) in image classification, its variants have yielded great success in many downstream vision tasks. Among those, the semantic segmentation task has also benefited greatly from the advance of ViT variants. However, most studies of the transformer for semantic segmentation only focus on designing efficient transformer encoders, rarely giving attention to designing the decoder. Several studies make attempts in using the transformer decoder as the segmentation decoder with class-wise learnable query. Instead, we aim to directly use the encoder features as the queries. This paper proposes the Feature Enhancing Decoder transFormer (FeedFormer) that enhances structural information using the transformer decoder. Our goal is to decode the high-level encoder features using the lowest-level encoder feature. We do this by formulating high-level features as queries, and the lowest-level feature as the key and value. This enhances the high-level features by collecting the structural information from the lowest-level feature. Additionally, we use a simple reformation trick of pushing the encoder blocks to take the place of the existing self-attention module of the decoder to improve efficiency. We show the superiority of our decoder with various light-weight transformer-based decoders on popular semantic segmentation datasets. Despite the minute computation, our model has achieved state-of-the-art performance in the performance computation trade-off. Our model FeedFormer-B0 surpasses SegFormer-B0 with 1.8% higher mIoU and 7.1% less computation on ADE20K, and 1.7% higher mIoU and 14.4% less computation on Cityscapes, respectively. Code will be released at: https://github.com/jhshim1995/FeedFormer. Jae-hun Shim, Hyunwoo Yu, Kyeongbo Kong, Suk-Ju Kang |
AAAI | 2 |
| 2022 | Overo: Sharing Private Audio RecordingsabstractThe use of smartphones as voice recorders has made it easy to record audios as proof of conversations, but sharing of such audio evidence incurs speech and voice privacy risks. However, protecting speech/voice privacy without losing audio authenticity is challenging. The conventional post-process redaction and voice conversion of audio recordings, which invalidate their original signatures, make the audio unverifiable and prone to tampering. In this paper, we present Overo, an audio recording/sharing solution that supports privacy processing without losing audio authenticity. Overo records a realtime audio stream in the standard AAC-encoded format and allows privacy post-processing prior to sharing of audios while keeping their original signatures valid (even after the post redaction and voice conversion), guaranteeing no tampering since the time of their recording. Therefore, users can post-process their recordings to desired levels of privacy on speech (what content to redact) and speakers (whose voice to disguise) at the time of audio release, and still prove their authenticity. Overo is readily implementable in today's commodity smartphones. Our prototype on iPhones/Android phones demonstrates the production of AAC-compliant, tamperproof, and self-authenticating audios with speech/voice privacy protected based on users' post-recording decisions. Jaemin Lim, Kiyeon Kim, Hyunwoo Yu, Suk-Bok Lee |
CCS | 3 |
| 2022 | Vision Transformer-Based Retina Vessel Segmentation with Deep Adaptive Gamma CorrectionabstractAccurate segmentation of the retina vessel is essential for the early diagnosis of eye-related diseases. Recently, convolutional neural networks have shown remarkable performance in retina vessel segmentation. However, the complexity of edge structural information and the changeable intensity distribution depending on retina images reduce the performance of the segmentation tasks. This paper proposes two novel deep learning-based modules, channel attention vision transformer (CAViT) and deep adaptive gamma correction (DAGC), to tackle these issues. The CAViT jointly applies the efficient channel attention (ECA) and the vision transformer (ViT), in which the channel attention module considers the interdependency among feature channels and the ViT discriminates meaningful edge structures by considering the global context. The DAGC module provides the optimal gamma correction value for each input image by jointly training a CNN model with the segmentation network so that all the retina images are mapped to a unified intensity distribution. The experimental results show that our proposed method achieves superior performance compared to conventional methods on widely used datasets, DRIVE and CHASE DB1. Hyunwoo Yu, Jae-hun Shim, Jaeho Kwak, Jou Won Song, Suk-Ju Kang |
ICASSP | 1 |
| 2018 | Pinto: Enabling Video Privacy for Commodity IoT CamerasabstractWith various IoT cameras today, sharing of their video evidences, while benefiting the public, threatens the privacy of individuals in the footage. However, protecting visual privacy without losing video authenticity is challenging. The conventional post-process blurring would open the door for posterior fabrication, whereas the realtime blurring results in poor quality, low-frame-rate videos due to the limited processing power of commodity cameras. This paper presents Pinto, a software-based solution for producing privacy-protected, forgery-proof, and high-frame-rate videos using low-end IoT cameras. Pinto records a realtime video stream at a fast rate and allows post-processing for privacy protection prior to sharing of videos while keeping their original, realtime signatures valid even after the post blurring, guaranteeing no content forgery since the time of their recording. Pinto is readily implementable in today's commodity cameras. Our prototype on three different embedded devices, each deployed in a specific application context---on-site, vehicular, and aerial surveillance---demonstrates the production of privacy-protected, forgery-proof videos with frame rates of 17--24 fps, comparable to those of HD videos. Hyunwoo Yu, Jaemin Lim, Kiyeon Kim, Suk-Bok Lee |
CCS | 1 |
| 2017 | ViewMap: Sharing Private In-Vehicle Dashcam Videos
Jaemin Lim, Hyunwoo Yu, Kiyeon Kim, Suk-Bok Lee |
NSDI | 3 |