Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Hyunwoo Yu

dblp:183/7958 · DBLP profile ↗
← Back
8ranked-venue papers
3as first author
6since 2021 · last 2024
0009-0009-4426-8272ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Security and privacy · 2 · 1 first-author · 1 since 2021Computer networks · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Segmentation and scene understanding · 41% Deep learning architectures and training · 30% Efficient and distributed learning · 14%
Network and information security
3 papers
Privacy and data protection · 74% Digital forensics and information hiding · 21% Systems and software security · 5%

Topics — the 16 heaviest of 17, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Segmentation and scene understanding
semantic segmentation
1.422024
Embedding-Free Transformer with Inference Spatial Reduction for Efficient Semantic Segmentation · ECCV (42) 2024
FeedFormer: Revisiting Transformer Decoder for Efficient Semantic Segmentation · AAAI 2023
Computer vision › Vision and language
cross-modal alignment
0.812024
Cross-Aware Early Fusion With Stage-Divided Vision and Language Transformer Encoders for Referring Image Segmentation · IEEE Trans. Multim. 2024
Machine learning › Efficient and distributed learning › inference efficiency
efficient transformer inference
0.812024
Embedding-Free Transformer with Inference Spatial Reduction for Efficient Semantic Segmentation · ECCV (42) 2024
Computer vision › Segmentation and scene understanding
referring image segmentation
0.812024
Cross-Aware Early Fusion With Stage-Divided Vision and Language Transformer Encoders for Referring Image Segmentation · IEEE Trans. Multim. 2024
Machine learning › Deep learning architectures and training
transformer
0.812024
Cross-Aware Early Fusion With Stage-Divided Vision and Language Transformer Encoders for Referring Image Segmentation · IEEE Trans. Multim. 2024
Machine learning › Deep learning architectures and training › transformer
transformer decoder
0.712023
FeedFormer: Revisiting Transformer Decoder for Efficient Semantic Segmentation · AAAI 2023
Privacy and data protection › data confidentiality › content privacy › multimedia privacy
video privacy
0.622018
Pinto: Enabling Video Privacy for Commodity IoT Cameras · CCS 2018
ViewMap: Sharing Private In-Vehicle Dashcam Videos · NSDI 2017
Privacy and data protection › data confidentiality › content privacy
audio privacy
0.612022
Overo: Sharing Private Audio Recordings · CCS 2022
Privacy and data protection
privacy-preserving data sharing
0.612022
Overo: Sharing Private Audio Recordings · CCS 2022
Digital forensics and information hiding
watermarking
0.612022
Overo: Sharing Private Audio Recordings · CCS 2022
Privacy and data protection
image privacy
0.312018
Pinto: Enabling Video Privacy for Commodity IoT Cameras · CCS 2018
Privacy and data protection › privacy-preserving computation › privacy-preserving multimedia processing
privacy-preserving video
0.312018
Pinto: Enabling Video Privacy for Commodity IoT Cameras · CCS 2018
Machine learning › Deep learning architectures and training › transformer
vision transformer
0.212023
FeedFormer: Revisiting Transformer Decoder for Efficient Semantic Segmentation · AAAI 2023
Systems and software security
data integrity
0.212022
Overo: Sharing Private Audio Recordings · CCS 2022
Digital forensics and information hiding › content authentication
video authentication
0.112018
Pinto: Enabling Video Privacy for Commodity IoT Cameras · CCS 2018
Vehicular, aerial and satellite networks
vehicular networks
0.112017
ViewMap: Sharing Private In-Vehicle Dashcam Videos · NSDI 2017

Methods — techniques the papers use, named apart from their topics

transformer · 1.4stage-divided encoders · 0.8spatial reduction attention · 0.8feature-based alignment · 0.8attention mechanism · 0.7voice conversion · 0.6redaction · 0.6realtime signature · 0.3post-processing blurring · 0.3
YearPublicationVenuePosition
2024 Embedding-Free Transformer with Inference Spatial Reduction for Efficient Semantic Segmentation
Hyunwoo Yu, Yubin Cho, Beoungwoo Kang, Seunghun Moon, Kyeongbo Kong, Suk-Ju Kang
ECCV (42)1
2024 MetaSeg: MetaFormer-based Global Contexts-aware Network for Efficient Semantic Segmentation
abstract
Beyond the Transformer, it is important to explore how to exploit the capacity of the MetaFormer, an architecture that is fundamental to the performance improvements of the Transformer. Previous studies have exploited it only for the backbone network. Unlike previous studies, we explore the capacity of the Metaformer architecture more extensively in the semantic segmentation task. We propose a powerful semantic segmentation network, MetaSeg, which leverages the Metaformer architecture from the backbone to the decoder. Our MetaSeg shows that the MetaFormer architecture plays a significant role in capturing the useful contexts for the decoder as well as for the backbone. In addition, recent segmentation methods have shown that using a CNN-based backbone for extracting the spatial information and a decoder for extracting the global information is more effective than using a transformer-based backbone with a CNN-based decoder. This motivates us to adopt the CNN-based backbone using the MetaFormer block and design our MetaFormer-based decoder, which consists of a novel self-attention module to capture the global contexts. To consider both the global contexts extraction and the computational efficiency of the self-attention for semantic segmentation, we propose a Channel Reduction Attention (CRA) module that reduces the channel dimension of the query and key into the one dimension. In this way, our proposed MetaSeg outperforms the previous state-of-the-art methods with more efficient computational costs on popular semantic segmentation and a medical image segmentation benchmark, including ADE20K, Cityscapes, COCO-stuff, and Synapse.
Beoungwoo Kang, Seunghun Moon, Yubin Cho, Hyunwoo Yu, Suk-Ju Kang
WACV4
2024 Cross-Aware Early Fusion With Stage-Divided Vision and Language Transformer Encoders for Referring Image Segmentation
abstract
Referring segmentation aims to segment a target object related to a natural language expression. Key challenges of this task are understanding the meaning of complex and ambiguous language expressions and determining the relevant regions in the image with multiple objects by referring to the expression. Recent models have focused on the early fusion with the language features at the intermediate stage of the vision encoder, but these approaches have a limitation that the language features cannot refer to the visual information. To address this issue, this paper proposes a novel architecture, Cross-aware early fusion with stage-divided Vision and Language Transformer encoders (CrossVLT), which allows both language and vision encoders to perform the early fusion for improving the ability of the cross-modal context modeling. Unlike previous methods, our method enables the vision and language features to refer to each other's information at each stage to mutually enhance the robustness of both encoders. Furthermore, unlike the conventional scheme that relies solely on the high-level features for the cross-modal alignment, we introduce a feature-based alignment scheme that enables the low-level to high-level features of the vision and language encoders to engage in the cross-modal alignment. By aligning the intermediate cross-modal features in all encoder stages, this scheme leads to effective cross-modal fusion. In this way, the proposed approach is simple but effective for referring image segmentation, and it outperforms the previous state-of-the-art methods on three public benchmarks.
Yubin Cho, Hyunwoo Yu, Suk-Ju Kang
IEEE Trans. Multim.2
2023 FeedFormer: Revisiting Transformer Decoder for Efficient Semantic Segmentation
abstract
With the success of Vision Transformer (ViT) in image classification, its variants have yielded great success in many downstream vision tasks. Among those, the semantic segmentation task has also benefited greatly from the advance of ViT variants. However, most studies of the transformer for semantic segmentation only focus on designing efficient transformer encoders, rarely giving attention to designing the decoder. Several studies make attempts in using the transformer decoder as the segmentation decoder with class-wise learnable query. Instead, we aim to directly use the encoder features as the queries. This paper proposes the Feature Enhancing Decoder transFormer (FeedFormer) that enhances structural information using the transformer decoder. Our goal is to decode the high-level encoder features using the lowest-level encoder feature. We do this by formulating high-level features as queries, and the lowest-level feature as the key and value. This enhances the high-level features by collecting the structural information from the lowest-level feature. Additionally, we use a simple reformation trick of pushing the encoder blocks to take the place of the existing self-attention module of the decoder to improve efficiency. We show the superiority of our decoder with various light-weight transformer-based decoders on popular semantic segmentation datasets. Despite the minute computation, our model has achieved state-of-the-art performance in the performance computation trade-off. Our model FeedFormer-B0 surpasses SegFormer-B0 with 1.8% higher mIoU and 7.1% less computation on ADE20K, and 1.7% higher mIoU and 14.4% less computation on Cityscapes, respectively. Code will be released at: https://github.com/jhshim1995/FeedFormer.
Jae-hun Shim, Hyunwoo Yu, Kyeongbo Kong, Suk-Ju Kang
AAAI2
2022 Overo: Sharing Private Audio Recordings
abstract
The use of smartphones as voice recorders has made it easy to record audios as proof of conversations, but sharing of such audio evidence incurs speech and voice privacy risks. However, protecting speech/voice privacy without losing audio authenticity is challenging. The conventional post-process redaction and voice conversion of audio recordings, which invalidate their original signatures, make the audio unverifiable and prone to tampering. In this paper, we present Overo, an audio recording/sharing solution that supports privacy processing without losing audio authenticity. Overo records a realtime audio stream in the standard AAC-encoded format and allows privacy post-processing prior to sharing of audios while keeping their original signatures valid (even after the post redaction and voice conversion), guaranteeing no tampering since the time of their recording. Therefore, users can post-process their recordings to desired levels of privacy on speech (what content to redact) and speakers (whose voice to disguise) at the time of audio release, and still prove their authenticity. Overo is readily implementable in today's commodity smartphones. Our prototype on iPhones/Android phones demonstrates the production of AAC-compliant, tamperproof, and self-authenticating audios with speech/voice privacy protected based on users' post-recording decisions.
Jaemin Lim, Kiyeon Kim, Hyunwoo Yu, Suk-Bok Lee
CCS3
2022 Vision Transformer-Based Retina Vessel Segmentation with Deep Adaptive Gamma Correction
abstract
Accurate segmentation of the retina vessel is essential for the early diagnosis of eye-related diseases. Recently, convolutional neural networks have shown remarkable performance in retina vessel segmentation. However, the complexity of edge structural information and the changeable intensity distribution depending on retina images reduce the performance of the segmentation tasks. This paper proposes two novel deep learning-based modules, channel attention vision transformer (CAViT) and deep adaptive gamma correction (DAGC), to tackle these issues. The CAViT jointly applies the efficient channel attention (ECA) and the vision transformer (ViT), in which the channel attention module considers the interdependency among feature channels and the ViT discriminates meaningful edge structures by considering the global context. The DAGC module provides the optimal gamma correction value for each input image by jointly training a CNN model with the segmentation network so that all the retina images are mapped to a unified intensity distribution. The experimental results show that our proposed method achieves superior performance compared to conventional methods on widely used datasets, DRIVE and CHASE DB1.
Hyunwoo Yu, Jae-hun Shim, Jaeho Kwak, Jou Won Song, Suk-Ju Kang
ICASSP1
2018 Pinto: Enabling Video Privacy for Commodity IoT Cameras
abstract
With various IoT cameras today, sharing of their video evidences, while benefiting the public, threatens the privacy of individuals in the footage. However, protecting visual privacy without losing video authenticity is challenging. The conventional post-process blurring would open the door for posterior fabrication, whereas the realtime blurring results in poor quality, low-frame-rate videos due to the limited processing power of commodity cameras. This paper presents Pinto, a software-based solution for producing privacy-protected, forgery-proof, and high-frame-rate videos using low-end IoT cameras. Pinto records a realtime video stream at a fast rate and allows post-processing for privacy protection prior to sharing of videos while keeping their original, realtime signatures valid even after the post blurring, guaranteeing no content forgery since the time of their recording. Pinto is readily implementable in today's commodity cameras. Our prototype on three different embedded devices, each deployed in a specific application context---on-site, vehicular, and aerial surveillance---demonstrates the production of privacy-protected, forgery-proof videos with frame rates of 17--24 fps, comparable to those of HD videos.
Hyunwoo Yu, Jaemin Lim, Kiyeon Kim, Suk-Bok Lee
CCS1
2017 ViewMap: Sharing Private In-Vehicle Dashcam Videos
Jaemin Lim, Hyunwoo Yu, Kiyeon Kim, Suk-Bok Lee
NSDI3