Shinyeong Noh

dblp:317/5434 · DBLP profile ↗
← Back
2ranked-venue papers
0as first author
2since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
1 paper
Visual content generation and editing · 87% Computational photography and imaging · 13%
Artificial intelligence
1 paper
Representation and self-supervised learning · 61% Vision and language · 30% Video understanding and tracking · 9%

Topics — the 7 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Visual content generation and editing › image generation
controllable image generation
1.012026
RatioMorph: Controllable Diffusion Framework for Automotive Viewpoint and Proportion Manipulation in Vehicle Design · AAAI 2026
Visual content generation and editing › image generation
diffusion-based image generation
1.012026
RatioMorph: Controllable Diffusion Framework for Automotive Viewpoint and Proportion Manipulation in Vehicle Design · AAAI 2026
Machine learning › Representation and self-supervised learning
contrastive learning
0.612022
Video-Text Representation Learning via Differentiable Weak Temporal Alignment · CVPR 2022
Machine learning › Representation and self-supervised learning › contrastive learning
multimodal contrastive learning
0.612022
Video-Text Representation Learning via Differentiable Weak Temporal Alignment · CVPR 2022
Computer vision › Vision and language › multimodal representation
video-language representation learning
0.612022
Video-Text Representation Learning via Differentiable Weak Temporal Alignment · CVPR 2022
Computational photography and imaging
depth estimation
0.312026
RatioMorph: Controllable Diffusion Framework for Automotive Viewpoint and Proportion Manipulation in Vehicle Design · AAAI 2026
Computer vision › Video understanding and tracking
temporal alignment
0.212022
Video-Text Representation Learning via Differentiable Weak Temporal Alignment · CVPR 2022

Methods — techniques the papers use, named apart from their topics

diffusion model · 1.0depth estimation · 1.0dynamic time warping · 0.6differentiable DTW · 0.6contrastive learning · 0.6
YearPublicationVenuePosition
2026 RatioMorph: Controllable Diffusion Framework for Automotive Viewpoint and Proportion Manipulation in Vehicle Design
abstract
Designing vehicle exteriors requires repeated refinement of key proportions and viewpoints, a process traditionally reliant on manual sketching, which is often time-consuming and inefficient in early concept stages. To accelerate the design process, we are exploring the potential of utilizing AI for ideation in these early stages. However, it remains a challenging task to control proportions and maintain a fixed perspective when generating images using AI. To address these limitations, we present RatioMorph, a controllable image generation system that enables manipulation of vehicle proportions and viewpoints when generating images by AI. RatioMorph comprises two core modules. Car2BoxNet is a depth estimation model that transforms real photographs into structured box-style depth maps that capture the geometric layout of the vehicle. Box2CarNet is a diffusion-based image generator fine-tuned to produce vehicle designs that adhere to the provided geometric conditions. Both Car2BoxNet and Box2CarNet are trained on a synthetic dataset curated through automated filtering based on geometric alignment and visual quality. Evaluated within a production-adjacent automotive design workflow, RatioMorph significantly reduced early-stage design iteration time and enabled exploratory workflows that were difficult with previous AI workflows. This work introduces a domain-specific, controllable diffusion-based generation system tailored for automotive design, enabling manipulation of vehicle viewpoint and proportion. It demonstrates strong potential to accelerate early-stage workflows and outlines a path toward industrial deployment, with phased integration into production environments currently underway.
Haeji Go, Jae-Hun Lee, Shinyeong Noh, Kayoung Kim, Kyuseong Lim, Jee Eun Song, Joowan Sung, Soonbeom Kwon, Myoungbok Shin, Junsang Park
AAAI3
2022 Video-Text Representation Learning via Differentiable Weak Temporal Alignment
abstract
Learning generic joint representations for video and text by a supervised method requires a prohibitively substantial amount of manually annotated video datasets. As a practical alternative, a large-scale but uncurated and narrated video dataset, HowTo100M, has recently been introduced. But it is still challenging to learn joint embeddings of video and text in a self-supervised manner, due to its ambiguity and non-sequential alignment. In this paper, we propose a novel multi-modal self-supervised framework Video-Text Temporally Weak Alignment-based Contrastive Learning (VT-TWINS) to capture significant information from noisy and weakly correlated data using a variant of Dynamic Time Warping (DTW). We observe that the standard DTW inherently cannot handle weakly correlated data and only considers the globally optimal alignment path. To address these problems, we develop a differentiable DTW which also reflects local information with weak temporal alignment. Moreover, our proposed model applies a contrastive learning scheme to learn feature representations on weakly correlated data. Our extensive experiments demonstrate that VT-TWINS attains significant improvements in multi-modal representation learning and outperforms various challenging downstream tasks. Code is available at https://github.com/mlvlab/VT-Twins.
Dohwan Ko, Joonmyung Choi, Juyeon Ko, Shinyeong Noh, Kyoung-Woon On, Eun-Sol Kim, Hyunwoo J. Kim
CVPR4