VLDB 2026 Research / reviewers in the wild / expert
Hijung Shin
dblp:41/9388 · also Hijung Valentina Shin
· DBLP profile ↗
20ranked-venue papers
4as first author
10since 2021 · last 2026
0000-0001-8798-4580ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 16 · 2 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 3 first-author · 1 since 2021Artificial intelligence and machine learning · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Visual Lyrics: Generating Animated Text for Music Lyric Videos with an Augmented Text EditorabstractAnimated lyric videos transform song lyrics into dynamic visual experiences, offering a powerful medium for artistic expression and audience engagement. However, creating these videos is challenging, requiring expertise in audio, typography, graphic design, and animation, making it inaccessible to novices. To address this challenge, we introduce Visual Lyrics, a proof-of-concept system for generating animated lyric videos controlled with an augmented text editor interface. We examined existing lyric videos to distill a taxonomy and design guidelines, informing the design of Visual Lyrics. Our key insight is a multimodal music analysis pipeline based on the taxonomy and leveraging LLM’s strong natural language understanding and code generation capabilities to synthesize creative and semantically meaningful animations. We collected a dataset of over 300 code-driven creative text animations to serve as inspiration for our LLM-driven pipeline, which we open source. In a user study, Visual Lyrics enabled novices to easily create high-quality animated lyric videos with high ratings of enjoyment, inspiration, and exploration. David Chuan-En Lin, Cuong Nguyen 0003, Hijung Shin, Nikolas Martelaro |
IUI | 3 |
| 2025 | Compositional Structures as Substrates for Human-AI Co-creation Environment: A Design Approach and A Case Study
Yining Cao, Yiyi Huang, Anh Truong, Hijung Shin, Haijun Xia |
CHI | 4 |
| 2025 | VideoDiff: Human-AI Video Co-Creation with AlternativesabstractTo make an engaging video, people sequence interesting moments and add visuals such as B-rolls or text. While video editing requires time and effort, AI has recently shown strong potential to make editing easier through suggestions and automation. A key strength of generative models is their ability to quickly generate multiple variations, but when provided with many alternatives, creators struggle to compare them to find the best fit. We propose VideoDiff, an AI video editing tool designed for editing with alternatives. With VideoDiff, creators can generate and review multiple AI recommendations for each editing process: creating a rough cut, inserting B-rolls, and adding text effects. VideoDiff simplifies comparisons by aligning videos and highlighting differences through timelines, transcripts, and video previews. Creators have the flexibility to regenerate and refine AI suggestions as they compare alternatives. Our study participants (N=12) could easily compare and customize alternatives, creating more satisfying results. Mina Huh, Kim Pimmel, Hijung Shin, Amy Pavel, Mira Dontcheva |
CHI | 4 |
| 2023 | SoundToons: Exemplar-Based Authoring of Interactive Audio-Driven Animation SpritesabstractAnimations can come to life when they are synchronized with relevant sounds. Yet, synchronizing animations to audio requires tedious key-framing or programming, which is difficult for novice creators. There are existing tools that support audio-driven live animation, but they focus primarily on speech and have little or no support for non-speech sounds. We present SoundToons, an exemplar-based authoring tool for interactive, audio-driven animation focusing on non-speech sounds. Our tool enables novice creators to author live animations to a wide variety of non-speech sounds, such as clapping and instrumental music. We support two types of audio interactions: (1) discrete interaction, which triggers animations when a discrete sound event is detected, and (2) continuous, which synchronizes an animation to continuous audio parameters. By employing an exemplar-based iterative authoring approach, we empower novice creators to design and quickly refine interactive animations. User evaluations demonstrate that novice users can author and perform live audio-driven animation intuitively. Moreover, compared to other input modalities such as trackpads or foot pedals, users preferred using audio as an intuitive way to drive animation. Toby Chong, Hijung Shin, Deepali Aneja, Takeo Igarashi |
IUI | 2 |
| 2023 | Automated Conversion of Music Videos into Lyric VideosabstractMusicians and fans often produce lyric videos, a form of music videos that showcase the song’s lyrics, for their favorite songs. However, making such videos can be challenging and time-consuming as the lyrics need to be added in synchrony and visual harmony with the video. Informed by prior work and close examination of existing lyric videos, we propose a set of design guidelines to help creators make such videos. Our guidelines ensure the readability of the lyric text while maintaining a unified focus of attention. We instantiate these guidelines in a fully automated pipeline that converts an input music video into a lyric video. We demonstrate the robustness of our pipeline by generating lyric videos from a diverse range of input sources. A user study shows that lyric videos generated by our pipeline are effective in maintaining text readability and unifying the focus of attention. Jiaju Ma, Anyi Rao, Li-Yi Wei, Rubaiat Habib Kazi, Hijung Shin, Maneesh Agrawala |
UIST | 5 |
| 2022 | Beyond Subtitles: Captioning and Visualizing Non-speech Sounds to Improve Accessibility of User-Generated VideosabstractCaptioning provides access to sounds in audio-visual content for people who are Deaf or Hard-of-hearing (DHH). As user-generated content in online videos grows in prevalence, researchers have explored using automatic speech recognition (ASR) to automate captioning. However, definitions of captions (as compared to subtitles) include non-speech sounds, which ASR typically does not capture as it focuses on speech. Thus, we explore DHH viewers’ and hearing video creators’ perspectives on captioning non-speech sounds in user-generated online videos using text or graphics. Formative interviews with 11 DHH participants informed the design and implementation of a prototype interface for authoring text-based and graphic captions using automatic sound event detection, which was then evaluated with 10 hearing video creators. Our findings include identifying DHH viewers’ interests in having important non-speech sounds included in captions, as well as various criteria for sound selection and the appropriateness of text-based versus graphic captions of non-speech sounds. Our findings also include hearing creators’ requirements for automatic tools to assist them in captioning non-speech sounds. Oliver Alonzo, Hijung Shin, Dingzeyu Li |
ASSETS | 2 |
| 2022 | CatchLive: Real-time Summarization of Live Streams with Stream Content and Interaction DataabstractLive streams usually last several hours with many viewers joining in the middle. Viewers who join in the middle often want to understand what has happened in the stream. However, catching up with the earlier parts is challenging because it is difficult to know which parts are important in the long, unedited stream while also keeping up with the ongoing stream. We present CatchLive, a system that provides a real-time summary of ongoing live streams by utilizing both the stream content and user interaction data. CatchLive provides viewers with an overview of the stream along with summaries of highlight moments with multiple levels of detail in a readable format. Results from deployments of three streams with 67 viewers show that CatchLive helps viewers grasp the overview of the stream, identify important moments, and stay engaged. Our findings provide insights into designing summarizations of live streams reflecting their characteristics. Saelyne Yang, Jisu Yim, Juho Kim 0001, Hijung Shin |
CHI | 4 |
| 2021 | Beyond Show of Hands: Engaging Viewers via Expressive and Scalable Visual Communication in Live StreamingabstractLive streaming is gaining popularity across diverse application domains in recent years. A core part of the experience is streamer-viewer interaction, which has been mainly text-based. Recent systems explored extending viewer interaction to include visual elements with richer expression and increased engagement. However, understanding expressive visual inputs becomes challenging with many viewers, primarily due to the relative lack of structure in visual input. On the other hand, adding rigid structures can limit viewer interactions to narrow use cases or decrease the expressiveness of viewer inputs. To facilitate the sensemaking of many visual inputs while retaining the expressiveness or versatility of viewer interactions, we introduce a visual input management framework (VIMF) and a system, VisPoll, that help streamers specify, aggregate, and visualize many visual inputs. A pilot evaluation indicated that VisPoll can expand the types of viewer interactions. Our framework provides insights for designing scalable and expressive visual communication for live streaming. John Joon Young Chung, Hijung Shin, Haijun Xia, Li-Yi Wei, Rubaiat Habib Kazi |
CHI | 2 |
| 2021 | Multi-level Correspondence via Graph Kernels for Editing Vector Graphics Designs
Hijung Shin, Jeremy Warner, Björn Hartmann, Celso Gomes, Holger Winnemöller, Wilmot Li |
Graphics Interface | 1 |
| 2021 | SymbolFinder: Brainstorming Diverse Symbols Using Local Semantic NetworksabstractVisual symbols are the building blocks for visual communication. They convey abstract concepts like reform and participation quickly and effectively. When creating graphics with symbols, novice designers often struggle to brainstorm multiple, diverse symbols because they fixate on a few associations instead of broadly exploring different aspects of the concept. We present SymbolFinder, an interactive tool for finding visual symbols for abstract concepts. SymbolFinder molds symbol-finding into a recognition rather than recall task by introducing the user to diverse clusters of words associated with the concept. Users can dive into these clusters to find related, concrete objects that symbolize the concept. We evaluate SymbolFinder with two studies: a comparative user study, demonstrating that SymbolFinder helps novices find more unique symbols for abstract concepts with significantly less effort than a popular image database and a case study demonstrating how SymbolFinder helped design students create visual metaphors for three cover illustrations of news articles. Savvas Petridis, Hijung Shin, Lydia B. Chilton |
UIST | 2 |
| 2020 | Temporal Segmentation of Creative Live StreamsabstractMany artists broadcast their creative process through live streaming platforms like Twitch and YouTube, and people often watch archives of these broadcasts later for learning and inspiration. Unfortunately, because live stream videos are often multiple hours long and hard to skim and browse, few can leverage the wealth of knowledge hidden in these archives. We present an approach for automatic temporal segmentation of creative live stream videos. Using an audio transcript and a log of software usage, the system segments the video into sections that the artist can optionally label with meaningful titles. We evaluate this approach by gathering feedback from expert streamers and comparing automatic segmentations to those made by viewers. We find that, while there is no one "correct" way to segment a live stream, our automatic method performs similarly to viewers, and streamers find it useful for navigating their streams after making slight adjustments and adding section titles. C. Ailie Fraser, Joy Kim, Hijung Shin, Joel Brandt, Mira Dontcheva |
CHI | 3 |
| 2020 | Generating Audio-Visual Slideshows from Text Articles Using Word ConcretenessabstractWe present a system that automatically transforms text articles into audio-visual slideshows by leveraging the notion of word concreteness, which measures how strongly a word or phrase is related to some perceptible concept. In a formative study we learn that people not only prefer such audio-visual slideshows but find that the content is easier to understand compared to text articles or text articles augmented with images. We use word concreteness to select search terms and find images relevant to the text. Then, based on the distribution of concrete words and the grammatical structure of an article, we time-align selected images with audio narration obtained through text-to-speech to produce audio-visual slideshows. In a user evaluation we find that our concreteness-based algorithm selects images that are highly relevant to the text. The quality of our slideshows is comparable to slideshows produced manually using standard video editing tools, and people strongly prefer our slideshows to those generated using a simple keyword-search based approach. Mackenzie Leake, Hijung Shin, Joy Kim, Maneesh Agrawala |
CHI | 2 |
| 2020 | Snapstream: Snapshot-based Interaction in Live Streaming for Visual ArtabstractLive streaming visual art such as drawing or using design software is gaining popularity. An important aspect of live streams is the direct and real-time communication between streamers and viewers. However, currently available text-based interaction limits the expressiveness of viewers as well as streamers, especially when they refer to specific moments or objects in the stream. To investigate the feasibility of using snapshots of streamed content as a way to enhance streamer-viewer interaction, we introduce Snapstream, a system that allows users to take snapshots of the live stream, annotate them, and share the annotated snapshots in the chat. Streamers can also verbally reference a specific snapshot during streaming to respond to viewers' questions or comments. Results from live deployments show that participants communicate more expressively and clearly with increased engagement using Snapstream. Participants used snapshots to reference part of the artwork, give suggestions on it, make fun images or memes, and log intermediate milestones. Our findings suggest that visual interaction enables richer experiences in live streaming. Saelyne Yang, Changyoon Lee, Hijung Shin, Juho Kim 0001 |
CHI | 3 |
| 2020 | Pose2Pose: pose selection and transfer for 2D character animationabstractAn artist faces two challenges when creating a 2D animated character to mimic a specific human performance. First, the artist must design and draw a collection of artwork depicting portions of the character in a suitable set of poses, for example arm and hand poses that can be selected and combined to express the range of gestures typical for that person. Next, to depict a specific performance, the artist must select and position the appropriate set of artwork at each moment of the animation. This paper presents a system that addresses these challenges by leveraging video of the target human performer. Our system tracks arm and hand poses in an example video of the target. The UI displays clusters of these poses to help artists select representative poses that capture the actor's style and personality. From this mapping of pose data to character artwork, our system can generate an animation from a new performance video. It relies on a dynamic programming algorithm to optimize for smooth animations that match the poses found in the video. Artists used our system to create four 2D characters and were pleased with the final automatically animated results. We also describe additional applications addressing audio-driven or text-based animations. Nora S. Willett, Hijung Shin, Zeyu Jin, Wilmot Li, Adam Finkelstein |
IUI | 2 |
| 2019 | B-Script: Transcript-based B-roll Video Editing with RecommendationsabstractIn video production, inserting B-roll is a widely used technique to enrich the story and make a video more engaging. However, determining the right content and positions of B-roll and actually inserting it within the main footage can be challenging, and novice producers often struggle to get both timing and content right. We present B-Script, a system that supports B-roll video editing via interactive transcripts. B-Script has a built-in recommendation system trained on expert-annotated data, recommending users B-roll position and content. To evaluate the system, we conducted a within-subject user study with 110 participants, and compared three interface variations: a timeline-based editor, a transcript-based editor, and a transcript-based editor with recommendations. Users found it easier and were faster to insert B-roll using the transcript-based interface, and they created more engaging videos when recommendations were provided. Bernd Huber, Hijung Shin, Bryan C. Russell, Oliver Wang, Gautham J. Mysore |
CHI | 2 |
| 2018 | On Learning Associations of Faces and Voices
Changil Kim 0001, Hijung Shin, Tae-Hyun Oh, Alexandre Kaspar, Mohamed A. Elgharib, Wojciech Matusik |
ACCV (5) | 2 |
| 2016 | Dynamic Authoring of Audio with Linked ScriptsabstractSpeech recordings are central to modern media from podcasts to audio books to e-lectures and voice-overs. Authoring these recordings involves an iterative back and forth process between script writing/editing and audio recording/editing. Yet, most existing tools treat the script and the audio separately, making the back and forth workflow very tedious. We present Voice Script, an interface to support a dynamic workflow for script writing and audio recording/editing. Our system integrates the script with the audio such that, as the user writes the script or records speech, edits to the script are translated to the audio and vice versa. Through informal user studies, we demonstrate that our interface greatly facilitates the audio authoring process in various scenarios. Hijung Shin, Wilmot Li, Frédo Durand |
UIST | 1 |
| 2016 | Reconciling Elastic and Equilibrium Methods for Static AnalysisabstractWe examine two widely used classes of methods for static analysis of masonry buildings: linear elasticity analysis using finite elements and equilibrium methods. It is often claimed in the literature that finite element analysis is less accurate than equilibrium analysis when it comes to masonry analysis; we examine and qualify this claimed inaccuracy, provide a systematic explanation for the discrepancy observed between their results, and present a unified formulation of the two approaches to stability analysis. We prove that both approaches can be viewed as equivalent, dual methods for getting the same answer to the same problem. We validate our observations with simulations and physical tilt experiments of structures. Hijung Shin, Christopher F. Porst, Etienne Vouga, John Ochsendorf, Frédo Durand |
ACM Trans. Graph. | 1 |
| 2015 | Visual transcripts: lecture notes from blackboard-style lecture videosabstractBlackboard-style lecture videos are popular, but learning using existing video player interfaces can be challenging. Viewers cannot consume the lecture material at their own pace, and the content is also difficult to search or skim. For these reasons, some people prefer lecture notes to videos. To address these limitations, we present Visual Transcripts , a readable representation of lecture videos that combines visual information with transcript text. To generate a Visual Transcript, we first segment the visual content of a lecture into discrete visual entities that correspond to equations, figures, or lines of text. Then, we analyze the temporal correspondence between the transcript and visuals to determine how sentences relate to visual entities. Finally, we arrange the text and visuals in a linear layout based on these relationships. We compare our result with a standard video player, and a state-of-the-art interface designed specifically for blackboard-style lecture videos. User evaluation suggests that users prefer our interface for learning and that our interface is effective in helping them browse or search through lecture videos. Hijung Shin, Floraine Berthouzoz, Wilmot Li, Frédo Durand |
ACM Trans. Graph. | 1 |
| 2012 | Structural optimization of 3D masonry buildingsabstractIn the design of buildings, structural analysis is traditionally performed after the aesthetic design has been determined and has little influence on the overall form. In contrast, this paper presents an approach to guide the form towards a shape that is more structurally sound. Our work is centered on the study of how variations of the geometry might improve structural stability. We define a new measure of structural soundness for masonry buildings as well as cables, and derive its closed-form derivative with respect to the displacement of all the vertices describing the geometry. We start with a gradient descent tool which displaces each vertex along the gradient. We then introduce displacement operators, imposing constraints such as the preservation of orientation or thickness; or setting additional objectives such as volume minimization. Emily Whiting, Hijung Shin, John Ochsendorf, Frédo Durand |
ACM Trans. Graph. | 2 |