VLDB 2026 Research / reviewers in the wild / expert
Zeyu Wang 0003
dblp:132/7882-3
· DBLP profile ↗
41ranked-venue papers
5as first author
39since 2021 · last 2026
0000-0001-5374-6330ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 20 · 3 first-author · 18 since 2021Human-computer interaction and ubiquitous computing · 18 · 2 first-author · 18 since 2021Artificial intelligence and machine learning · 10 · 10 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | EvDiff3D: Event-Aware Diffusion Repair for High-Fidelity Event-Based 3D ReconstructionabstractEvent cameras are bio-inspired sensors that capture visual information through asynchronous brightness changes, offering distinct advantages including high temporal resolution and wide dynamic range. While prior research has investigated event-based 3D reconstruction for extreme scenarios, existing methods face inherent limitations and fail to fully exploit the unique characteristics of event data. In this paper, we present EvDiff3D, a novel two-stage 3D reconstruction framework that integrates event-based geometric constraints with an event-aware diffusion prior for appearance refinement. Our key insight lies in bridging the gap between physically grounded event-based reconstruction and data-driven appearance repair through a unified cyclical pipeline. In the first stage, we reconstruct a coarse 3D scene under supervision from event loss and event-based monocular depth constraints to preserve structural fidelity. The second stage fine-tunes an event-aware diffusion model based on a pretrained video diffusion model as a repair prior to enhance the appearance in under-constrained regions. Based on the diffusion model, our pipeline operates within a reconstruction-generation cycle that progressively refines both geometry and appearance using only event data. Extensive experiments on synthetic and real-world datasets demonstrate that EvDiff3D significantly outperforms existing methods in perceptual quality and structural consistency. Kanghao Chen, Lin Wang 0025, Zeyu Wang 0003 |
AAAI | 5 |
| 2026 | EmoVid: A Multimodal Emotion Video Dataset for Emotion-Centric Video Understanding and GenerationabstractEmotion plays a pivotal role in video-based expression, but existing video generation systems predominantly focus on low-level visual metrics while neglecting affective dimensions. Although emotion analysis has made progress in the visual domain, the video community lacks dedicated resources to bridge emotion understanding with generative tasks, particularly for stylized and non-realistic contexts. To address this gap, we introduce EmoVid, the first multimodal, emotion-annotated video dataset specifically designed for artistic media, which includes cartoon animations, movie clips, and animated stickers. Each video is annotated with emotion labels, visual attributes (brightness, colorfulness, hue), and text captions. Through systematic analysis, we uncover spatial and temporal patterns linking visual features to emotional perceptions across diverse video forms. Building on these insights, we develop an emotion-conditioned video generation technique by fine-tuning the Wan2.1 model. The results show a significant improvement in both quantitative metrics and the visual quality of generated videos for text-to-video and image-to-video tasks. EmoVid establishes a new benchmark and protocol for affective video computing. Our work not only offers valuable insights into visual emotion analysis in artistic videos but also provides practical methods for enhancing emotional expression in video generation. The extended version and the dataset are available on our project page. Zongyang Qiu, Bingyuan Wang, Xingbei Chen, Yingqing He, Zeyu Wang 0003 |
AAAI | 5 |
| 2026 | Gen-Diaolou: An Integrated AI-Assisted Interactive System for Diachronic Understanding and Preservation of the Kaiping DiaolouabstractThe Kaiping Diaolou and Villages, a UNESCO World Heritage Site, exemplify hybrid Chinese and Western architecture shaped by migration culture. However, architectural heritage engagement often faces authenticity debates, resource constraints, and limited participatory approaches. This research explores current challenges of leveraging Artificial Intelligence (AI) for architectural heritage, and how AI-assisted interactive systems can foster cultural heritage understanding and preservation awareness. We conducted a formative study (N=14) to uncover empirical insights from heritage stakeholders that inform design. These insights informed the design of Gen-Diaolou, an integrated AI-assisted interactive system that supports heritage understanding and preservation. A pilot study (N=18) and a museum field study (N=26) provided converging evidence suggesting that Gen-Diaolou may support visitors’ diachronic understanding and preservation awareness, and together informed design implications for future human–AI collaborative systems for digital cultural heritage engagement. More broadly, this work bridges the research gap between passive heritage systems and unconstrained creative tools in the HCI domain. Xuanchen Lu, Bingyuan Wang, Lujin Zhang, Zeyu Wang 0003, David Kei-Man Yip |
CHI | 6 |
| 2026 | SketchDynamics: Exploring Free-Form Sketches for Dynamic Intent Expression in Animation GenerationabstractSketching provides an intuitive way to convey dynamic intent in animation authoring (i.e., how elements change over time and space), making it a natural medium for automatic content creation. Yet existing approaches often constrain sketches to fixed command tokens or predefined visual forms, overlooking their free-form nature and the central role of humans in shaping intention. To address this, we introduce an interaction paradigm where users convey dynamic intent to a vision–language model via free-form sketching, instantiated here in a sketch storyboard to motion graphics workflow. We implement an interface and improve it through a three-stage study with 24 participants. The study shows how sketches convey motion with minimal input, how their inherent ambiguity requires users to be involved for clarification, and how sketches can visually guide video refinement. Our findings reveal the potential of sketch–AI interaction to bridge the gap between intention and outcome, and demonstrate its applicability to 3D animation and video generation. Boyu Li 0007, Lin-Ping Yuan, Zeyu Wang 0003, Hongbo Fu 0001 |
CHI | 3 |
| 2026 | GatheringSense: AI-Generated Imagery and Embodied Experiences for Understanding Literati GatheringsabstractChinese literati gatherings (Wenren Yaji), as a situated form of Chinese traditional culture, remain underexplored in depth. Although generative AI supports powerful multimodal generation, current cultural applications largely emphasize aesthetic reproduction and struggle to convey the deeper meanings of cultural rituals and social frameworks. Based on embodied cognition, we propose an AI-driven dual-path framework for cultural understanding, which we instantiate through GatheringSense, a literati-gathering experience. We conduct a mixed-methods study (N = 48) to compare how AI-generated multimodal content and embodied participation complement each other in supporting the understanding of literati gatherings and fostering cultural resonance. Our results show that AI-generated content effectively improves the readability of cultural symbols and initial emotional attraction, yet limitations in physical coherence and micro-level credibility may affect users’ satisfaction. In contrast, embodied experience significantly deepens participants’ understanding of ritual rules and social roles, and increases their psychological closeness and presence. Based on these findings, we offer empirical evidence and five transferable design implications for generative experience in cultural heritage. Bingyuan Wang, Hongcheng Guo, Zeyu Wang 0003 |
CHI | 5 |
| 2026 | SmartTracer: Interactive tracing-based stroke extraction for complex line art
Zhongyue Guan, Zeyu Wang 0003 |
Comput. Aided Geom. Des. | 3 |
| 2026 | Layer3D: A 3D Layered Representation for Multiview Vector GraphicsabstractAbstract We present Layer3D, a novel 3D neural representation that models objects as collections of decomposable neural implicit primitives. These primitives enable the generation of layered images with consistent correspondences across viewpoints, establishing a flexible framework for multiview vector graphics decomposition. Integrated into a text‐to‐3D pipeline via Score Distillation Sampling (SDS), Layer3D learns to generate primitives with diverse shape topologies while preserving structural coherence. To ensure front‐to‐back ordering required for 2D flat graphics, our method incorporates front‐to‐back rendering and shape regularization constraints. Experimental results demonstrate that Layer3D consistently produces meaningful, topology‐diverse layers across multiple views, thereby facilitating intuitive and effective layer‐based vector editing. Zhongyue Guan, Zeyu Wang 0003 |
Comput. Graph. Forum | 3 |
| 2026 | Simulating the Real World: A Unified Survey of Multimodal Generative ModelsabstractUnderstanding and replicating the real world is a critical challenge in Artificial General Intelligence (AGI) research. To achieve this, many existing approaches, such as world models, aim to capture the fundamental principles governing the physical world, enabling more accurate simulations and meaningful interactions. However, current methods often treat different modalities, including 2D (images), videos, 3D, and 4D representations, as independent domains, overlooking their interdependencies. Additionally, these methods typically focus on isolated dimensions of reality without systematically integrating their connections. In this survey, we present a unified survey for multimodal generative models that investigate the progression of data dimensionality in real-world simulation. Specifically, this survey starts from 2D generation (appearance), then moves to video (appearance+dynamics) and 3D generation (appearance+ geometry), and finally culminates in 4D generation that integrate all dimensions. To the best of our knowledge, this is the first attempt to systematically unify the study of 2D, video, 3D, and 4D generation within a single framework. To guide future research, we provide a comprehensive review of datasets, evaluation metrics, and future directions to foster insights for newcomers. This survey serves as a bridge to advance the study of multimodal generative models and real-world simulation within a unified framework. Longguang Wang, Yuwei Guo 0002, Yukai Shi, Anyi Rao, Zeyu Wang 0003, Hui Xiong 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 9 |
| 2026 | LayerInbetween: Occlusion-Aware Stroke Correspondence and Inbetweening with Automatic LayeringabstractThe Hong Kong University of Science and Technology (Guangzhou), China and The Hong Kong University of Science and Technology, China Establishing one-to-one stroke correspondences is fundamental to vector-based animation inbetweening. Animators may face great challenges when handling occlusion, as occluded strokes must be drawn explicitly in keyframes and manually hidden frame by frame after stroke interpolation. To reduce tedious effort, we present LayerInbetween , an occlusion-aware framework for vector stroke correspondence and automatic inbetweening. It performs automatic layering to guide stroke tracing and correspondence finding for occluded strokes, and to resolve occlusion with layers in the inbetween frames. To predict occluded strokes, we propose a Global-Local Layer Transformation (GLLT) module that progressively improves the spatial alignment of strokes across keyframes via layer guidance, thereby indicating their potential positions. Our framework is trained on a synthetic dataset comprising 17k+ pairs of keyframes with occlusion and their stroke correspondences. Extensive experiments demonstrate the effectiveness of LayerInbetween compared with existing methods and its generalization capabilities to various types of drawings. In addition to its superior performance, our vector-based inbetweening method enables more flexible editing of 2D animation than raster-based video generation. Code and data for this paper are available at https://github.com/MarkMoHR/LayerInbetween. Haoran Mo, Zhongyue Guan, Zeyu Wang 0003 |
ACM Trans. Graph. | 4 |
| 2026 | DoodleAssist: Progressive Interactive Line Art Generation With Latent Distribution AlignmentabstractCreating high-quality line art in a fast and controlled manner plays a crucial role in anime production and concept design. We present DoodleAssist, an interactive and progressive line art generation system controlled by sketches and prompts, which helps both experts and novices concretize their design intentions or explore possibilities. Built upon a controllable diffusion model, our system performs progressive generation based on the last generated line art, synthesizing regions corresponding to drawn or modified strokes while keeping the remaining ones unchanged. To facilitate this process, we propose a latent distribution alignment mechanism to enhance the transition between the two regions and allow seamless blending, thereby alleviating issues of region incoherence and line discontinuity. Finally, we also build a user interface that allows the convenient creation of line art through interactive sketching and prompts. Qualitative and quantitative comparisons against existing approaches and an in-depth user study demonstrate the effectiveness and usability of our system. Our system can benefit various applications such as anime concept design, drawing assistant, and creativity support for children. Haoran Mo, Yulin Shen 0001, Edgar Simo-Serra, Zeyu Wang 0003 |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2026 | Direct vs. Score-based Selection: Understanding the Heisenberg Effect in Target Acquisition Across Input Modalities in Virtual RealityabstractTarget selection is a fundamental interaction in virtual reality (VR). But the act of confirming a selection, such as a button press or pinch, can disturb the tracked pose and shift the intended target, which is referred to as the Heisenberg Effect. Prior research has mainly investigated controller input. However, it remains unclear how the effect manifests in the bare-hand input and how score-based techniques may mitigate the effect in different spatial variations. To fill the gap, we conduct a within-subject study to examine the Heisenberg Effect across two input modalities (i.e., controller and hand) and two selection mechanisms (i.e., direct and score-based). Our results show that hand input is more susceptible to the Heisenberg Effect, with direct selection more influenced by target width and score-based selection more sensitive to target density. Based on previous vote-oriented technique and our temporal analysis, we introduce weighted VOTE, a history-based intention accuracy model for target voting, that reweights recent interaction intent to counteract input disturbances. Our evaluation shows the method improves selection accuracy compared to baseline techniques. Finally, we discuss future directions for adaptive selection methods. Linjie Qiu, Duotun Wang, Boyu Li 0007, Jiawei Li 0009, Yulin Shen 0001, Zeyu Wang 0003, Mingming Fan 0001 |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2025 | DEGAS: Detailed Expressions on Full-Body Gaussian AvatarsabstractAlthough neural rendering has made significant ad-vances in creating lifelike, animatable full-body and head avatars, incorporating detailed expressions into full-body avatars remains largely unexplored. We present DEGAS, the first 3D Gaussian Splatting (3DGS)-based modeling method for full-body avatars with rich facial expressions. Trained on multiview videos of a given subject, our method learns a conditional variational autoencoder that takes both the body motion and facial expression as driving signals to generate Gaussian maps in the UV layout. To drive the facial expressions, instead of the commonly used 3D Mor-phable Models (3DMMs) in 3D head avatars, we propose to adopt the expression latent space trained solely on 2D portrait images, bridging the gap between 2D talking faces and 3D avatars. Leveraging the rendering capability of 3DGS and the rich expressiveness of the expression latent space, the learned avatars can be reenacted to reproduce photo-realistic rendering images with subtle and accurate facial expressions. Experiments on an existing dataset and our newly proposed dataset offull-body talking avatars demonstrate the efficacy of our method. We also propose an audio-driven extension of our method with the help of 2D talking faces, opening new possibilities for interactive AI agents. Project page: https://initialneil.github.io/DEGAS. Zhijing Shao, Duotun Wang, Qing-Yao Tian, Yao-Dong Yang, Hengyu Meng, Yu Zhang 0166, Kang Zhang 0001, Zeyu Wang 0003 |
3DV | 10 |
| 2025 | HeadEvolver: Text to Head Avatars via Expressive and Attribute-Preserving Mesh DeformationabstractCurrent text-to-avatar methods often rely on implicit representations (e.g., NeRF, SDF, and DMTet), leading to 3D content that artists cannot easily edit and animate in graphics software. This paper introduces a novel framework for generating stylized head avatars from text guidance, which leverages locally learnable mesh deformation and 2D diffusion priors to achieve high-quality digital assets for attribute-preserving manipulation. Given a template mesh, our method represents mesh deformation with perface Jacobians and adaptively modulates local deformation using a learnable vector field. This vector field enables anisotropic scaling while preserving the rotation of vertices, which can better express identity and geometric details. We also employ landmark- and contour-based regularization terms to balance the expressiveness and plausibility of generated head avatars from multiple views without relying on any specific shape prior. Our framework can generate realistic shapes and textures that can be further edited via text, while supporting seamless editing using the preserved attributes from the template mesh, such as 3DMM parameters, blendshapes, and UV coordinates. Extensive experiments demonstrate that our framework can generate diverse and expressive head avatars with high-quality meshes that artists can easily manipulate in 3D graphics software, facilitating downstream applications such as efficient asset creation and animation with preserved attributes. Duotun Wang, Hengyu Meng, Zhijing Shao, Qianxi Liu, Lin Wang 0025, Mingming Fan 0001, Xiaohang Zhan, Zeyu Wang 0003 |
3DV | 9 |
| 2025 | DiT4Edit: Diffusion Transformer for Image EditingabstractDespite recent advances in UNet-based image editing, methods for shape-aware object editing in high-resolution images are still lacking. Compared to UNet, Diffusion Transformers (DiT) demonstrate superior capabilities to effectively capture the long-range dependencies among patches, leading to higher-quality image generation. In this paper, we propose DiT4Edit, the first Diffusion Transformer-based image editing framework. Specifically, DiT4Edit uses the DPM-Solver inversion algorithm to obtain the inverted latents, reducing the number of steps compared to the DDIM inversion algorithm commonly used in UNet-based frameworks. Additionally, we design unified attention control and patch merging, tailored for transformer computation streams. This integration allows our framework to generate higher-quality edited images faster. Our design leverages the advantages of DiT, enabling it to surpass UNet structures in image editing, especially in high-resolution and arbitrary-size images. Extensive experiments demonstrate the strong performance of DiT4Edit in various editing scenarios, highlighting the potential of diffusion transformers for image editing. Kunyu Feng, Yue Ma 0016, Bingyuan Wang, Qifeng Chen 0001, Zeyu Wang 0003 |
AAAI | 7 |
| 2025 | GS-ID: Illumination Decomposition on Gaussian Splatting via Adaptive Light Aggregation and Diffusion-Guided Material Priors
Yulin Shen 0001, Zeyu Wang 0003 |
ICCV | 4 |
| 2025 | Text2VDM: Text to Vector Displacement Maps for Expressive and Interactive 3D SculptingabstractProfessional 3D asset creation often requires diverse sculpting brushes to add surface details and geometric structures. Despite recent progress in 3D generation, producing reusable sculpting brushes compatible with artists' workflows remains an open and challenging problem. These sculpting brushes are typically represented as vector displacement maps (VDMs), which existing models cannot easily generate compared to natural images. This paper presents Text2VDM, a novel framework for text-to-VDM brush generation through the deformation of a dense planar mesh guided by score distillation sampling (SDS). The original SDS loss is designed for generating full objects and struggles with generating desirable sub-object structures from scratch in brush generation. We refer to this issue as semantic coupling, which we address by introducing weighted blending of prompt tokens to SDS, resulting in a more accurate target distribution and semantic guidance. Experiments demonstrate that Text2VDM can generate diverse, high-quality VDM brushes for sculpting surface details and geometric structures. Our generated brushes can be seamlessly integrated into mainstream modeling software, enabling various applications such as mesh stylization and real-time interactive modeling. Hengyu Meng, Duotun Wang, Zhijing Shao, Zeyu Wang 0003 |
ICCV | 5 |
| 2025 | Event-Guided Consistent Video Enhancement with Modality-Adaptive Diffusion PipelineabstractRecent advancements in low-light video enhancement (LLVE) have increasingly leveraged both RGB and event cameras to improve video quality under challenging conditions. However, existing approaches share two key drawbacks. First, they are tuned for steady low-light scenes, so their performance drops when illumination varies. Second, they assume every sensing modality is always available, while real systems may lose or corrupt one of them. These limitations make the methods brittle in dynamic, real-world settings. In this paper, we propose EVDiffuser, a novel framework for consistent LLVE that integrates RGB and event data through a modality-adaptive diffusion pipeline. By harnessing the powerful priors of video diffusion models, EVDiffuser enables consistent video enhancement and generalization to diverse scenarios under varying illumination, where RGB or events may even be absent. Specifically, we first design a modality-agnostic conditioning mechanism based on a diffusion pipeline by treating the two modalities as optional conditions, which is fine-tuned using augmented and integrated datasets. Furthermore, we introduce a modality-adaptive guidance rescaling that dynamically adjusts the contribution of each modality according to sensor-specific characteristics. Additionally, we establish a benchmark that accounts for varying illumination and diverse real-world scenarios, facilitating future research on consistent event-guided LLVE. Our experiments demonstrate state-of-the-art performance across challenging scenarios (i.e., varying illumination) and sensor-based settings (e.g., event-only, RGB-only), highlighting the generalization of our framework. Kanghao Chen, Guoqiang Liang 0003, Lutao Jiang, Zeyu Wang 0003, Ying-Cong Chen |
NeurIPS | 5 |
| 2025 | GaussianShopVR: Facilitating Immersive 3D Authoring Using Gaussian Splatting in VR
Yulin Shen 0001, Boyu Li 0007, Jiayang Huang, David Kei-Man Yip, Zeyu Wang 0003 |
UIST | 5 |
| 2025 | VideoCraft: A Mixed Reality-empowered Video Generation Workflow with Spatial Layer Editing for Concept Video Creation
Boyu Li 0007, Linping Yuan, Zeyu Wang 0003 |
UIST | 3 |
| 2025 | Expanding Virtual Production Frontiers: AI-Driven Workflows for Enhanced Cinematic CreationabstractAlthough virtual production (VP) offers cinematic and immersive storytelling by aligning a high-end camera with multiple LED screens through a central server, the creation of high-quality 3D scenery with real-time interactions remains complex and resource-intensive. This paper explores the potential of expanding cinematic virtual production scenes through three innovative AI-driven approaches: (1) AI-generated 360° panoramas from text prompts to produce immersive backgrounds; (2) Direct text/AI-generated image-to-3D environment conversion using mesh generation; (3) An end-to-end AI pipeline for rapid stylized scene construction. We demonstrate each workflow through real-time avatar-interactive shooting scenarios. Our approach bridges technical and artistic domains and aims to show how these AI-driven workflows could accelerate scene creation, enable novel cinematic experiences, and reveal the future potential of visual content generation. Junrong Song, Hongcheng Guo, Lujin Zhang, Zeyu Wang 0003, David Kei-Man Yip |
VINCI | 4 |
| 2025 | MagicScroll: Enhancing Immersive Storytelling with Controllable Scroll Image GenerationabstractScroll images are a unique medium commonly used in virtual reality (VR) providing an immersive visual storytelling experience. Despite rapid advances in diffusion-based image generation, it remains an open research question to generate scroll images suitable for immersive, coherent, and controllable storytelling in VR. This paper proposes a multi-layered, diffusion-based scroll image generation framework with a novel semantic-aware denoising process. We incorporate layout prediction and style control modules to generate coherent scroll images of any aspect ratio. Based on the scroll image generation framework, we use different multi-window strategies to render diverse visual forms such as chains, rings, and forks for VR storytelling. Quantitative and qualitative evaluations demonstrate that our techniques can significantly enhance text-image consistency and visual coherence in scroll image generation, as well as the level of immersion and engagement of VR storytelling. We will release our source code to facilitate better collaborations on immersive storytelling between AI researchers and creative practitioners. https://magicscroll.github.io/ Bingyuan Wang, Hengyu Meng, Lanjiong Li, Yue Ma 0016, Qifeng Chen 0001, Zeyu Wang 0003 |
VR | 8 |
| 2025 | LineDrawer: Stroke-level process reconstruction of complex line art based on human perception
Zhongyue Guan, Zeyu Wang 0003 |
Comput. Graph. | 3 |
| 2025 | DesignMemo: Integrating Discussion Context into Online Collaboration with Enhanced Design Rationale TrackingabstractRemote collaborative design has become increasingly popular, but current design tools often overlook the importance of contextual communication during synchronized design activities, which is critical for understanding the rationale and decisions behind design choices. In this paper, we introduce DesignMemo, a proof-of-concept system that integrates the verbal context of remote discussions into visual design history tracking. The system automatically labels the visual elements with an annotation, which is linked to a certain transcript of the meeting, so that the user can easily recall the context of the visual design by clicking the element. The system also integrates an LLM agent for annotation-oriented summarization based on global context tracking, so users can quickly follow the rationale of the design without reading the lengthy transcript. Our user study with 24 participants suggests that the ability to track communication context makes the iterative design process smoother and more efficient. Boyu Li 0007, Linjie Qiu, Duotun Wang, Qianxi Liu, Ryo Suzuki 0001, Mingming Fan 0001, Zeyu Wang 0003 |
Proc. ACM Hum. Comput. Interact. | 7 |
| 2025 | FocalSelect: Improving Occluded Objects Acquisition with Heuristic Selection and Disambiguation in Virtual RealityabstractIn recent years, various head-worn virtual reality (VR) techniques have emerged to enhance object selection for occluded or distant targets. However, many approaches focus solely on ray-casting inputs, restricting their use with other input methods, such as bare hands. Additionally, some techniques speed up selection by changing the user's perspective or modifying the scene context, which may complicate interactions when users plan to resume or manipulate the scene afterward. To address these challenges, we present FocalSelect, a heuristic selection technique that builds 3D disambiguation through head-hand coordination and scoring-based functions. Our interaction design adheres to the principle that the intended selection range is a small sector of the headset's viewing frustum, allowing optimal targets to be identified within this scope. We also introduce a density-aware adjustable occlusion plane for effective depth culling of rendered objects. Two experiments are conducted to assess the adaptability of FocalSelect across different input modalities and its performance against five selection techniques. The results indicate that FocalSelect enhances selection experiences in occluded and remote scenarios while preserving the spatial context among objects. This preservation helps maintain users' understanding of the original scene and facilitates further manipulation. We also explore potential applications and enhancements to demonstrate more practical implementations of FocalSelect. Duotun Wang, Linjie Qiu, Boyu Li 0007, Qianxi Liu, Xiaoying Wei, Jianhao Chen 0002, Zeyu Wang 0003, Mingming Fan 0001 |
IEEE Trans. Vis. Comput. Graph. | 7 |
| 2025 | EmotionLens: Interactive visual exploration of the circumplex emotion space in literary works via affective word cloudsabstractEmotion (e.g., valence and arousal) is an important factor in literature (e.g., poetry and prose), and has rich values for plotting the life and knowledge of historical figures and appreciating the aesthetics of literary works. Currently, digital humanities and computational literature apply data statistics extensively in emotion analysis but lack visual analytics for efficient exploration. To fill the gap, we propose a user-centric approach that integrates advanced machine learning models and intuitive visualization for emotion analysis in literature. We make three main contributions. First, we consolidate a new emotion dataset of literary works in different periods, literary genres, and language contexts, augmented with fine-grained valence and arousal labels. Next, we design an interactive visual analytic system named EmotionLens , which allows users to perform multi-granularity (e.g., individual, group, society) and multi-faceted (e.g., distribution, chronology, correlation) analyses of literary emotions, supporting both exploratory and confirmatory approaches in digital humanities. Specifically, we introduce a novel affective word cloud with augmented word weight, position, and color, to facilitate literary text analysis from an emotional perspective. To validate the usability and effectiveness of EmotionLens , we provide two consecutive case studies, two user studies, and interviews with experts from different domains. Our results show that EmotionLens bridges literary text, emotion, and various other attributes, enables efficient knowledge discovery in massive data, and facilitates raising and validating domain-specific hypotheses in literature. Bingyuan Wang, Wei Zeng 0004, Zeyu Wang 0003 |
Vis. Informatics | 6 |
| 2024 | Get Your Hands Dirty? A Comparative Study of Tool Usage and Perceptual Engagement in Physical and Digital SculptingabstractThe creation of 3D content, crucial in various applications, is often challenging and time-intensive. While digital tools are prevalent for 3D content creation, traditional clay sculpting offers an embodied experience that fosters artists’ perceptual engagement with physical space, enhancing their interactive and cognitive connection with the creation process. We conducted an eight-day live sculpting session at an art academy, systematically comparing the creative workflows of eight professional artists in both physical and digital mediums. Our qualitative and quantitative analysis include artists’ differences in tool usage between physical and digital sculpting, variations in visual and tactile perceptual engagement, and the potential for future integration of the two modalities. Our study provides insights into the benefits of physical and digital sculpting and may inform future design of hybrid interfaces for 3D content creation. Hengyu Meng, Yanan Jin, Zeyu Wang 0003 |
Creativity & Cognition | 6 |
| 2024 | Neural Canvas: Supporting Scenic Design Prototyping by Integrating 3D Sketching and Generative AIabstractWe propose Neural Canvas, a lightweight 3D platform that integrates sketching and a collection of generative AI models to facilitate scenic design prototyping. Compared with traditional 3D tools, sketching in a 3D environment helps designers quickly express spatial ideas, but it does not facilitate the rapid prototyping of scene appearance or atmosphere. Neural Canvas integrates generative AI models into a 3D sketching interface and incorporates four types of projection operations to facilitate 2D-to-3D content creation. Our user study shows that Neural Canvas is an effective creativity support tool, enabling users to rapidly explore visual ideas and iterate 3D scenic designs. It also expedites the creative process for both novices and artists who wish to leverage generative AI technology, resulting in attractive and detailed 3D designs created more efficiently than using traditional modeling tools or individual generative AI platforms. Yulin Shen 0001, Yifei Shen 0002, Jiawen Cheng, Chutian Jiang, Mingming Fan 0001, Zeyu Wang 0003 |
CHI | 6 |
| 2024 | Toward Making Virtual Reality (VR) More Inclusive for Older Adults: Investigating Aging Effect on Target Selection and Manipulation Tasks in VRabstractRecent studies show the promise of VR in improving physical, cognitive, and emotional health of older adults. However, prior work on optimizing object selection and manipulation performance in VR was mostly conducted among younger adults. It remains unclear how older adults would perform such tasks compared to younger adults and the challenges they might face. To fill in this gap, we conducted two studies with both older and younger adults to understand their performances and user experiences of object selection and manipulation in VR respectively. Based on the results, we delineated interaction difficulties that older adults exhibited in VR and identified multiple factors, such as headset-related neck fatigue, extra head movements from out-of-view interactions, and slow spatial perceptions, that significantly decreased the motor performance of older adults. We further proposed design recommendations for improving the accessibility of direct interaction experiences in VR for older adults. Zhiqing Wu, Duotun Wang, Shumeng Zhang, Yuru Huang, Zeyu Wang 0003, Mingming Fan 0001 |
CHI | 5 |
| 2024 | SplattingAvatar: Realistic Real-Time Human Avatars With Mesh-Embedded Gaussian SplattingabstractWe present SplattingAvatar, a hybrid 3D representation of photorealistic human avatars with Gaussian Splatting em-bedded on a triangle mesh, which renders over 300 FPS on a modern GPU and 30 FPS on a mobile device. We disentangle the motion and appearance of a virtual human with explicit mesh geometry and implicit appearance modeling with Gaus-sian Splatting. The Gaussians are defined by barycentric coordinates and displacement on a triangle mesh as Phong surfaces. We extend lifted optimization to simultaneously op-timize the parameters of the Gaussians while walking on the triangle mesh. SplattingAvatar is a hybrid representation of virtual humans where the mesh represents low-frequency motion and surface deformation, while the Gaussians take over the high-frequency geometry and detailed appearance. Un-like existing deformation methods that rely on an MLP-based linear blend skinning (LBS) field for motion, we control the rotation and translation of the Gaussians directly by mesh, which empowers its compatibility with various animation techniques, e.g., skeletal animation, blend shapes, and mesh editing. Trainable from monocular videos for both full-body and head avatars, SplattingAvatar shows state-of-the-art ren-dering quality across multiple datasets. Code and data are available at https://github.com/initialneil/SplattingAvatar. Zhijing Shao, Duotun Wang, Xiangru Lin, Yu Zhang 0166, Mingming Fan 0001, Zeyu Wang 0003 |
CVPR | 8 |
| 2024 | LaSe-E2V: Towards Language-guided Semantic-aware Event-to-Video ReconstructionabstractEvent cameras harness advantages such as low latency, high temporal resolution, and high dynamic range (HDR), compared to standard cameras. Due to the distinct imaging paradigm shift, a dominant line of research focuses on event-to-video (E2V) reconstruction to bridge event-based and standard computer vision. However, this task remains challenging due to its inherently ill-posed nature: event cameras only detect the edge and motion information locally. Consequently, the reconstructed videos are often plagued by artifacts and regional blur, primarily caused by the ambiguous semantics of event data. In this paper, we find language naturally conveys abundant semantic information, rendering it stunningly superior in ensuring semantic consistency for E2V reconstruction. Accordingly, we propose a novel framework, called LaSe-E2V, that can achieve semantic-aware high-quality E2V reconstruction from a language-guided perspective, buttressed by the text-conditional diffusion models. However, due to diffusion models' inherent diversity and randomness, it is hardly possible to directly apply them to achieve spatial and temporal consistency for E2V reconstruction. Thus, we first propose an Event-guided Spatiotemporal Attention (ESA) module to condition the event data to the denoising pipeline effectively. We then introduce an event-aware mask loss to ensure temporal coherence and a noise initialization strategy to enhance spatial consistency. Given the absence of event-text-video paired data, we aggregate existing E2V datasets and generate textual descriptions using the tagging models for training and evaluation. Extensive experiments on three datasets covering diverse challenging scenarios (e.g., fast motion, low light) demonstrate the superiority of our method. Demo videos for the results are attached to the project page. Kanghao Chen, Jiazhou Zhou, Zeyu Wang 0003, Lin Wang 0025 |
NeurIPS | 4 |
| 2024 | AniCraft: Crafting Everyday Objects as Physical Proxies for Prototyping 3D Character Animation in Mixed RealityabstractWe introduce AniCraft, a mixed reality system for prototyping 3D character animation using physical proxies crafted from everyday objects. Unlike existing methods that require specialized equipment to support the use of physical proxies, AniCraft only requires affordable markers, webcams, and daily accessible objects and materials. AniCraft allows creators to prototype character animations through three key stages: selection of virtual characters, fabrication of physical proxies, and manipulation of these proxies to animate the characters. This authoring workflow is underpinned by diverse physical proxies, manipulation types, and mapping strategies, which ease the process of posing virtual characters and mapping user interactions with physical proxies to animated movements of virtual characters. We provide a range of cases and potential applications to demonstrate how diverse physical proxies can inspire user creativity. User experiments show that our system can outperform traditional animation methods for rapid prototyping. Furthermore, we provide insights into the benefits and usage patterns of different materials, which lead to design implications for future research. Boyu Li 0007, Linping Yuan, Qianxi Liu, Yulin Shen 0001, Zeyu Wang 0003 |
UIST | 6 |
| 2023 | PointShopAR: Supporting Environmental Design Prototyping Using Point Cloud in Augmented RealityabstractWe present PointShopAR, a novel tablet-based system for AR environmental design using point clouds as the underlying representation. It integrates point cloud capture and editing in a single AR workflow to help users quickly prototype design ideas in their spatial context. We hypothesize that point clouds are well suited for prototyping, as they can be captured more rapidly than textured meshes and then edited immediately in situ on the capturing device. We based the design of PointShopAR on the practical needs of six architects in a formative study. Our system supports a variety of point cloud editing operations in AR, including selection, transformation, hole filling, drawing, morphing, and animation. We evaluate PointShopAR through a remote study on usability and an in-person study on environmental design support. Participants were able to iterate design rapidly, showing the merits of an integrated capture and editing workflow with point clouds in AR environmental design. Zeyu Wang 0003, Cuong Nguyen 0003, Paul Asente, Julie Dorsey |
CHI | 1 |
| 2023 | From Expanded Cinema to Extended Reality: How AI Can Expand and Extend Cinematic ExperiencesabstractThis paper explores the concept of expanded cinema and its relationship to extended reality (XR), focusing on the potential of artificial intelligence (AI) to expand and extend expressive possibilities. Expanded cinema refers to experimental film and multimedia art forms that challenge the conventions of traditional cinema by creating immersive and interactive experiences for audiences. XR, on the other hand, blurs the line between physical and virtual reality, offering immersive storytelling experiences. Both expanded cinema and XR aim to push the boundaries of traditional norms and create immersive experiences through the integration of technology, interactivity, and cross-sensory elements. The paper emphasizes the role of AI in optimizing 3D scene creation for XR and enhancing the overall experience through a case study. It also presents several AI-based techniques, such as generative models and AI-assisted rendering, that facilitate efficient and effective 3D content creation. Additionally, it explores the use of AI plugins in 3D modeling software and the generation of 3D models and textures from 2D images using techniques like GANs and VAEs. The incorporation of AI to extend and expand opens up new possibilities for immersive experiences in the future. Junrong Song, Bingyuan Wang, Zeyu Wang 0003, David Kei-Man Yip |
VINCI | 3 |
| 2023 | Simonstown: An AI-facilitated Interactive Story of Love, Life, and PandemicabstractWe present an interactive story named Simonstown that demonstrates the love and life of ordinary people in the fictional setting of a fatal pandemic. Technically, the artwork integrates different Artificial Intelligence (AI) technologies in the whole production pipeline, including concept formation, creation, and presentation stages; artistically, this interactive film explores the relationship between human and environment in the contemporary context, especially infused with advanced technologies in daily life. The project serves as a demonstration and case study of AI-facilitated interactive storytelling, including better control with AI and how they integrate with live image projects, as well as using the stand-alone camera for real-time synchronization. Our results highlight the significant contribution of AI in visualizing intricate story branching, translation, and adaptation, presenting AI visualization as a distinct, specialized, and well-suited tool for interactive filmmaking. Bingyuan Wang, Pinxi Zhu, Hao Li 0176, David Kei-Man Yip, Zeyu Wang 0003 |
VINCI | 5 |
| 2023 | Naturality: A Natural Reflection of Chinese CalligraphyabstractWe present a machine learning-based interactive video installation powered by CLIP and diffusion models and inspired by the concept of naturality in traditional Chinese calligraphy. The artwork explores contemporary interpretations of this traditional concept through practical methods in Artificial Intelligence Generated Content (AIGC). Technically, the algorithms are based on state-of-the-art perceptual and generative models, incorporating multi-dimensional controls over text-to-image and image-to-image translation; conceptually, this real-time art installation extends the discussion brought by Xu Bing’s pieces Book from the Sky and Square Word Calligraphy. The project explores the possibility of AIGC in bridging human creativity and natural randomness, as well as a shifting creative paradigm enhanced by AI knowledge, perception, and association. Bingyuan Wang, Kang Zhang 0001, Zeyu Wang 0003 |
VINCI | 3 |
| 2023 | OdorV-Art: An Initial Exploration of An Olfactory Intervention for Appreciating Style Information of Artworks in Virtual MuseumabstractStyle information, such as tone, mood, and genre of artworks, is important for museum visitors to appreciate them better. However, such information can be challenging for non-art specialists to comprehend in the short period that they view artworks. The sense of smell is instrumental for humans to assist their image memory, color, emotion, and shape association. However, it is rarely used in the appreciation of artworks. Taking Western landscape painting as an example, this research explores the following research questions (RQs): 1) How does the intervention of the sense of smell improve the acquisition of style information in paintings? 2) How does the intervention of the sense of smell enhance the immersion in painting appreciation? To answer RQs, we first recruited seven art specialists to participate in a co-design workshop to design a prototype of the virtual museum with olfactory intervention. We then conducted an experiment with 12 non-specialists who viewed several paintings in the VR museum while being exposed to olfactory stimuli that were designed to be correlated with the style information of the paintings. We found potential effects of smell stimuli on enhancing the perception of style information for non-art specialists. Moreover, we found that olfactory intervention has both positive and negative impacts on immersiveness. Finally, we provide design implications for future virtual museum design with olfactory stimuli. Shumeng Zhang, Shihan Fu, Zeyu Wang 0003, Mingming Fan 0001 |
VINCI | 6 |
| 2021 | DistanciAR: Authoring Site-Specific Augmented Reality Experiences for Remote EnvironmentsabstractMost augmented reality (AR) authoring tools only support the author’s current environment, but designers often need to create site-specific experiences for a different environment. We propose DistanciAR, a novel tablet-based workflow for remote AR authoring. Our baseline solution involves three steps. A remote environment is captured by a camera with LiDAR; then, the author creates an AR experience from a different location using AR interactions; finally, a remote viewer consumes the AR content on site. A formative study revealed understanding and navigating the remote space as key challenges with this solution. We improved the authoring interface by adding two novel modes: Dollhouse, which renders a bird’s-eye view, and Peek, which creates photorealistic composite images using captured images. A second study compared this improved system with the baseline, and participants reported that the new modes made it easier to understand and navigate the remote scene. Zeyu Wang 0003, Cuong Nguyen 0003, Paul Asente, Julie Dorsey |
CHI | 1 |
| 2021 | The Role of Subsurface Scattering in Glossiness PerceptionabstractThis study investigates the potential impact of subsurface light transport on gloss perception for the purposes of broadening our understanding of visual appearance in computer graphics applications. Gloss is an important attribute for characterizing material appearance. We hypothesize that subsurface scattering of light impacts the glossiness perception. However, gloss has been traditionally studied as a surface-related quality and the findings in the state-of-the-art are usually based on fully opaque materials, although the visual cues of glossiness can be impacted by light transmission as well. To address this gap and to test our hypothesis, we conducted psychophysical experiments and found that subjects are able to tell the difference in terms of gloss between stimuli that differ in subsurface light transport but have identical surface qualities and object shape. This gives us a clear indication that subsurface light transport contributes to a glossy appearance. Furthermore, we conducted additional experiments and found that the contribution of subsurface scattering to gloss varies across different shapes and levels of surface roughness. We argue that future research on gloss should include transparent and translucent media and to extend the perceptual models currently limited to surface scattering to more general ones inclusive of subsurface light transport. Davit Gigilashvili, Zeyu Wang 0003, Marius Pedersen, Jon Yngve Hardeberg, Holly E. Rushmeier |
ACM Trans. Appl. Percept. | 3 |
| 2021 | Tracing versus freehand for evaluating computer-generated drawingsabstractNon-photorealistic rendering (NPR) and image processing algorithms are widely assumed as a proxy for drawing. However, this assumption is not well assessed due to the difficulty in collecting and registering freehand drawings. Alternatively, tracings are easier to collect and register, but there is no quantitative evaluation of tracing as a proxy for freehand drawing. In this paper, we compare tracing, freehand drawing, and computer-generated drawing approximation (CGDA) to understand their similarities and differences. We collected a dataset of 1,498 tracings and freehand drawings by 110 participants for 100 image prompts. Our drawings are registered to the prompts and include vector-based timestamped strokes collected via stylus input. Comparing tracing and freehand drawing, we found a high degree of similarity in stroke placement and types of strokes used over time. We show that tracing can serve as a viable proxy for freehand drawing because of similar correlations between spatio-temporal stroke features and labeled stroke types. Comparing hand-drawn content and current CGDA output, we found that 60% of drawn pixels corresponded to computer-generated pixels on average. The overlap tended to be commonly drawn content, but people's artistic choices and temporal tendencies remained largely uncaptured. We present an initial analysis to inform new CGDA algorithms and drawing applications, and provide the dataset for use by the community. Zeyu Wang 0003, Sherry Qiu, Nicole Feng, Holly E. Rushmeier, Leonard McMillan, Julie Dorsey |
ACM Trans. Graph. | 1 |
| 2019 | AniCode: authoring coded artifacts for network-free personalized animations
Zeyu Wang 0003, Shiyu Qiu, Natallia Trayan, Alexander Ringlein, Julie Dorsey, Holly E. Rushmeier |
Vis. Comput. | 1 |
| 2016 | Perceptual enhancement for stereoscopic videos based on horopter consistencyabstractAudience discomfort, such as eye strain and dizziness, is one of the urgent issues that virtual reality and 3D movie technologies should tackle. Except for inappropriate horizontal and vertical disparity, one major problem is that people's binocular vergence and focal length in the cinema remain inconsistent from normal visual habits. Psychologists discovered the horopter and Panum's fusional area to describe zero-disparity points projected on the retinas based on accommodation-convergence consistency. In this paper, inspired by these concepts, we propose a stereoscopic effect correction system for perceptual enhancement according to fixated region and scene information. As a preprocessing step, tracking and stereo matching algorithms are implemented to prepare cues for further transformation in 3D space. Then in order to accomplish certain visual effects, we describe a geometric framework for disparity refinement and image warping based on parameter adjustment of the virtual stereoscopic rig. For evaluation, subjective experiments have been conducted to prove the effectiveness of our method. Therefore, our work provides a possibility to improve the audience experience from a formerly underexplored perspective. Zeyu Wang 0003, Xiaohan Jin, Renju Li, Hongbin Zha, Katsushi Ikeuchi |
VRST | 1 |