Bingyuan Wang

dblp:245/3819 · DBLP profile ↗
← Back
13ranked-venue papers
5as first author
13since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 1 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 6 · 3 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 5 since 2021
YearPublicationVenuePosition
2026 EmoVid: A Multimodal Emotion Video Dataset for Emotion-Centric Video Understanding and Generation
abstract
Emotion plays a pivotal role in video-based expression, but existing video generation systems predominantly focus on low-level visual metrics while neglecting affective dimensions. Although emotion analysis has made progress in the visual domain, the video community lacks dedicated resources to bridge emotion understanding with generative tasks, particularly for stylized and non-realistic contexts. To address this gap, we introduce EmoVid, the first multimodal, emotion-annotated video dataset specifically designed for artistic media, which includes cartoon animations, movie clips, and animated stickers. Each video is annotated with emotion labels, visual attributes (brightness, colorfulness, hue), and text captions. Through systematic analysis, we uncover spatial and temporal patterns linking visual features to emotional perceptions across diverse video forms. Building on these insights, we develop an emotion-conditioned video generation technique by fine-tuning the Wan2.1 model. The results show a significant improvement in both quantitative metrics and the visual quality of generated videos for text-to-video and image-to-video tasks. EmoVid establishes a new benchmark and protocol for affective video computing. Our work not only offers valuable insights into visual emotion analysis in artistic videos but also provides practical methods for enhancing emotional expression in video generation. The extended version and the dataset are available on our project page.
Zongyang Qiu, Bingyuan Wang, Xingbei Chen, Yingqing He, Zeyu Wang 0003
AAAI2
2026 Gen-Diaolou: An Integrated AI-Assisted Interactive System for Diachronic Understanding and Preservation of the Kaiping Diaolou
abstract
The Kaiping Diaolou and Villages, a UNESCO World Heritage Site, exemplify hybrid Chinese and Western architecture shaped by migration culture. However, architectural heritage engagement often faces authenticity debates, resource constraints, and limited participatory approaches. This research explores current challenges of leveraging Artificial Intelligence (AI) for architectural heritage, and how AI-assisted interactive systems can foster cultural heritage understanding and preservation awareness. We conducted a formative study (N=14) to uncover empirical insights from heritage stakeholders that inform design. These insights informed the design of Gen-Diaolou, an integrated AI-assisted interactive system that supports heritage understanding and preservation. A pilot study (N=18) and a museum field study (N=26) provided converging evidence suggesting that Gen-Diaolou may support visitors’ diachronic understanding and preservation awareness, and together informed design implications for future human–AI collaborative systems for digital cultural heritage engagement. More broadly, this work bridges the research gap between passive heritage systems and unconstrained creative tools in the HCI domain.
Xuanchen Lu, Bingyuan Wang, Lujin Zhang, Zeyu Wang 0003, David Kei-Man Yip
CHI4
2026 GatheringSense: AI-Generated Imagery and Embodied Experiences for Understanding Literati Gatherings
abstract
Chinese literati gatherings (Wenren Yaji), as a situated form of Chinese traditional culture, remain underexplored in depth. Although generative AI supports powerful multimodal generation, current cultural applications largely emphasize aesthetic reproduction and struggle to convey the deeper meanings of cultural rituals and social frameworks. Based on embodied cognition, we propose an AI-driven dual-path framework for cultural understanding, which we instantiate through GatheringSense, a literati-gathering experience. We conduct a mixed-methods study (N = 48) to compare how AI-generated multimodal content and embodied participation complement each other in supporting the understanding of literati gatherings and fostering cultural resonance. Our results show that AI-generated content effectively improves the readability of cultural symbols and initial emotional attraction, yet limitations in physical coherence and micro-level credibility may affect users’ satisfaction. In contrast, embodied experience significantly deepens participants’ understanding of ritual rules and social roles, and increases their psychological closeness and presence. Based on these findings, we offer empirical evidence and five transferable design implications for generative experience in cultural heritage.
Bingyuan Wang, Hongcheng Guo, Zeyu Wang 0003
CHI2
2025 DiT4Edit: Diffusion Transformer for Image Editing
abstract
Despite recent advances in UNet-based image editing, methods for shape-aware object editing in high-resolution images are still lacking. Compared to UNet, Diffusion Transformers (DiT) demonstrate superior capabilities to effectively capture the long-range dependencies among patches, leading to higher-quality image generation. In this paper, we propose DiT4Edit, the first Diffusion Transformer-based image editing framework. Specifically, DiT4Edit uses the DPM-Solver inversion algorithm to obtain the inverted latents, reducing the number of steps compared to the DDIM inversion algorithm commonly used in UNet-based frameworks. Additionally, we design unified attention control and patch merging, tailored for transformer computation streams. This integration allows our framework to generate higher-quality edited images faster. Our design leverages the advantages of DiT, enabling it to surpass UNet structures in image editing, especially in high-resolution and arbitrary-size images. Extensive experiments demonstrate the strong performance of DiT4Edit in various editing scenarios, highlighting the potential of diffusion transformers for image editing.
Kunyu Feng, Yue Ma 0016, Bingyuan Wang, Qifeng Chen 0001, Zeyu Wang 0003
AAAI3
2025 Magiccolor: Multi-Instance Sketch Colorization
Yinhan Zhang, Bingyuan Wang
ICCV3
2025 MagicScroll: Enhancing Immersive Storytelling with Controllable Scroll Image Generation
abstract
Scroll images are a unique medium commonly used in virtual reality (VR) providing an immersive visual storytelling experience. Despite rapid advances in diffusion-based image generation, it remains an open research question to generate scroll images suitable for immersive, coherent, and controllable storytelling in VR. This paper proposes a multi-layered, diffusion-based scroll image generation framework with a novel semantic-aware denoising process. We incorporate layout prediction and style control modules to generate coherent scroll images of any aspect ratio. Based on the scroll image generation framework, we use different multi-window strategies to render diverse visual forms such as chains, rings, and forks for VR storytelling. Quantitative and qualitative evaluations demonstrate that our techniques can significantly enhance text-image consistency and visual coherence in scroll image generation, as well as the level of immersion and engagement of VR storytelling. We will release our source code to facilitate better collaborations on immersive storytelling between AI researchers and creative practitioners. https://magicscroll.github.io/
Bingyuan Wang, Hengyu Meng, Lanjiong Li, Yue Ma 0016, Qifeng Chen 0001, Zeyu Wang 0003
VR1
2025 EmotionLens: Interactive visual exploration of the circumplex emotion space in literary works via affective word clouds
abstract
Emotion (e.g., valence and arousal) is an important factor in literature (e.g., poetry and prose), and has rich values for plotting the life and knowledge of historical figures and appreciating the aesthetics of literary works. Currently, digital humanities and computational literature apply data statistics extensively in emotion analysis but lack visual analytics for efficient exploration. To fill the gap, we propose a user-centric approach that integrates advanced machine learning models and intuitive visualization for emotion analysis in literature. We make three main contributions. First, we consolidate a new emotion dataset of literary works in different periods, literary genres, and language contexts, augmented with fine-grained valence and arousal labels. Next, we design an interactive visual analytic system named EmotionLens , which allows users to perform multi-granularity (e.g., individual, group, society) and multi-faceted (e.g., distribution, chronology, correlation) analyses of literary emotions, supporting both exploratory and confirmatory approaches in digital humanities. Specifically, we introduce a novel affective word cloud with augmented word weight, position, and color, to facilitate literary text analysis from an emotional perspective. To validate the usability and effectiveness of EmotionLens , we provide two consecutive case studies, two user studies, and interviews with experts from different domains. Our results show that EmotionLens bridges literary text, emotion, and various other attributes, enables efficient knowledge discovery in massive data, and facilitates raising and validating domain-specific hypotheses in literature.
Bingyuan Wang, Wei Zeng 0004, Zeyu Wang 0003
Vis. Informatics1
2023 From Expanded Cinema to Extended Reality: How AI Can Expand and Extend Cinematic Experiences
abstract
This paper explores the concept of expanded cinema and its relationship to extended reality (XR), focusing on the potential of artificial intelligence (AI) to expand and extend expressive possibilities. Expanded cinema refers to experimental film and multimedia art forms that challenge the conventions of traditional cinema by creating immersive and interactive experiences for audiences. XR, on the other hand, blurs the line between physical and virtual reality, offering immersive storytelling experiences. Both expanded cinema and XR aim to push the boundaries of traditional norms and create immersive experiences through the integration of technology, interactivity, and cross-sensory elements. The paper emphasizes the role of AI in optimizing 3D scene creation for XR and enhancing the overall experience through a case study. It also presents several AI-based techniques, such as generative models and AI-assisted rendering, that facilitate efficient and effective 3D content creation. Additionally, it explores the use of AI plugins in 3D modeling software and the generation of 3D models and textures from 2D images using techniques like GANs and VAEs. The incorporation of AI to extend and expand opens up new possibilities for immersive experiences in the future.
Junrong Song, Bingyuan Wang, Zeyu Wang 0003, David Kei-Man Yip
VINCI2
2023 Simonstown: An AI-facilitated Interactive Story of Love, Life, and Pandemic
abstract
We present an interactive story named Simonstown that demonstrates the love and life of ordinary people in the fictional setting of a fatal pandemic. Technically, the artwork integrates different Artificial Intelligence (AI) technologies in the whole production pipeline, including concept formation, creation, and presentation stages; artistically, this interactive film explores the relationship between human and environment in the contemporary context, especially infused with advanced technologies in daily life. The project serves as a demonstration and case study of AI-facilitated interactive storytelling, including better control with AI and how they integrate with live image projects, as well as using the stand-alone camera for real-time synchronization. Our results highlight the significant contribution of AI in visualizing intricate story branching, translation, and adaptation, presenting AI visualization as a distinct, specialized, and well-suited tool for interactive filmmaking.
Bingyuan Wang, Pinxi Zhu, Hao Li 0176, David Kei-Man Yip, Zeyu Wang 0003
VINCI1
2023 Naturality: A Natural Reflection of Chinese Calligraphy
abstract
We present a machine learning-based interactive video installation powered by CLIP and diffusion models and inspired by the concept of naturality in traditional Chinese calligraphy. The artwork explores contemporary interpretations of this traditional concept through practical methods in Artificial Intelligence Generated Content (AIGC). Technically, the algorithms are based on state-of-the-art perceptual and generative models, incorporating multi-dimensional controls over text-to-image and image-to-image translation; conceptually, this real-time art installation extends the discussion brought by Xu Bing’s pieces Book from the Sky and Square Word Calligraphy. The project explores the possibility of AIGC in bridging human creativity and natural randomness, as well as a shifting creative paradigm enhanced by AI knowledge, perception, and association.
Bingyuan Wang, Kang Zhang 0001, Zeyu Wang 0003
VINCI1
2023 Sequential POI Recommend Based on Personalized Federated Learning
Baisong Liu, Xueyuan Zhang, Jiangcheng Qin, Bingyuan Wang
Neural Process. Lett.5
2022 Ranking-based Federated POI Recommendation with Geographic Effect
abstract
Point of Interest (POI) recommendation system rec-ommends places in which users have never been to but may be in-terested. Traditionally, it centrally collects contextual information and interaction data to model users' preferences, which raises many privacy concerns. The current studies habitually sacrifice the recommendation performance to cope with privacy con-cerns. To protect users' privacy while ensuring the performance of the POI recommendation system, we propose a Ranking-based Federated POI Recommendation with Geographic Effect (RFPG). The RFPG allows users to reserve their private data on local devices to secure privacy. It adaptively constructs an active region to model the geographic effect, enhancing users' personalized preference modeling. In addition, we design a probability-based negative sampling method to protect privacy further and improve recommendation performance. This method calculates the probability of a POI being a negative sample through POI geographic distribution, then combines the positive samples to construct a local triplet training dataset. Theoretical analysis and experiments on two real datasets demonstrate that our proposed RFPG improves the performance of the POI recommendation while protecting users' private data.
Baisong Liu, Xueyuan Zhang, Jiangcheng Qin, Bingyuan Wang, Jiangbo Qian
IJCNN5
2022 Exploiting high-order behaviour patterns for cross-domain sequential recommendation
abstract
The cross-domain sequential recommendation aims to predict the next item based on a sequence of recorded user behaviours in multiple domains. We propose a novel Cross-domain Sequential Recommendation approach with Graph-Collaborative Filtering (CsrGCF) to alleviate the sparsity issue of user-interaction data. Specifically, we design time-aware and relation-aware graph attention mechanisms with collaborative filtering to exploit high-order behaviour patterns of users for promising results in both domains. Time-aware Graph Attention mechanism (TGAT) is designed to learn the inter-domain sequence-level representation of items. Relationship-aware Graph Attention mechanism (RGAT) is proposed to learn collaborative items' and users' feature representations. Moreover, to simultaneously improve the recommendation performance in the two domains, a Cross-domain Feature Bidirectional Transfer module (CFBT) is proposed, transferring user's common sharing features in both domains and retaining user's domain-specific features in a specific domain. Finally, cross-domain and sequential information jointly recommend the next items that users like. We conduct extensive experiments on two real-world datasets that show that CsrGCF outperforms several state-of-the-art baselines in terms of Recall and MRR. These demonstrate the necessity of exploiting high-order behaviour patterns of users for a cross-domain sequential recommendation. Meanwhile, retaining domain-specific features is an important step in the process of cross-domain feature bidirectional transferring.
Bingyuan Wang, Baisong Liu, Xueyuan Zhang, Jiangcheng Qin, Jiangbo Qian
Connect. Sci.1