Weitao You

dblp:216/7211 · also Wei-tao You · DBLP profile ↗
← Back
27ranked-venue papers
5as first author
25since 2021 · last 2026
0000-0002-9625-5547ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 13 · 1 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 3 first-author · 11 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author
YearPublicationVenuePosition
2026 IEvoAgent: Evolving Conversational Agent based on User Implicit Feedback
abstract
Yichen Cai, Jiayang Li, Junyuan Qiu, Jingya Guo, Weitao You, Changyuan Yang, Lingyun Sun, Pei Chen. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Yichen Cai 0005, Jiayang Li 0003, Junyuan Qiu, Jingya Guo, Weitao You, Chang-yuan Yang, Lingyun Sun, Pei Chen 0005
ACL (1)5
2026 Seeing Eye to Eye: Enabling Cognitive Alignment Through Shared First-Person Perspective in Human-AI Collaboration: Seeing Eye to Eye
abstract
Despite advances in multimodal AI, current vision-based assistants often remain inefficient in collaborative tasks. We identify two key gulfs: a communication gulf, where users must translate rich parallel intentions into verbal commands due to the channel mismatch, and an understanding gulf, where AI struggles to interpret subtle embodied cues. To address these, we propose Eye2Eye, a framework that leverages first-person perspective as a channel for human-AI cognitive alignment. It integrates three components: (1) joint attention coordination for fluid focus alignment, (2) revisable memory to maintain evolving common ground, and (3) reflective feedback allowing users to clarify and refine AI’s understanding. We implement this framework in an AR prototype and evaluate it through a user study and a post-hoc pipeline evaluation. Results show that Eye2Eye significantly reduces task completion time and interaction load while increasing trust, demonstrating its components work in concert to improve collaboration.
Zhuyu Teng, Pei Chen 0005, Yichen Cai 0005, Ruoqing Lu, Zhaoqu Jiang, Jiayang Li 0003, Weitao You, Lingyun Sun
CHI7
2026 ImmersiProtor: A Collaborative Mixed-Prototype Tool Integrating Spatial Augmented Reality and Component-layered Generation
abstract
Conceptual design is a critical stage in product development, which is a co-design process involving multidisciplinary collaboration based on prototypes. In this paper, we aim to propose a novel prototype paradigm that combines the distinct strengths of generative artificial intelligence (GAI) and spatial augmented reality (SAR), leveraging the expressive potential of SAR and the creative potential of GAI for co-design. To achieve this, we initially conducted a formative study with designers to explore how these technologies could be effectively combined to facilitate co-design. Based on our findings, we introduce ImmersiProtor, a prototype tool integrating multi-view SAR and component-layered GAI for co-design. On one hand, ImmersiProtor allows design team members to freely create and modify physical prototypes while automatically generating multi-view and high-fidelity renderings that are projected onto the surfaces of the physical prototype using SAR technology, enabling immersive communication and intuitive evaluation. On the other hand, ImmersiProtor introduces a component-layered generation and collaboration mode, offering both personal and shared team component resources. It ensures that individual team members can explore ideas independently without interference, while also supporting concept integration, evaluation, and iteration. We implemented ImmersiProtor, which involves a web-based application and an SAR design space. We conducted a user study to verify ImmersiProtor’s usability in supporting prototype and collaboration. Our results highlighted ImmersiProtor’s inherent strengths in enhancing intuition, promoting collaboration, and strengthening GAI controllability. We also explored the effect of mixed interaction on design and critically discuss its best practices for the HCI community.
Chaoyi Lin, Weitao You, Lingyun Sun, Pei Chen 0005
CHI4
2025 Personalized Dynamic Music Emotion Recognition with Dual-Scale Attention-Based Meta-Learning
abstract
Dynamic Music Emotion Recognition (DMER) aims to predict the emotion of different moments in music, playing a crucial role in music information retrieval. The existing DMER methods struggle to capture long-term dependencies when dealing with sequence data, which limits their performance. Furthermore, these methods often overlook the influence of individual differences on emotion perception, even though everyone has their own personalized emotional perception in the real world. Motivated by these issues, we explore more effective sequence processing methods and introduce the Personalized DMER (PDMER) problem, which requires models to predict emotions that align with personalized perception. Specifically, we propose a Dual-Scale Attention-Based Meta-Learning (DSAML) method. This method fuses features from a dual-scale feature extractor and captures both short and long-term dependencies using a dual-scale attention transformer, improving the performance in traditional DMER. To achieve PDMER, we design a novel task construction strategy that divides tasks by annotators. Samples in a task are annotated by the same annotator, ensuring consistent perception. Leveraging this strategy alongside meta-learning, DSAML can predict personalized perception of emotions with just one personalized annotation sample. Our objective and subjective experiments demonstrate that our method can achieve state-of-the-art performance in both traditional DMER and PDMER.
Dengming Zhang, Weitao You, Lingyun Sun, Pei Chen 0005
AAAI2
2025 GVMGen: A General Video-to-Music Generation Model with Hierarchical Attentions
abstract
Composing music for video is essential yet challenging, leading to a growing interest in automating music generation for video applications. Existing approaches often struggle to achieve robust music-video correspondence and generative diversity, primarily due to inadequate feature alignment methods and insufficient datasets. In this study, we present General Video-to-Music Generation model (GVMGen), designed for generating high-related music to the video input. Our model employs hierarchical attentions to extract and align video features with music in both spatial and temporal dimensions, ensuring the preservation of pertinent features while minimizing redundancy. Remarkably, our method is versatile, capable of generating multi-style music from different video inputs, even in zero-shot scenarios. We also propose an evaluation model along with two novel objective metrics for assessing video-music alignment. Additionally, we have compiled a large-scale dataset comprising diverse types of video-music pairs. Experimental results demonstrate that GVMGen surpasses previous models in terms of music-video correspondence, music quality generative diversity, and application universality.
Heda Zuo, Weitao You, Junxian Wu 0003, Shihong Ren, Pei Chen 0005, Mingxu Zhou, Yujia Lu, Lingyun Sun
AAAI2
2025 CoExploreDS: Framing and Advancing Collaborative Design Space Exploration Between Human and AI
Pei Chen 0005, Zhuoyi Cheng, Yichen Cai 0005, Jiayang Li 0003, Weitao You, Lingyun Sun
CHI6
2025 Exploring the Design of Human Speech Indicators to Enhance Waiting Experience in Voice User Interface
abstract
Waiting for system loading is a common scenario that often diminishes user experience, leading to dissatisfaction. Well-established visual indicators like progress bars can not directly apply to the interactions with voice assistants (VAs) like Siri. As VAs continue to rise in popularity, this research aims to explore the design of auditory indicators, particularly human speech, for optimizing waiting experiences in Voice User Interfaces (VUIs). We first organized focus groups (N=35) to identify design considerations for speech indicators, uncovering design opportunities in integrating explanations and humor. Subsequently, we conducted an empirical study (N=30) to evaluate the effects of speech indicators with two levels of explanation and humor on the waiting experience, measured by attention, perceived time, pleasure, and overall satisfaction, during both short and long loading durations. Our findings suggest significant potential for incorporating explanations and humor into VUIs, offering actionable insights for designing effective speech indicators that improve waiting experiences.
Wenan Li, Junnan Yu, Yehong Zhou, Jinlei Shi, Weitao You, Zhibin Zhou 0002
CHI5
2025 IEDS: Exploring an Intelli-Embodied Design Space Combining Designer, AR, and GAI to Support Industrial Conceptual Design
Pei Chen 0005, Zhaoqu Jiang, Xuelong Xie, Weitao You, Lingyun Sun
CHI7
2025 Intoner: For Chinese Poetry Intoning Synthesis
abstract
Chinese Poetry Intoning, with improvised melodies devoid of fixed musical scores, is crucial for emotional expression and prosodic rendition. However, this cultural heritage faces challenges in propagation due to scant audio records and a scarcity of domain experts. Existing text-to-speech models lack the ability to generate melodious audio, while singing-voice-synthesis models rely on predetermined musical scores, which are all unsuitable for intoning synthesis. Hence, we introduce Chinese Poetry Intoning Synthesis (PIS) as a novel task to reproduce intoning audio and preserve this age-old cultural art. Corresponding to this task, we summarize three-level principles from poetry metrical patterns and construct a diffusion PIS model Intoner based on them. We also collect a multi-style Chinese poetry intoning dataset of text-audio pairs accompanied by feature annotations. Experimental results show that our model effectively learns diverse intoning styles and contents which can synthesize more melodious and vibrant intoning audio. To the best of our knowledge, we are the first to work on poetry intoning synthesis task.
Heda Zuo, Liyao Sun, Zeyu Lai, Weitao You, Pei Chen 0005, Lingyun Sun
IJCAI4
2025 Controllable Video-to-Music Generation with Multiple Time-Varying Conditions
abstract
Music enhances video narratives and emotions, driving demand for automatic video-to-music (V2M) generation. However, existing V2M methods relying solely on visual features or supplementary textual inputs generate music in a black-box manner, often failing to meet user expectations. To address this challenge, we propose a novel multi-condition guided V2M generation framework that incorporates multiple time-varying conditions for enhanced control over music generation. Our method uses a two-stage training strategy that enables learning of V2M fundamentals and audiovisual temporal synchronization while meeting users' needs for multi-condition control. In the first stage, we introduce a fine-grained feature selection module and a progressive temporal alignment attention mechanism to ensure flexible feature alignment. For the second stage, we develop a dynamic conditional fusion module and a control-guided decoder module to integrate multiple conditions and accurately guide the music composition process. Extensive experiments demonstrate that our method outperforms existing V2M pipelines in both subjective and objective evaluations, significantly enhancing control and alignment with user expectations.
Junxian Wu 0003, Weitao You, Heda Zuo, Dengming Zhang, Pei Chen 0005, Lingyun Sun
ACM Multimedia2
2025 Spatial-Temporal Decomposition and Alignment in Controllable Video-to-Music Generation
abstract
Achieving high-quality output alongside enhanced controllability is crucial in video-to-music generation, especially for optimizing user experience in real-life application scenarios. Most existing studies emphasize generative quality, but often overlooking the vital aspect of controllability. Therefore, the generated music cannot be easily fine-tuned or modified to meet users' expectations. In this paper, we delve into the spatial-temporal decomposition and alignment in controllable video-to-music generation. We first introduce a novel video-music decomposition and transformation approach in both spatial and temporal domain, and enhance the cross-modal correspondence through feature alignment and flow-matching based alignment. Furthermore, our method attains unsupervised controllability during training via feature-free guidance. Experimental results demonstrate that our model achieves state-of-the-art results in overall generative quality. Moreover, its controllability significantly outperforms existing models, making it exceptionally well-suited to accommodate users' flexible and diverse control requirements.
Weitao You, Heda Zuo, Junxian Wu 0003, Dengming Zhang, Zhibin Zhou 0002, Lingyun Sun
ACM Multimedia1
2025 GPSdesign: Integrating Generative AI with Problem-Solution Co-Evolution Network to Support Product Conceptual Design
abstract
In conceptual design, designers often face the challenge of navigating vast design spaces to define ambiguous problems and generate feasible solutions. Recent advancements in generative artificial intelligence (GenAI) offer new opportunities to support this process. However, formative research revealed that designers struggle to simultaneously advance both problem and solution spaces when using GenAI in conceptual design, leading to increased communication load and diminished solution practicality. This study explores the integration of GenAI with the problem-solution co-evolution model to facilitate the construction of a structured design space. We propose a GenAI-supported method for expanding and evaluating the design space and developed the GPSdesign system based on this method. Compared with a baseline system, GPSdesign fosters greater design space divergence, retrospection, and structured construction, while improving design efficiency and solution quality.
Pei Chen 0005, Yexinrui Wu, Zhuoshu Li, Mingxu Zhou, Weitao You, Lingyun Sun
Int. J. Hum. Comput. Interact.7
2025 CONDA: Introducing Context-Aware Decision Making Assistant in Virtual Reality for Interior Renovation
abstract
Customized interiors enhance quality of life and self-expression, driving demand for VR-based design solutions. However, scant research exists on exploiting contextual cues in VR to aid decision making. Consequently, we propose CONDA, a context-aware assistant which leveraging LLMs to support interior renovation decisions. Specifically, we reconstruct users’ homes in VR and provide CONDA with stylistic details and spatial layouts, allowing it to predict furniture labels based on the decision scenario. Besides, we devise various modes to comprehensively express users’ purchasing preferences. Finally, CONDA recommend compatible items based on the label matching algorithm, and generate multi-dimensional explanations. A 30-user study reveals contextual completeness and preference diversity critically influence recommendation quality and decision behaviors, with 90% praising CONDA’s performance and all expressing daily-use intent. Overall, we validated the efficacy and practicality of CONDA, deriving universal design insights for VR decision-support systems and establishing new research directions.CCS ConceptsHuman-centered computing → Virtual realityComputing methodologies → Natural language generationApplied computing → Computer-aided design
Yizhan Shao, Weitao You, Ziqing Zheng, Yinyu Lu, Chang-yuan Yang, Zhibin Zhou 0002
Int. J. Hum. Comput. Interact.2
2025 Using a Configurational Approach to Examine the Impacts of Vehicle Appearance Perception on Pedestrian Acceptance of the External Human-Machine Interfaces on Autonomous Vehicles
abstract
The interaction between autonomous vehicles (AVs) and pedestrians has gained significant attention, leading to the exploration of external human-machine interfaces (eHMIs) equipped on AVs to facilitate effective communication. While existing research suggests that perceptions of vehicle appearances may influence interactions between pedestrians and AVs, a comprehensive study on the eHMIs related to AV appearance remains lacking. Therefore, we conducted a virtual reality (VR) experiment to investigate how AV appearances affect pedestrians’ acceptance regarding Awareness, Intent, and Harmony during interactions with various eHMIs. Leveraging the fuzzy set qualitative comparative analysis (fsQCA) method, we identified specific combinations of AV appearances and eHMIs that yield either high or low-performance interactions. For example, our findings reveal that text displays exhibit high performance in terms of awareness on AVs that are aggressive and ordinary. Furthermore, we distilled design guidelines to provide actionable suggestions for the design of eHMIs, fostering the acceptance of AVs among pedestrians.
Zhibin Zhou 0002, Yitao Fan, Wenan Li, Hao Jiang 0046, Weitao You, Lingyun Sun
Int. J. Hum. Comput. Interact.5
2025 DesignManager: An Agent-Powered Copilot for Designers to Integrate AI Design Tools into Creative Workflows
abstract
Creative design is an inherently complex and iterative process characterized by continuous exploration, evaluation, and refinement. While recent advances in generative AI have demonstrated remarkable potential in supporting specific design tasks, there remains a critical gap in understanding how these technologies can enhance the holistic design process rather than just isolated stages. This paper introduces DesignManager, a novel AI-powered design support system that aims to transform how designers collaborate with AI throughout their creative workflow. Through a formative study examining designers' current practices with generative AI, we identified key challenges and opportunities in integrating AI into the creative design process. Based on these insights, we developed DesignManager as an interactive copilot system that provides node-based visualization of design evolution, enabling designers to track, modify, and branch their design processes while maintaining meaningful dialogue-based collaboration. The system offers two collaboration modes: DesignManager-guiding and Designer-guiding. Designers can engage in conversational interactions with the DesignManager to obtain design inspiration and tool recommendations, and proactively advance the design progress. The system employs an agent framework to manage decoupled contextual information emerged during the design process, facilitating deep understanding of designers' needs and providing context-aware assistance. Our technical evaluation validated the effectiveness of context decoupling and the use of agent framework, while the open-ended user study with experts demonstrated that DesignManager successfully supports intuitive intention expression, flexible process control, and deeper creative articulation. This work contributes to the understanding of how AI can evolve from task-specific tools to collaborative partners in creative design processes.
Weitao You, Yinyu Lu, Zirui Ma, Nan Li 0085, Mingxu Zhou, Pei Chen 0005, Lingyun Sun
ACM Trans. Graph.1
2025 PaRUS: A Virtual Reality Shopping Method Focusing on Contextual Information between Products and Real Usage Scenes
abstract
The development of AR and VR technologies is enhancing users' online shopping experiences in various ways. However, in existing VR shopping applications, shopping contexts merely refer to the products and virtual malls or metaphorical scenes where users select products. This leads to the defect that users can only imagine rather than intuitively feel whether the selected products are suitable for their real usage scenes, resulting in a significant discrepancy between their expectations before and after the purchase. To address this issue, we propose PaRUS, a VR shopping approach that focuses on the context between products and their real usage scenes. PaRUS begins by rebuilding the virtual scenario of the products' real usage scene through a new semantic scene reconstruction pipeline (manual operation needed), which preserves both the structured scene and textured object models in the scene. Afterwards, intuitive visualization of how the selected products fit the reconstructed virtual scene is provided. We conducted two user studies to evaluate how PaRUS impacts user experience, behavior, and satisfaction with their purchase. The results indicated that PaRUS significantly reduced the perceived performance risk and improved users' trust and expectation with their results of purchase.
Yinyu Lu, Weitao You, Ziqing Zheng, Yizhan Shao, Chang-yuan Yang, Zhibin Zhou 0002
IEEE Trans. Vis. Comput. Graph.2
2024 ProtoDreamer: A Mixed-prototype Tool Combining Physical Model and Generative AI to Support Conceptual Design
abstract
Prototyping serves as a critical phase in the industrial conceptual design process, enabling exploration of problem space and identification of solutions. Recent advancements in large-scale generative models have enabled AI to become a co-creator in this process. However, designers often consider generative AI challenging due to the necessity to follow computer-centered interaction rules, diverging from their familiar design materials and languages. Physical prototype is a commonly used design method, offering unique benefits in prototype process, such as intuitive understanding and tangible testing. In this study, we propose ProtoDreamer, a mixed-prototype tool that synergizes generative AI with physical prototype to support conceptual design. ProtoDreamer allows designers to construct preliminary prototypes using physical materials, while AI recognizes these forms and vocal inputs to generate diverse design alternatives. This tool empowers designers to tangibly interact with prototypes, intuitively convey design intentions to AI, and continuously draw inspiration from the generated artifacts. An evaluation study confirms ProtoDreamer’s utility and strengths in time efficiency, creativity support, defects exposure, and detailed thinking facilitation.
Pei Chen 0005, Xuelong Xie, Chaoyi Lin, Lianyan Liu, Zhuoshu Li, Weitao You, Lingyun Sun
UIST7
2024 StyleFactory: Towards Better Style Alignment in Image Creation through Style-Strength-Based Control and Evaluation
abstract
Generative AI models have been widely used for image creation. However, generating images that are well-aligned with users’ personal styles on aesthetic features (e.g., color and texture) can be challenging due to the poor style expression and interpretation between humans and models. Through a formative study, we observed that participants showed a clear subjective perception of the desired style and variations in its strength, which directly inspired us to develop style-strength-based control and evaluation. Building on this, we present StyleFactory, an interactive system that helps users achieve style alignment. Our interface enables users to rank images based on their strengths in the desired style and visualizes the strength distribution of other images in that style from the model’s perspective. In this way, users can evaluate the understanding gap between themselves and the model, and define well-aligned personal styles for image creation through targeted iterations. Our technical evaluation and user study demonstrate that StyleFactory accurately generates images in specific styles, effectively facilitates style alignment in image creation workflow, stimulates creativity, and enhances the user experience in human-AI interactions.
Mingxu Zhou, Dengming Zhang, Weitao You, Chenghao Pan, Tianyu Lao, Pei Chen 0005
UIST3
2024 The Doctrine of the Mean: Chinese Calligraphy with Moderate Visual Complexity Elicits High Aesthetic Preference
abstract
Chinese calligraphy is a symbol of cultural heritage; it is widely used in visual design (e.g., film posters and modern fashion) for its aesthetic value. In human-computer interaction, visual design complexity strongly influences users’ preferences. Indeed, the relevance of visual complexity to aesthetic preference has been confirmed. However, the visual complexity of calligraphy artworks has seldom been investigated and the process of how visual complexity affects aesthetic preference remains vague. Therefore, in this work we proposed a computational method to evaluate complexity and conducted several perception studies. Results showed that layout features and calligraphy style (regular script, running script, and cursive script) affected visual complexity. The level of visual complexity (low, medium, and high) affected aesthetic preference but calligraphy style did not. Furthermore, Chinese calligraphy with moderate visual complexity evokes strong aesthetic preference. The present findings can help designers redesign websites and interfaces for high aesthetic preference and can provide insights for developing the theoretical design of calligraphic art for advanced interaction in cultural heritage, such as in interactive systems for teaching calligraphy based on aesthetic cognition.
Kaixin Han, Weitao You, Shuhui Shi, Huanghuang Deng, Lingyun Sun
Int. J. Hum. Comput. Interact.2
2024 The Impression of Round and Square: Chinese Calligraphy Aesthetics in Modern Type Design
abstract
Round and square are important concepts in Chinese aesthetics. The legacy of Chinese calligraphy aesthetics has led the way to modern design. In HCI, interface typefaces influence users’ reading experiences. Users’ aesthetic preferences may influence their reading experiences as well. The investigation of how human aesthetic preferences are affected by character visual features (e.g., character size and character weight) in type design has received substantial attention. However, it is difficult to accurately study human aesthetic preferences through these visual features in Chinese calligraphy-type design. How Chinese character visual features affect human aesthetic preferences is not clear. Inspired by graphic design, different visual features (e.g., color and texture) can cause changes in impression expressions (e.g., modern and classical), thus affecting human aesthetic preferences. In this article, we introduced a character impression analysis method to investigate how human aesthetic preferences are affected by Chinese character visual features. First, we proposed two impression expressions of character (round and square) from the concepts of Chinese aesthetics; Then, for these two impression expressions, we modeled four visual features of calligraphy characters from the principles of Gestalt psychology. Finally, we conducted two studies to investigate how visual features affect impression expressions and explored the relationship between aesthetic preferences. The results clearly explain how character visual features affect human aesthetic preferences. The present study can give designers a deep understanding of how visual features affect human aesthetic preferences based on character impression analysis. It can guide designers to make better type design decisions to achieve a high aesthetic preference for users. Furthermore, the findings of the present study can potentially provide new insights into Latin type design.
Kaixin Han, Weitao You, Shuhui Shi, Lingyun Sun
Int. J. Hum. Comput. Interact.2
2024 HierVid: Lowering the Barriers to Entry of Interactive Video Making with a Hierarchical Authoring System
abstract
Interactive videos have been applied to various areas due to their engagement potential and efficiency improvement of information communication. However, creating interactive videos can be challenging because of a lack of novice-oriented guidance in current platforms, and the logic-building process when authoring interactive videos. To address these challenges, we obtained insights from four creativity support tool designers, proposed a series of hierarchical interactive video structures based on existing narrative structures, and presented the HierVid system. The system is designed as a Template-Module-Unit Mode-based hierarchical authoring platform grounded on three design requirements, and we conducted two user studies to evaluate HierVid. The results showed that novice users could get started to use and understand the functions easily, and the system allowed users to use and explore freely, with an enhanced efficiency compared to the bilibili platform. In conclusion, our research and design of HierVid offer guidance and support for novice users, making interactive video authoring quicker and more accessible.
Weitao You, Zhuoyi Cheng, Zirui Ma, Guang Yang 0022, Zhibin Zhou 0002, Lingyun Sun
Int. J. Hum. Comput. Interact.1
2024 LanT: finding experts for digital calligraphy character restoration
Kaixin Han, Weitao You, Huanghuang Deng, Lingyun Sun, Jinyu Song, Zijin Hu, Heyang Yi
Multim. Tools Appl.2
2024 Correction: LanT: finding experts for digital calligraphy character restoration
Kaixin Han, Weitao You, Huanghuang Deng, Lingyun Sun, Jinyu Song, Zijin Hu, Heyang Yi
Multim. Tools Appl.2
2024 Automatic Generation of Interactive Nonlinear Video for Online Apparel Shopping Navigation
abstract
We present an automatic generation pipeline of interactive nonlinear video for online apparel shopping navigation. Our approach was inspired by Google's “Messy Middle” theory, which suggests that people mentally are faced with two tasks—exploration and evaluation—before purchasing online. Given a set of apparel product presentation videos, our navigation UI organizes them to optimize users' product exploration and automatically generates interactive videos for users' product evaluation. To support automatic methods, we proposed a video clustering similarity ($\operatorname{CSIM}$) and a camera movement similarity ($\operatorname{MSIM}$), as well as a comparative video generation algorithm for product recommendation, presentation, and comparison. To evaluate our pipeline's effectiveness, we conducted several user studies. The results showed that our pipeline can help users complete the consumption process more efficiently, making it easier for them to understand and choose a product.
Weitao You, Juntao Ji, Lingyun Sun, Chang-yuan Yang, Mi Yu, Shi Chen 0005
IEEE Trans. Multim.1
2024 Hearing with the eyes: modulating lyrics typography for music visualization
Kaixin Han, Weitao You, Shuhui Shi, Lingyun Sun
Vis. Comput.2
2020 Automatic synthesis of advertising images according to a specified style
abstract
Images are widely used by companies to advertise their products and promote awareness of their brands. The automatic synthesis of advertising images is challenging because the advertising message must be clearly conveyed while complying with the style required for the product, brand, or target audience. In this study, we proposed a data-driven method to capture individual design attributes and the relationships between elements in advertising images with the aim of automatically synthesizing the input of elements into an advertising image according to a specified style. To achieve this multi-format advertisement design, we created a dataset containing 13 280 advertising images with rich annotations that encompassed the outlines and colors of the elements, in addition to the classes and goals of the advertisements. Using our probabilistic models, users guided the style of synthesized advertisements via additional constraints (e.g., context-based keywords). We applied our method to a variety of design tasks, and the results were evaluated in several perceptual studies, which showed that our method improved users’ satisfaction by 7.1% compared to designs generated by nonprofessional students, and that more users preferred the coloring results of our designs to those generated by the color harmony model and Colormind.
Weitao You, Hao Jiang 0046, Zhi-Yuan Yang, Chang-yuan Yang, Lingyun Sun
Frontiers Inf. Technol. Electron. Eng.1
2018 Crowdsourcing intelligent design
abstract
Design intelligence, namely, artificial intelligence to solve creative problems and produce creative ideas, has improved rapidly with the new generation artificial intelligence. However, existing methods are more skillful in learning from data and have limitations in creating original ideas different from the training data. Crowdsourcing offers a promising method to produce creative designs by combining human inspiration and machines’ computational ability. We propose a crowdsourcing intelligent design method called ‘flexible crowdsourcing design’. Design ideas produced through crowdsourcing design can be unreliable and inconsistent because they rely solely on selection among participants’ submissions of ideas. In contrast, the flexible crowdsourcing design method employs a cultivation procedure that integrates the ideas from crowd participants and cultivates these ideas to improve design quality at the same time. We introduce a series of studies to show how flexible crowdsourcing design can produce original design ideas consistently. Specifically, we will describe the typical procedure of flexible crowdsourcing design, the refined crowdsourcing tasks, the factors that affect the idea development process, the method for calculating idea development potential, and two applications of the flexible crowdsourcing design method. Finally, it summarizes the design capabilities enabled by crowdsourcing intelligent design. This method enhances the performance of crowdsourcing design and supports the development of design intelligence.
Wei Xiang 0008, Lingyun Sun, Weitao You, Chang-yuan Yang
Frontiers Inf. Technol. Electron. Eng.3