Chang-yuan Yang

dblp:216/7120 · also Changyuan Yang · DBLP profile ↗
← Back
20ranked-venue papers
1as first author
15since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 6 since 2021Artificial intelligence and machine learning · 6 · 1 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 5 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 3Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 ThinkPersona: Thinking with Persona Graphs for Faithful Individualized Role-Playing
abstract
Yichen Cai, Pei Chen, Jiayang Li, Jingya Guo, Zejian Li, Changyuan Yang, Lingyun Sun. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Yichen Cai 0005, Pei Chen 0005, Jiayang Li 0003, Jingya Guo, Zejian Li, Chang-yuan Yang, Lingyun Sun
ACL (1)6
2026 IEvoAgent: Evolving Conversational Agent based on User Implicit Feedback
abstract
Yichen Cai, Jiayang Li, Junyuan Qiu, Jingya Guo, Weitao You, Changyuan Yang, Lingyun Sun, Pei Chen. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Yichen Cai 0005, Jiayang Li 0003, Junyuan Qiu, Jingya Guo, Weitao You, Chang-yuan Yang, Lingyun Sun, Pei Chen 0005
ACL (1)6
2025 Distribution Backtracking Builds A Faster Convergence Trajectory for Diffusion Distillation
abstract
Accelerating the sampling speed of diffusion models remains a significant challenge. Recent score distillation methods distill a heavy teacher model into a student generator to achieve one-step generation, which is optimized by calculating the difference between two score functions on the samples generated by the student model. However, there is a score mismatch issue in the early stage of the score distillation process, since existing methods mainly focus on using the endpoint of pre-trained diffusion models as teacher models, overlooking the importance of the convergence trajectory between the student generator and the teacher model. To address this issue, we extend the score distillation process by introducing the entire convergence trajectory of the teacher model and propose $\textbf{Dis}$tribution $\textbf{Back}$tracking Distillation ($\textbf{DisBack}$). DisBask is composed of two stages: $\textit{Degradation Recording}$ and $\textit{Distribution Backtracking}$. $\textit{Degradation Recording}$ is designed to obtain the convergence trajectory by recording the degradation path from the pre-trained teacher model to the untrained student generator. The degradation path implicitly represents the intermediate distributions between the teacher and the student, and its reverse can be viewed as the convergence trajectory from the student generator to the teacher model. Then $\textit{Distribution Backtracking}$ trains the student generator to backtrack the intermediate distributions along the path to approximate the convergence trajectory of the teacher model. Extensive experiments show that DisBack achieves faster and better convergence than the existing distillation method and achieves comparable or better generation performance, with an FID score of 1.38 on the ImageNet 64$\times$64 dataset. DisBack is easy to implement and can be generalized to existing distillation methods to boost performance.
Shengyuan Zhang, Ling Yang 0006, Zejian Li, Chenye Meng, Chang-yuan Yang, Guang Yang 0022, Lingyun Sun
ICLR6
2025 Inversion-DPO: Precise and Efficient Post-Training for Diffusion Models
abstract
Recent advancements in diffusion models (DMs) have been propelled by alignment methods that post-train models to better conform to human preferences. However, these approaches typically require computation-intensive training of a base model and a reward model, which not only incurs substantial computational overhead but may also compromise model accuracy and training efficiency. To address these limitations, we propose Inversion-DPO, a novel alignment framework that circumvents reward modeling by reformulating Direct Preference Optimization (DPO) with DDIM inversion for DMs. Our method conducts intractable posterior sampling in Diffusion-DPO with the deterministic inversion from winning and losing samples to noise and thus derive a new post-training paradigm. This paradigm eliminates the need for auxiliary reward models or inaccurate appromixation, significantly enhancing both precision and efficiency of training. We apply Inversion-DPO to a basic task of text-to-image generation and a challenging task of compositional image generation. Extensive experiments show substantial performance improvements achieved by Inversion-DPO compared to existing post-training methods and highlight the ability of the trained generative models to generate high-fidelity compositionally coherent images. For the post-training of compostitional image geneation, we curate a paired dataset consisting of 11,140 images with complex structural annotations and comprehensive scores, designed to enhance the compositional capabilities of generative models. Inversion-DPO explores a new avenue for efficient, high-precision alignment in diffusion models, advancing their applicability to complex realistic generation tasks. Our code is available at https://github.com/MIGHTYEZ/Inversion-DPO
Zejian Li, Yize Li 0001, Chenye Meng, Zhongni Liu, Ling Yang 0006, Shengyuan Zhang, Guang Yang 0022, Chang-yuan Yang, Lingyun Sun
ACM Multimedia8
2025 CONDA: Introducing Context-Aware Decision Making Assistant in Virtual Reality for Interior Renovation
abstract
Customized interiors enhance quality of life and self-expression, driving demand for VR-based design solutions. However, scant research exists on exploiting contextual cues in VR to aid decision making. Consequently, we propose CONDA, a context-aware assistant which leveraging LLMs to support interior renovation decisions. Specifically, we reconstruct users’ homes in VR and provide CONDA with stylistic details and spatial layouts, allowing it to predict furniture labels based on the decision scenario. Besides, we devise various modes to comprehensively express users’ purchasing preferences. Finally, CONDA recommend compatible items based on the label matching algorithm, and generate multi-dimensional explanations. A 30-user study reveals contextual completeness and preference diversity critically influence recommendation quality and decision behaviors, with 90% praising CONDA’s performance and all expressing daily-use intent. Overall, we validated the efficacy and practicality of CONDA, deriving universal design insights for VR decision-support systems and establishing new research directions.CCS ConceptsHuman-centered computing → Virtual realityComputing methodologies → Natural language generationApplied computing → Computer-aided design
Yizhan Shao, Weitao You, Ziqing Zheng, Yinyu Lu, Chang-yuan Yang, Zhibin Zhou 0002
Int. J. Hum. Comput. Interact.5
2025 PaRUS: A Virtual Reality Shopping Method Focusing on Contextual Information between Products and Real Usage Scenes
abstract
The development of AR and VR technologies is enhancing users' online shopping experiences in various ways. However, in existing VR shopping applications, shopping contexts merely refer to the products and virtual malls or metaphorical scenes where users select products. This leads to the defect that users can only imagine rather than intuitively feel whether the selected products are suitable for their real usage scenes, resulting in a significant discrepancy between their expectations before and after the purchase. To address this issue, we propose PaRUS, a VR shopping approach that focuses on the context between products and their real usage scenes. PaRUS begins by rebuilding the virtual scenario of the products' real usage scene through a new semantic scene reconstruction pipeline (manual operation needed), which preserves both the structured scene and textured object models in the scene. Afterwards, intuitive visualization of how the selected products fit the reconstructed virtual scene is provided. We conducted two user studies to evaluate how PaRUS impacts user experience, behavior, and satisfaction with their purchase. The results indicated that PaRUS significantly reduced the perceived performance risk and improved users' trust and expectation with their results of purchase.
Yinyu Lu, Weitao You, Ziqing Zheng, Yizhan Shao, Chang-yuan Yang, Zhibin Zhou 0002
IEEE Trans. Vis. Comput. Graph.5
2024 Reducing Spatial Fitting Error in Distillation of Denoising Diffusion Models
abstract
Denoising Diffusion models have exhibited remarkable capabilities in image generation. However, generating high-quality samples requires a large number of iterations. Knowledge distillation for diffusion models is an effective method to address this limitation with a shortened sampling process but causes degraded generative quality. Based on our analysis with bias-variance decomposition and experimental observations, we attribute the degradation to the spatial fitting error occurring in the training of both the teacher and student model in the distillation. Accordingly, we propose Spatial Fitting-Error Reduction Distillation model (SFERD). SFERD utilizes attention guidance from the teacher model and a designed semantic gradient predictor to reduce the student's fitting error. Empirically, our proposed model facilitates high-quality sample generation in a few function evaluations. We achieve an FID of 5.31 on CIFAR-10 and 9.39 on ImageNet 64x64 with only one step, outperforming existing diffusion methods. Our study provides a new perspective on diffusion distillation by highlighting the intrinsic denoising ability of models.
Shengzhe Zhou, Zejian Li, Shengyuan Zhang, Lefan Hou, Chang-yuan Yang, Guang Yang 0022, Lingyun Sun
AAAI5
2024 Designing the Conversational Agent: Asking Follow-up Questions for Information Elicitation
abstract
Conversational Agents (CAs) can facilitate information elicitation in various scenarios, such as semi-structured interviews. Current CAs can ask predetermined questions but lack skills for asking follow-up questions. Thus, we designed three approaches for CAs to automatically ask follow-up questions, i.e., follow-ups on concepts, follow-ups on related concepts, and general follow-ups. To investigate their effects, we conducted a user study (N=26) in which a CA interviewer asked follow-up questions generated by algorithms and crafted by human wizards. Our results showed that the CA's follow-up questions were readable and effective in information elicitation. The follow-ups on concepts and related concepts achieved a lower drop rate and better relevance, while the general follow-ups elicited more informative responses. Further qualitative analysis of the human-CA interview data revealed algorithm drawbacks and identified follow-up question techniques used by the human wizards. We provided design implications for improving information elicitation of future CAs based on the results.
Jiaxiong Hu, Jingya Guo, Ningjing Tang, Xiaojuan Ma, Chang-yuan Yang, Ying-Qing Xu
Proc. ACM Hum. Comput. Interact.6
2024 Automatic Generation of Interactive Nonlinear Video for Online Apparel Shopping Navigation
abstract
We present an automatic generation pipeline of interactive nonlinear video for online apparel shopping navigation. Our approach was inspired by Google's “Messy Middle” theory, which suggests that people mentally are faced with two tasks—exploration and evaluation—before purchasing online. Given a set of apparel product presentation videos, our navigation UI organizes them to optimize users' product exploration and automatically generates interactive videos for users' product evaluation. To support automatic methods, we proposed a video clustering similarity ($\operatorname{CSIM}$) and a camera movement similarity ($\operatorname{MSIM}$), as well as a comparative video generation algorithm for product recommendation, presentation, and comparison. To evaluate our pipeline's effectiveness, we conducted several user studies. The results showed that our pipeline can help users complete the consumption process more efficiently, making it easier for them to understand and choose a product.
Weitao You, Juntao Ji, Lingyun Sun, Chang-yuan Yang, Mi Yu, Shi Chen 0005
IEEE Trans. Multim.4
2023 "I Never Envy Anyone, for I Have Already Built a Kingdom With My Fingertips": Exploring Teenagers' Experience in Chat-based Cosplay Community
abstract
This paper reports an interview study about the practice of teenagers’ chat-based cosplay in China. Findings reveal the four primary motivations of the participants and their main practice in chat-based cosplay. We found that adolescents perceived character presentation and portrayal as a central aspect of chat-based cosplay and they devoted significant effort to refine their characters to achieve higher character consistency. We highlighted the positive feedback loop between social relationships and story creation in chat-based cosplay community. In addition, we identified the influence and negative experiences on adolescents in the chat-based cosplay community.
Yaohua Bu, Suqi Lou, Shi Chen 0005, Lingyun Sun, Chang-yuan Yang
IDC6
2023 Learning Object Consistency and Interaction in Image Generation from Scene Graphs
abstract
This paper is concerned with synthesizing images conditioned on a scene graph (SG), a set of object nodes and their edges of interactive relations. We divide existing works into image-oriented and code-oriented methods. In our analysis, the image-oriented methods do not consider object interaction in spatial hidden feature. On the other hand, in empirical study, the code-oriented methods lose object consistency as their generated images miss certain objects in the input scene graph. To alleviate these two issues, we propose Learning Object Consistency and Interaction (LOCI). To preserve object consistency, we design a consistency module with a weighted augmentation strategy for objects easy to be ignored and a matching loss between scene graphs and image codes. To learn object interaction, we design an interaction module consisting of three kinds of message propagation between the input scene graph and the learned image code. Experiments on COCO-stuff and Visual Genome datasets show our proposed method alleviates the ignorance of objects and outperforms the state-of-the-art on visual fidelity of generated images and objects.
Yangkang Zhang, Chenye Meng, Zejian Li, Pei Chen 0005, Guang Yang 0022, Chang-yuan Yang, Lingyun Sun
IJCAI6
2023 Cultural Self-Adaptive Multimodal Gesture Generation Based on Multiple Culture Gesture Dataset
abstract
Co-speech gesture generation is essential for multimodal chatbots and agents. Previous research extensively studies the relationship between text, audio, and gesture. Meanwhile, to enhance cross-culture communication, culture-specific gestures are crucial for chatbots to learn cultural differences and incorporate cultural cues. However, culture-specific gesture generation faces two challenges: lack of large-scale, high-quality gesture datasets that include diverse cultural groups, and lack of generalization across different cultures. Therefore, in this paper, we first introduce a Multiple Culture Gesture Dataset (MCGD), the largest freely available gesture dataset to date. It consists of ten different cultures, over 200 speakers, and 10,000 segmented sequences. We further propose a Cultural Self-adaptive Gesture Generation Network (CSGN) that takes multimodal relationships into consideration while generating gestures using a cascade architecture and learnable dynamic weight. The CSGN adaptively generates gestures with different cultural characteristics without the need to retrain a new network. It extracts cultural features from the multimodal inputs or a cultural style embedding space with a designated culture. We broadly evaluate our method across four large-scale benchmark datasets. Empirical results show that our method achieves multiple cultural gesture generation and improves comprehensiveness of multimodal inputs. Our method improves the state-of-the-art average FGD from 53.7 to 48.0 and culture deception rate (CDR) from 33.63% to 39.87%.
Jingyu Wu, Shi Chen 0005, Shuyu Gan, Chang-yuan Yang, Lingyun Sun
ACM Multimedia5
2023 Robust discriminant latent variable manifold learning for rotating machinery fault diagnosis
Chang-yuan Yang, Qinkai Han
Eng. Appl. Artif. Intell.1
2023 What makes virtual intimacy...intimate? Understanding the Phenomenon and Practice of Computer-Mediated Paid Companionship
abstract
Virtual romance service (VRS), as a notable commodification of intimacy, is currently emerging in China. Such service is not similar to the kind of intimacy that fans and idols generate through parasocial relationships, but behaves as the direct dyadic intimacy between service providers (virtual lovers) and buyers (customers). To gain a deep understanding of computer-mediated paid companionship, we study emerging user behaviors in VRS through a mixed-method study, including a survey (N = 178) and a follow-up semi-structured interview (N = 22) with both virtual lovers and customers to learn about their motivations, perceptions, and how virtual lovers provide online paid companionship to meet customers' emotional needs. We found three behavioral strategies of virtual lovers and the fact that they provide service in surface and deep acting and real feeling. Customers see VRS as a way to obtain affective benefits with reduced affective cost. We also found that VRS customers paid for the tangible benefits of an idealized romantic partner, rather than long-term commitment and emotional investment, and we identified key characteristics that VRS reduces from intimate relationships that fit its pay-per-use feature. We conclude by discussing the nature of virtual lovers and design implications for computer-mediated paid companionship.
Shi Chen 0005, Lingyun Sun, Chang-yuan Yang
Proc. ACM Hum. Comput. Interact.4
2021 Shing: A Conversational Agent to Alert Customers of Suspected Online-payment Fraud with Empathetical Communication Skills
abstract
Alerting customers on suspected online-payment fraud and persuade them to terminate transactions is increasingly requested with the rapid growth of digital finance worldwide. We explored the feasibility of using a conversational agent (CA) to fulfill this request. Shing, a voice-based CA, proactively initializes and repairs the conversation with empathetical communication skills in order to alert customers when a suspected online-payment fraud is detected, collects important information for fraud scrutiny and persuades customers to terminate the transaction once the fraud is confirmed. We evaluated our system by comparing it with a rule-based CA with regards to customer response and perceptions in a real-world context where our systems took 144,795 phone calls in total in which 83,019 (57.3%) natural breakdowns happened. Results showed that more customers stopped risky transactions after conversing with Shing. They seemed more willing to converse with Shing for more dialogue turns and provide transaction details. Our work presents practical implications for the design of proactive CA.
Jingya Guo, Jiajing Guo, Chang-yuan Yang, Yanjing Wu, Lingyun Sun
CHI3
2020 Automatic synthesis of advertising images according to a specified style
abstract
Images are widely used by companies to advertise their products and promote awareness of their brands. The automatic synthesis of advertising images is challenging because the advertising message must be clearly conveyed while complying with the style required for the product, brand, or target audience. In this study, we proposed a data-driven method to capture individual design attributes and the relationships between elements in advertising images with the aim of automatically synthesizing the input of elements into an advertising image according to a specified style. To achieve this multi-format advertisement design, we created a dataset containing 13 280 advertising images with rich annotations that encompassed the outlines and colors of the elements, in addition to the classes and goals of the advertisements. Using our probabilistic models, users guided the style of synthesized advertisements via additional constraints (e.g., context-based keywords). We applied our method to a variety of design tasks, and the results were evaluated in several perceptual studies, which showed that our method improved users’ satisfaction by 7.1% compared to designs generated by nonprofessional students, and that more users preferred the coloring results of our designs to those generated by the color harmony model and Colormind.
Weitao You, Hao Jiang 0046, Zhi-Yuan Yang, Chang-yuan Yang, Lingyun Sun
Frontiers Inf. Technol. Electron. Eng.4
2020 A music-driven system for generating apparel display video
Hui Zhang 0064, Yingping Cao, Xiaoyi Huang, Chang-yuan Yang, Lingyun Sun
Multim. Tools Appl.6
2019 Intelligent design of multimedia content in Alibaba
abstract
Multimedia content is an integral part of Alibaba’s business ecosystem and is in great demand. The production of multimedia content usually requires high technology and much money. With the rapid development of artificial intelligence (AI) technology in recent years, to meet the design requirements of multimedia content, many AI auxiliary tools for the production of multimedia content have emerged and become more and more widely used in Alibaba’s business ecology. Related applications include mainly auxiliary design, graphic design, video generation, and page production. In this report, a general pipeline of the AI auxiliary tools is introduced. Four representative tools applied in the Alibaba Group are presented for the applications mentioned above. The value brought by multimedia content design combined with AI technology has been well verified in business through these tools. This reflects the great role played by AI technology in promoting the production of multimedia content. The application prospects of the combination of multimedia content design and AI are also indicated.
Kuilong Liu, Chang-yuan Yang, Guang Yang 0022
Frontiers Inf. Technol. Electron. Eng.3
2018 The PMEmo Dataset for Music Emotion Recognition
abstract
Music Emotion Recognition (MER) has recently received considerable attention. To support the MER research which requires large music content libraries, we present the PMEmo dataset containing emotion annotations of 794 songs as well as the simultaneous electrodermal activity (EDA) signals. A Music Emotion Experiment was well-designed for collecting the affective-annotated music corpus of high quality, which recruited 457 subjects.
Hui Zhang 0064, Chang-yuan Yang, Lingyun Sun
ICMR4
2018 Crowdsourcing intelligent design
abstract
Design intelligence, namely, artificial intelligence to solve creative problems and produce creative ideas, has improved rapidly with the new generation artificial intelligence. However, existing methods are more skillful in learning from data and have limitations in creating original ideas different from the training data. Crowdsourcing offers a promising method to produce creative designs by combining human inspiration and machines’ computational ability. We propose a crowdsourcing intelligent design method called ‘flexible crowdsourcing design’. Design ideas produced through crowdsourcing design can be unreliable and inconsistent because they rely solely on selection among participants’ submissions of ideas. In contrast, the flexible crowdsourcing design method employs a cultivation procedure that integrates the ideas from crowd participants and cultivates these ideas to improve design quality at the same time. We introduce a series of studies to show how flexible crowdsourcing design can produce original design ideas consistently. Specifically, we will describe the typical procedure of flexible crowdsourcing design, the refined crowdsourcing tasks, the factors that affect the idea development process, the method for calculating idea development potential, and two applications of the flexible crowdsourcing design method. Finally, it summarizes the design capabilities enabled by crowdsourcing intelligent design. This method enhances the performance of crowdsourcing design and supports the development of design intelligence.
Wei Xiang 0008, Lingyun Sun, Weitao You, Chang-yuan Yang
Frontiers Inf. Technol. Electron. Eng.4